/var/log/fdm.log on HA master node:Fdm[2108683] [Originator@6876 sub=Cluster opID=clusterManager.cpp:980-601c5b31] Marking slave host-####### as unreachable/var/log show a gap in logging)/var/log/vobd.log reports datastore heartbeat timeout, and then recovers heartbeat within one to two seconds:vobd[2098147] [vmfsCorrelator] 391182884031us: [vob.vmfs.heartbeat.timedout] <Datastore UUID> <Datastore Name>
vobd[2098147] [vmfsCorrelator] 391184193077us: [vob.vmfs.heartbeat.recovered] Reclaimed heartbeat for volume <Datastore UUID> (<Datastore Name>): [Timeout] [HB state abcdef02 offset 3837952 gen 81 stampUS 391184176867 uuid <UUID> jrnl <FB 25165830> drv 24.82]VMware vSphere ESXi (all versions)
HA failover occurs when an ESXi host becomes unresponsive. The heartbeat timeouts and lost access to datastores can be a cause of the host unresponsiveness, which subsequently triggers HA failover.
While underlying hardware issues (such as faulty components) are a common cause of this behavior, it is equally important to consider environmental factors. Transient network instability, such as network flapping or switch issues, can induce temporary loss of datastore connectivity (e.g., All Paths Down - APD conditions), leading to host unresponsiveness and triggering unnecessary vSphere HA recovery workflows.
Thorough troubleshooting should include both physical hardware diagnostics and an analysis of the network fabric connectivity to the storage to isolate the trigger of the connectivity loss.
/var/log/fdm.log, /var/log/vobd.log, and /var/log/vmkwarning.log) to correlate timestamps of HA failovers with storage or network events.