A VMware vSAN 8.0.x stretched cluster may experience a network partition and subsequent virtual machine (VM) restarts after migrating a Layer 3 (L3) Gateway to a new physical switch. This typically occurs when the primary site loses communication with the secondary site and the witness appliance.
vCenter reports a vSAN Cluster Partition health alarm
The /var/run/log/hostd.log on primary site ESXi hosts record Lost access to volume events.Event 2573287 : Lost access to volume 5#######-########-####-############ (########-####-####-####-############) due to connectivity issues. Recovery attempt is in progress and outcome will be reported shortly.
Storage accessibility status changes to False for virtual machines specifically on primary site hosts. Below events are recorded in /var/run/log/hostd.log:UpdateStorageAccessibilityStatusInt: Vm's storage accessibility status changed to false
Virtual machines on primary site hosts become invalid and unexpectedly fail over to secondary site hosts via vSphere HA.
The object or item referred to could not be found.VMware vSAN 8.0.x
vSAN Stretched Cluster configuration
A Layer 3 (L3) Gateway migration disconnected the primary site from both the secondary site and the witness appliance.
Network Partition: According to /var/run/log/vsansystem.log, a network partition occurred, dropping the primary site active nodeCount. Primary site hosts could only communicate locally.[vSAN@6876 sub=VsanSystemProvider opId=CMMDSMembershipUpdate-953e] Complete, nodeCount: 5, runtime info: (vim.vsan.host.VsanRuntimeInfo)
Loss of Quorum: In an 11-node configuration, the Primary site holds 5 out of 11 quorum votes. When communication to the other 6 votes (5 Secondary nodes + 1 Witness) is lost, the Primary site loses quorum.
Automatic Failover: The secondary site retained connectivity to the witness appliance, forming a 6-node majority. Because site mirroring was active, the secondary site maintained a complete copy of all data components alongside quorum. Consequently, storage became inaccessible on the primary site, marking local virtual machines as invalid and prompting vSphere HA to restart the workloads on the secondary site.
Coordinate with the network administration team to analyze physical switch logs. ESXi vmkernel interfaces register connectivity loss but cannot determine external root causes like routing convergence delays or dropped packets.
Verify routing and gateway configurations on the new switch to ensure vSAN traffic (vmk3 for data, vmk1 for witness) is properly routed between sites.
Schedule all L3 gateway alterations and major network routing shifts strictly within an approved maintenance window to prevent unmitigated production downtime.
For software downloads, see Download Broadcom products, patches and software
If the issue persists, contact Broadcom Support. See Contact Broadcom Support