Following a planned Top-of-Rack (TOR) switch firmware update, the following symptoms were observed in the vSAN environment:
Data Unavailability: Virtual Machines (VMs) running on the vSAN datastore experienced significant I/O errors and became unresponsive for approximately 4 minutes.
Health Alarms: vSAN Health Service reported "Network Partition" and "Unicast Connectivity" failures.
Host Isolation: In the vsansystem.log, the node count dropped significantly (e.g., from 7 nodes to 1 node).
VMware vSAN 8.x
VMware vSAN 9.x
Network NIC Teaming and Failover > Failback is set it to Yes
TOR switch Activity (Firmware or OS upgrade).
The issue was caused by the vSAN network teaming policy being set to "Failback: Yes".
To prevent network partitions during future maintenance windows, implement the following changes:
Modify Teaming and Failover Policy
Change the vSAN Port Group "Failback" setting from Yes to No. This prevents ESXi from automatically returning traffic to a recovered uplink until a manual failback is performed by an administrator after verification.
NOTICE: Modify the NIC teaming settings to desired once the activity is completed.