vSAN Storage I/O Errors and Cluster Partition During Switch Maintenance
search cancel

vSAN Storage I/O Errors and Cluster Partition During Switch Maintenance

book

Article ID: 443458

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

Following a planned Top-of-Rack (TOR) switch firmware update, the following symptoms were observed in the vSAN environment:

  • Data Unavailability: Virtual Machines (VMs) running on the vSAN datastore experienced significant I/O errors and became unresponsive for approximately 4 minutes.

  • Health Alarms: vSAN Health Service reported "Network Partition" and "Unicast Connectivity" failures.

  • Host Isolation: In the vsansystem.log, the node count dropped significantly (e.g., from 7 nodes to 1 node).

Environment

  • VMware vSAN 8.x

  • VMware vSAN 9.x

  • Network NIC Teaming and Failover > Failback is set it to Yes

  • TOR switch Activity (Firmware or OS upgrade).

Cause

The issue was caused by the vSAN network teaming policy being set to "Failback: Yes".

  • When the primary switch (SW-01) completed its firmware reboot, its physical ports transitioned to an "Up" state before the switch control plane was fully initialized or the configuration (VLANs, MTU) was loaded. Because "Failback" was enabled, the ESXi hosts immediately moved vSAN traffic from the healthy secondary switch back to the unconfigured ports on the primary switch. This resulted in a network partition as the hosts could no longer communicate with each other over the vSAN network.

Resolution

To prevent network partitions during future maintenance windows, implement the following changes:

NOTICE: Modify the NIC teaming settings to desired once the activity is completed.

  • On the physical switch layer, keep the ports in a shutdown state during the activity. Bring up the ports port validating the switch availability and healthy configuration.