High vSAN RT Latency and intermittent VM Inaccessibility
search cancel

High vSAN RT Latency and intermittent VM Inaccessibility

book

Article ID: 447502

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

  • The vSAN cluster frequently destabilizes, resulting in low health scores and repeated partitioning.
  • Virtual machines become inaccessible, and data migration/vMotion fails, preventing host maintenance.
  • Logs reveal massive communication delays (Round-Trip latency).
  • The cluster repeatedly drops members and partitions, despite no changes being made to the ESXi host configurations.
  • Rebooting the ESXi hosts does not result in any improvements.
  • Manually toggling or swapping the vmnics in use resolves the issue.

Environment

VMware vSAN all versions.

Cause

The issue is in the upstream physical network layer.

Because the state is held upstream, rebooting the ESXi hosts fails to resolve the issue, as the reboot does not force the physical switch to clear the specific network state associated with the uplinks.

Resolution

The network layer needs to be investigated.

Refreshing the physical link (Link-Down/Link-Up) resolves issues where the upstream network has cached a bad state or encountered a buffer/protocol hang that is not sensitive to host-side software restarts.