High write latency and packet loss on vSAN ESA stretched clusters
search cancel

High write latency and packet loss on vSAN ESA stretched clusters

book

Article ID: 451292

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

Virtual machines running on vSAN Express Storage Architecture (ESA) stretched clusters may experience high storage write latency and intermittent stuns during write-intensive operations. 

Environment

VMware vSAN Express Storage Architecture (ESA)

Stretched Cluster Topology

Cause

Write amplification generated by RAID-5 Erasure Coding with Secondary Failures to Tolerate (SFTT=1) can saturate Inter-Switch Link (ISL) network bandwidth across stretched sites. This saturation leads to physical packet drops on the ISL, triggering frequent TCP Selective Acknowledgment (SACK) retransmissions.

Resolution

  • Network Infrastructure: Verify that the physical network bandwidth available on the Inter-Switch Links (ISLs) is sufficient to handle the combined vSAN storage traffic and management overhead across sites. (50gb - 100gb uplinks are highly recommended.) 

  • Optimize vSAN Storage Policy:

    • Set Primary Failures to Tolerate (PFTT) = 1

    • Set Secondary Failures to Tolerate (SFTT) = 0

    • Note: This configuration reduces site-to-site write amplification across the stretched cluster.

  • Consider upgrading to 50-100 gb uplinks. (Can engage PSO to ensure uplinks are appropriate for cluster workload.)

Additional Information

Performance Recommendations for vSAN ESA

Physical NIC Requirements for vSAN

Best Practices for vSAN Networking