When running latency-sensitive, high-frequency write/rewrite or UNMAP space reclamation workloads (such as trading session persistence, database logging, or rapid temporary file overwrites) on a vSAN Express Storage Architecture (ESA) cluster with Deduplication & Compression enabled, virtual machine disks experience transient latency spikes.
Virtual machine disk read and write latencies spiking between 400 ms and over 2,000 ms (2+ seconds).
Performance degradation correlating directly with high-frequency, small-block, latency-sensitive write/rewrite workloads running concurrently with active TRIM/UNMAP space reclamation.
vSAN Express Storage Architecture (ESA) with Deduplication & Compression enabled
A background space-cleanup process in vSAN ESA stalls during file deletion operations, causing incoming compute write operations to queue and temporarily freeze while waiting for space reclamation processing.
Broadcom is working on a fix that will be included in an upcoming release. Please contact Broadcom support for further investigation and assistance.
To stabilize virtual machine write latency and prevent guest I/O from being impacted by UNMAP throttling, reduce the UNMAP throttle timeout parameters to 1 millisecond across all hosts in the vSAN ESA cluster.
Run the following commands via SSH
esxcfg-advcfg -s 1 /VSAN/zDOMUnmapTimeoutMsec
esxcfg-advcfg -s 1 /VSAN/zDOMUnmapMaxTimeoutMsec
Subscribe to this Article for updates.