ESXi hosts experience storage latency or lost access to volumes during high I/O bursts (VDI Login Storms) - VMware vSphere ESXi
search cancel

ESXi hosts experience storage latency or lost access to volumes during high I/O bursts (VDI Login Storms) - VMware vSphere ESXi

book

Article ID: 396366

calendar_today

Updated On:

Products

VMware vSphere ESXi VMware vSphere ESX 7.x VMware vSphere ESX 8.x

Issue/Introduction

ESXi hosts intermittently report "Lost access to volume" messages, particularly during high-I/O operations such as Horizon VDI image recompose tasks, login storms, or backup windows.

  • Primary Indicators:
    • vSphere Client Alerts: Lost access to volume <DatastoreName> due to connectivity issues. Recovery attempt is in progress.
    • Virtual Machine Impact: Extreme sluggishness, extended login times (5–20+ minutes), or guest OS hangs/stuns.
    • Performance Deterioration: /var/log/vmkernel.log shows significant latency spikes:
       
      WARNING: ScsiDeviceIO: 1513: Device naa.### performance has deteriorated. I/O latency increased from average value of 666 microseconds to 826700 microseconds.
       
    • Driver Aborts: The Emulex Fibre Channel (lpfc) or QLogic (qlnativefc) drivers report command failures:
       
      lpfc: lpfc_handle_status:5637: 0:(0):3271: FCP cmd x89 failed <2/354> sid x521d03, did x520304, oxid x363 iotag x689 Abort Requested Host Abort Req
       
    • RC lock messages: In the vmkernel.log for the ESXi hosts connected to VMFS datastores.
       

       

       
      Res6: 2937: 'datastore': RC Lock not free for type 1, return TXN FULL
       
       
       
    • Heartbeat Timeouts: The /var/log/vobd.log confirms datastore heartbeat failures:
       
      HBX: 3063: '': HB at offset 3702784 - Waiting for timed out HB: [HB state abcdef02 offset 3702784 gen 261 stampUS 2658722121357]
       

Environment

VMware vSphere ESXi 7.x
VMware vSphere ESXi 8.x

Cause

This issue is typically caused by storage subsystem saturation. During burst I/O events (like VDI recompose), the storage array controllers or SAN fabric can exceed their available buffer credits or queue depths.

When the storage target fails to acknowledge I/O within the SCSI timeout period (typically 30–60 seconds), the ESXi host bus adapter (HBA) is forced to abort the stalled commands to free up queue slots. This "storage locking" prevents the management agents (hostd/vpxa) from completing tasks, leading to vCenter disconnections and VM unresponsiveness.

Resolution

To mitigate storage saturation and command aborts, implement the following optimizations:

1.Throttle Concurrent Operations

  • Horizon/VDI: Lower the maximum concurrent provisioning and power-state operations in the Horizon Console to reduce burst IOPS pressure on the array.
  • VAAI: Ensure VMware vStorage APIs for Array Integration (VAAI) are enabled to offload block cloning tasks directly to the storage hardware.

2. Optimize Multipathing and Queue Depth [Consult storage vendor for best practice]

  • Path Selection: Change the Path Selection Policy (PSP) for all production LUNs to Round Robin (VMW_PSP_RR).
  • I/O Limit: Set the Round Robin IOPS limit to 1 (default is 1000) to ensure more frequent path switching and better queue distribution: esxcli storage nmp psp round robin device config set --type=iops --iops=1 --device=naa.###
  • LUN Layout: Distribute virtual machines across multiple smaller LUNs rather than a single large volume to maximize queue parallelism across the SAN fabric.

3. Host and Hardware Configuration

  • Physical Layer Audit: Inspect the Fibre Channel fabric for physical errors. Check SFP modules, fiber cables, and switch ports for CRC errors or link resets that may exacerbate command timeouts.
  • Firmware/Driver Alignment: Ensure HBA firmware and ESXi drivers are aligned with the latest Broadcom Compatibility Guide (BCG) recommendations.

Additional Information