Fibre Channel HBA Volume Loss: ESXi 8.0 on HPE ProLiant Gen11
search cancel

Fibre Channel HBA Volume Loss: ESXi 8.0 on HPE ProLiant Gen11

book

Article ID: 449082

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

ESXi 8.0 hosts running on HPE ProLiant Gen11 hardware experience intermittent or total loss of access to VMFS datastores. This results in all virtual machines on the impacted host becoming unresponsive or powering off.

Symptoms:

  • vCenter Server reports the error: Lost access to volume <Datastore_Name> due to connectivity issues.
  • "All Paths Down (APD)" and "Heartbeat timeouts"

  • Virtual machines on the affected host are grayed out or inaccessible.

  • The host may recover only after a hard reboot.

  • Simultaneous failure of all virtual machines on a single host; peer cluster members remain healthy.

  • vmkernel logs show evidence of command queue saturation and transport failures:
ScsiDeviceIO: 4115: Cmd(0x45d988625dc8) 0x8a, CmdSN 0x800e000c from world 3129234 to dev "naa.####" failed H:0x0 D:0x8 P:0x0WARNING: NMP: nmp_DeviceRequestFastDeviceProbe:235: NMP device "naa.####" state in doubtDevice going All paths down

Environment

  • Product: VMware vSphere ESXi 8.0
  • Hardware: HPE ProLiant DL360 Gen11 (and other Gen11 models)
  • Protocol: Fibre Channel (FC)

Cause

The connectivity loss is caused by transport-layer instability at the Fibre Channel Host Bus Adapter (HBA) level. This is frequently associated with outdated system BIOS on HPE Gen11 platforms and physical layer errors (CRC errors) on the FC vmhba, leading to SCSI H:0x8 (Busy/Retry) status and eventual All Paths Down (APD) states.

Resolution

Perform the following steps to resolve and prevent recurrence:

  1. Update System Firmware: Update the HPE ProLiant Gen11 BIOS to the latest version (check Broadcom HCL) . This addresses known stability and power management issues in the Gen11 chipset.
  2. Verify Hardware Health: Check the FC HBA statistics for CRC errors or link failures. On the ESXi shell, use the following command to check for errors:

    "/usr/lib/vmware/vmkmgmt_nic/vmkmgmt_nic -sh vmhba<X>"

    If non-zero CRC errors are detected, replace the associated SFP modules and Fibre Channel cables.
  3. Update ESXi: Ensure the ESXi host is updated to the latest patch release. For download instructions, see Download Broadcom Products and Software.
  4. Storage Vendor Alignment: Contact your storage array vendor to investigate failed VAAI commands or significant VM stun operations (latency >15ms) identified in the host logs.

Additional Information

For assistance with log collection or further hardware analysis, see Contact Broadcom Support.