On FC fabric maintenance, multiple devices unexpectedly enter APD state
search cancel

On FC fabric maintenance, multiple devices unexpectedly enter APD state

book

Article ID: 446567

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

  • Planned fabric maintenance on half of a Fibre Channel fabric is expected to take down half of the ESXi paths to the FC storage.

  • This loss of path redundancy occurs for some devices on some ESXi hosts.

  • However, some devices unexpectedly enter an All Paths Down (APD) state.    

  • Virtual machines become unresponsive or experience Guest OS crashes during SAN maintenance (e.g., battery replacement or controller updates).
  • ESXi vmkernel.log contains the following log snippets:
      • WARNING: lpfc: lpfc_dev_loss_tmo_handler:537: vmhba3 #### Devloss timeout on WWPN ####
      • [esx.problem.storage.redundancy.degraded] Path redundancy to storage device #### degraded.
      • Device or filesystem with identifier [naa.####] has entered the All Paths Down state.

Environment

  • VMware vSphere ESXi (all versions)

Cause

  • Some devices only have paths via the half of the fabric, on which maintenace is being performed.

  • When the maintenance task brings these paths down, devices enter an APD state. 

Resolution

Prior to and during fabric maintenance, follow these steps to ensure path availability:

  1. Verify that each ESXi host has the expected number of paths to all storage devices.

  2. Audit the SAN zoning to ensure each Host Bus Adapter (HBA) is zoned to at least one port on both redundant Storage Processors.

  3. Verify that HBA 1 is connected to Fabric A and HBA 2 is connected to Fabric B.

  4. Confirm that the storage array ports are split across both fabrics to avoid a single point of failure.