Unresponsive management tasks and FCPIO_DATA_CNT_MISMATCH errors on ESXi host
search cancel

Unresponsive management tasks and FCPIO_DATA_CNT_MISMATCH errors on ESXi host

book

Article ID: 442961

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

  • An ESXi host shows as Not Responding or disconnected in vCenter Server.
  • vMotion migrations to or from the host hang at 17%.
  • Virtual machines residing on the affected host may become inaccessible or freeze.

  • Reviewing the vmkernel.log reveals Cisco nfnic HBA driver errors indicating an I/O data payload mismatch:

    WARNING: nfnic: <2>: fnic_fcpio_icmnd_cmpl_handler: 2060: sc: 0x45d9fc63ba80 tag: 0x516 hdr status: FCPIO_DATA_CNT_MISMATCH IO failure!

  • Immediately following the HBA errors, the VMware Native Multipathing Plugin (NMP) reports the storage path state is in doubt and requests a fast path update:

    WARNING: NMP: nmp_DeviceRequestFastDeviceProbe:235: NMP device "naa.########################" state in doubt; requested fast path state update...

  • The logs may also show preceding Permanent Device Loss (PDL) events for unrelated LUNs prior to the host lockup:

    VMW_SATP_ALUA: satp_alua_issueCommandOnPath:1019: Path "vmhba1:C0:T46:L193" (PERM LOSS) command 0x12 failed with status Device is permanently unavailable.

     

Environment

VMware vSphere ESXi 8.0.x

Cause

This issue occurs due to underlying Layer 1 physical fabric degradation or storage array congestion. The Cisco nfnic adapter detects that the data payload received does not match the expected size and drops the I/O to protect data integrity. This results in a deadlock where management agents (hostd) block while waiting for storage commands to complete.

Resolution

Perform a hardware-level hard reboot (cold boot) of the ESXi host to flush stalled I/O queues and restore management connectivity.

To prevent recurrence, investigate the physical storage network:

  1. SAN Switches: Check for CRC errors, Invalid Tx Words, or Link Failures on ports connected to the host.
  2. Storage Array: Review controller logs for failovers or hardware faults.
  3. Firmware: Verify the combination of ESXi version, Cisco UCS firmware, and nfnic driver version is certified on the Cisco UCS HCL.

Additional Information

If you need further assistance, see Contact Support. Scroll to the bottom of the page and click on your respective region.