Virtual machines residing on the affected host may become inaccessible or freeze.
Reviewing the vmkernel.log reveals Cisco nfnic HBA driver errors indicating an I/O data payload mismatch:
WARNING: nfnic: <2>: fnic_fcpio_icmnd_cmpl_handler: 2060: sc: 0x45d9fc63ba80 tag: 0x516 hdr status: FCPIO_DATA_CNT_MISMATCH IO failure!
Immediately following the HBA errors, the VMware Native Multipathing Plugin (NMP) reports the storage path state is in doubt and requests a fast path update:
WARNING: NMP: nmp_DeviceRequestFastDeviceProbe:235: NMP device "naa.########################" state in doubt; requested fast path state update...
The logs may also show preceding Permanent Device Loss (PDL) events for unrelated LUNs prior to the host lockup:
VMW_SATP_ALUA: satp_alua_issueCommandOnPath:1019: Path "vmhba1:C0:T46:L193" (PERM LOSS) command 0x12 failed with status Device is permanently unavailable.
VMware vSphere ESXi 8.0.x
This issue occurs due to underlying Layer 1 physical fabric degradation or storage array congestion. The Cisco nfnic adapter detects that the data payload received does not match the expected size and drops the I/O to protect data integrity. This results in a deadlock where management agents (hostd) block while waiting for storage commands to complete.
Perform a hardware-level hard reboot (cold boot) of the ESXi host to flush stalled I/O queues and restore management connectivity.
To prevent recurrence, investigate the physical storage network:
nfnic driver version is certified on the Cisco UCS HCL.If you need further assistance, see . Scroll to the bottom of the page and click on your respective region.