VMs hung and unresponsive on ESXi host due to nfnic storage driver deadlock
search cancel

VMs hung and unresponsive on ESXi host due to nfnic storage driver deadlock

book

Article ID: 427482

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

  • During normal operations, all virtual machines on a single ESXi host become hung and unresponsive, requiring a forced hardware reboot of the host.
  • Logs similar to the following are observed in /var/run/log/vmkernel.log without preceding path loss events prior to link reset:

YYYY-MM-DDTHH:MM:SS.831Z In(182) vmkernel: cpu0:2098004)nfnic: <1>: INFO: fdls_tgt_send_adisc: 1316: sending ADISC to tgt: 0x10700
YYYY-MM-DDTHH:MM:SS.831Z In(182) vmkernel: cpu0:2098004)nfnic: <1>: INFO: fdls_tgt_send_adisc: 1316: sending ADISC to tgt: 0x10600
YYYY-MM-DDTHH:MM:SS.831Z In(182) vmkernel: cpu0:2098004)nfnic: <1>: INFO: fdls_process_gpn_ft_rsp: 2626: iport->state: 4
YYYY-MM-DDTHH:MM:SS.831Z In(182) vmkernel: cpu0:2098004)nfnic: <1>: INFO: fdls_process_tgt_adisc_rsp: 2310: ADISC accepted from target: 0x11601. TGT now in ready state. Target logged in
YYYY-MM-DDTHH:MM:SS.831Z In(182) vmkernel: cpu0:2098004)nfnic: <1>: INFO: fdls_process_tgt_adisc_rsp: 2310: ADISC accepted from target: 0x11801. TGT now in ready state. Target logged in
YYYY-MM-DDTHH:MM:SS.377Z In(182) vmkernel: cpu17:2098077)nfnic: <2>: INFO: fnic_abort_cmd: 3862: Abort cmd called for Tag: 0x263 issued time: 15309 ms CMD_STATE: FNIC_IOREQ_ABTS_PENDING CDB Opcode: 0x2a sc:0x45da4735d2c0 flags: 0x43 lun: 246 target: 0x10200
YYYY-MM-DDTHH:MM:SS.377Z Wa(180) vmkwarning: cpu17:2098077)WARNING: nfnic: <2>: fnic_ab

Environment

  • VMware vSphere ESXi (all versions)
  • Cisco UCS with Cisco nfnic driver releases prior to version 5.0.0.48

Cause

A storage driver deadlock occurs when the Cisco nfnic driver resets links while I/O commands are in flight, triggering a new fabric login. When the driver attempts to abort timed-out commands and receives no response, it re-issues the abort request. This double abort condition results in an unrecoverable driver deadlock.

Resolution

  1. Verify the installed nfnic driver version on the affected ESXi host using ESXCLI or vSphere Lifecycle Manager.
  2. Upgrade the Cisco nfnic driver to version 5.0.0.48 or later.
  3. Contact Cisco Support to obtain the recommended hardware firmware and driver combination for your specific Cisco UCS blade/rack server model.

Additional Information

Please refer the following: