High VM latency, corrupted application data, or lost storage connection due to FC frame drops and high "Invalid Tx Word Count"
search cancel

High VM latency, corrupted application data, or lost storage connection due to FC frame drops and high "Invalid Tx Word Count"

book

Article ID: 413619

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

Fibre Channel (FC) storage adapters may report "Transient storage condition" alerts or lose connectivity to datastores.
VMs running on Fiber channel datastores have poor performance (high latency) on one or more ESXi hosts.

This occurs when the physical layer signaling is interrupted or frames are corrupted in transit. 
Symptoms can include the following: 

  • Application data on the VMs may be corrupted. 
  • VMs reporting poor performance  (high latency) on one or more ESXi hosts. 
  • ESXi host reports "Transient storage condition" in vmkernel.log.
  • SCSI commands fail with Host Status H:0x2 (dropped frames) and H:0x8 (driver aborts).
  • Datastores appear offline or unresponsive during I/O tasks.
  • Virtual machines may abruptly restart.
  • Lost access to volume events are observed on the "vCenter UI > Datastore > Monitor > Events" and "/var/run/log/hostd.log".

2026-07-14T16:30:21.004Z In(166) Hostd[2099635]: [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 218628 : Lost access to volume 69ed09bb-de78a555-2737-405b7fd831fe (IBOX3_EDC_ONT_PRODSTORE11) due to connectivity issues. Recovery attempt is in progress and outcome will be reported shortly.

  • Dead paths or unavailable targets are seen on the storage device.

esxcfg-mpath -b

naa.################################ : NFINIDAT Fibre Channel Disk (naa..################################)
vmhba4:C0:T0:L25 LUN:25 state:active fc Adapter: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:## Target: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:##
vmhba4:C0:T1:L25 LUN:25 state:active fc Adapter: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:## Target: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:##
vmhba3:C0:T1:L25 LUN:25 state:dead fc Adapter: Unavailable Target: Unavailable
vmhba3:C0:T0:L25 LUN:25 state:dead fc Adapter: Unavailable Target: Unavailable
vmhba1:C0:T3:L25 LUN:25 state:active fc Adapter: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:## Target: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:##
vmhba1:C0:T2:L25 LUN:25 state:active fc Adapter: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:## Target: WWNN: ##:##:##:##:##:##:##:## WWPN: ##:##:##:##:##:##:##:##

    • Dead paths are observed with SCSI sense code "H:0x1" which indicates "LUN disconnect".

    2026-07-14T16:29:06.312Z In(182) vmkernel: cpu50:2098305)lpfc: lpfc_rportStats:3118: 0:(0) Compression log for fcp target 5, path is dead, IO errors: busy 0,retry 120, no_connect 506282, fcperr 65, tmo 28
    2026-07-14T16:29:26.133Z In(182) vmkernel: cpu67:2098752)NMP: nmp_ThrottleLogForDevice:3893: Cmd 0x12 (0x45db56ae1c80, 0) to dev "naa.#######################"" on path "vmhba3:C0:T1:L25" Failed:
    2026-07-14T16:29:26.133Z In(182) vmkernel: cpu67:2098752)NMP: nmp_ThrottleLogForDevice:3898: H:0x1 D:0x0 P:0x0 . Act:NONE. cmdId.initiator=0x453b2769bb58 CmdSN 0x0SCSi commands failures are seen with sense code "H:0x0 D:0x28 P:0x0"

  • In /var/run/log/vmkernel.log , There scsi sense code " D:0x28" which means TASK_SET_FULL.

This status is returned when the LUN prevents accepting SCSI commands from initiators due to lack of resources, namely the queue depth on the array.

If configured, this code will activate when device status TASK SET FULL (0x28) is return for failed commands and essentially throttles back the I/O until the array stops returning this status.
These status codes may indicate congestion at the LUN level or at the port (or ports) on the array. When congestion is detected, VMkernel throttles the LUN queue depth(reduces to half of the original value). The VMkernel attempts to gradually restore the queue depth when congestion conditions subside.

2026-07-14T16:30:07.825Z In(182) vmkernel: cpu84:2098770)ScsiDeviceIO: 4644: Cmd(0x45db56bfaa80) 0x8a, CmdSN 0x364 from world 36106452 to dev "naa.#######################" failed H:0x0 D:0x28 P:0x0

Environment

VMware vSphere ESXi (all versions). 

Cause

This is caused to underlying SAN fabric issues i.e any physical layer instability in the SAN fabric.


Invalid Tx Word Count indicates a signaling error between the HBA and the switch, while Invalid CRC Count indicates hardware issues on SAN fabric layer.


These are typically caused by faulty SFPs, damaged FC cables, or loose connections.



Evidence of such issues may include:

  • Dropped frames, with repeated logging in /var/log/vmkernel.log similar to:

    qlnativefc: vmhba1(de:0.1): qlnativefcStatusEntry:1924:(3:8) Dropped frame(s) detected (37440 of 36864 bytes)

  • High "Invalid Tx Word Count" and "Invalid CRC errors reported on the vmhba adapters by running the command:

esxcli storage san fc stats get

Tx Frames: 0
Rx Frames: 0
Lip Count: 0
Error Frames: 3072
Dumped Frames: 0
Link Failure Count: 1
Loss of Signal Count: 0
PrimSeq Protocol Err Count: 0
Invalid Tx Word Count: 1066134556
Invalid CRC Count: 3072
Input Requests: 0
Output Requests: 0
Control Requests: 0

 

  • Aborted I/O due to slow I/O / command timeouts, e.g.:

/var/log/vmkernel.log includes logging similar to: 

VSCSI: 3285: handle 9366752458711311(GID:8463)(vscsi1:2):processing reset for handle ... state 1381192707
qlnativefc: vmhba1(de:0.1): qlnativefcTaskMgmt:2325:Task Mgmt virt reset
qlnativefc: vmhba1(de:0.1): qlnativefcEhVirtualReset:3239:C0:T0:L5: VIRTUAL RESET ISSUED.
qlnativefc: vmhba1(de:0.1): qlnativefcEhVirtualReset:3265:Command aborted on target=0x32x, lun=0x05 - SCSI command timeout counter incremented to 6081

  • /O aborts ( SCSI sense code H:0x5) and driver aborts (H:0x8) which are observed at the same time when lost access volume was reported . This indicates that due to underlying SAN fabric issue ,Storage is unable to handle the load leading to task full events and IO's not getting completed on time.

2026-07-14T16:30:21.002Z In(182) vmkernel: cpu42:2097336)ScsiDeviceIO: 4681: Cmd(0x45bb657c7440) 0x89, cmdId.initiator=0x430bb676d800 CmdSN 0x912b from world 2097288 to dev "naa.#######################" failed H:0x5 D:0x0 P:0x0 Cancelled from device layer.
2026-07-14T16:30:21.002Z In(182) vmkernel: cpu42:2097336)Cmd count Active:32 Queued:184
2026-07-14T16:30:21.002Z In(182) vmkernel: cpu42:2097336)ScsiDeviceIO: 4616: Cmd(0x45bb6564f240) 0x2a, cmdId.initiator=0x430bb676d800 CmdSN 0x912a from world 2097288 to dev "naa.#######################" failed H:0x8 D:0x0 P:0x0 Cancelled from device layer
2026-07-14T16:30:21.003Z In(182) vmkernel: cpu38:2098912)ScsiPath: 7084: Cancelled Cmd(0x45bb64b7b240) 0x2a, cmdId.initiator=0x430bb676d800 CmdSN 0x9128 from world 2097288 to path "vmhba3:C0:T5:L25". Cmd count Active:0 Queued:31.
2026-07-14T16:30:21.003Z In(182) vmkernel: cpu38:2098912)NMP: nmp_ThrottleLogForDevice:3893: Cmd 0x2a (0x45bb64b7b240, 2097288) to dev "naa.#######################" on path "vmhba3:C0:T5:L25" Failed:
2026-07-14T16:30:21.003Z Wa(180) vmkwarning: cpu38:2098912)WARNING: NMP: nmp_DeviceRequestFastDeviceProbe:235: NMP device "naa.#######################" state in doubt; requested fast path state update...

Resolution

Use the following steps to monitor HBA statistics and isolate the fault to the HBA, SFP, or cable:

  1. Log in to the ESXi shell via SSH.
  2. Run the following command to retrieve statistics for the affected adapter (replace vmhbaX with your adapter ID): 
    esxcli storage san fc stats get -A vmhbaX
  3. Review the FcStat output for the following counters:
    • Invalid Tx Word Count: Physical signaling/light interruption.
    • Invalid CRC Count: Frame data corruption.

Example output:

FcStat:
   Adapter: vmhbaX
   Tx Frames: #######
   Rx Frames: #######
   Lip Count: #
   Error Frames: ###### 
   Dumped Frames: #
   Link Failure Count: # 
   Loss of Signal Count: #
   PrimSeq Protocol Err Count: #
   Invalid Tx Word Count: ###  <===
   Invalid CRC Count: ####  <===
   Input Requests: #
   Output Requests: #
   Control Requests: #
  1. Run the command multiple times to determine if the counters are incrementing in real-time.
  2. If counters increment:
    • Inspect and reseat the FC cable.
    • Replace the SFP on the HBA or the connected switch port.
    • Verify HBA firmware is compliant with the VMware HCL.
  3. If the issue persists after hardware replacement, contact SAN switch or storage array vendor for upstream fabric diagnostics.