VM Unresponsive Due to Datastore Connectivity and I/O Stalls
search cancel

VM Unresponsive Due to Datastore Connectivity and I/O Stalls

book

Article ID: 410368

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

Symptoms:

  • VM become unresponsive and appear hung.

  • vmware.log confirms a hard CPU reset due to prolonged I/O stalls:

2###-0#-0#T10:32:01.461Z In(05) vcpu-0 - Checkpoint_Unstun: vm stopped for 6948 us
2###-0#-0#T10:32:01.461Z In(05) vcpu-0 - CPU reset: hard (mode Emulation)
2###-0#-0#T10:32:01.461Z In(05) vcpu-1 - CPU reset: hard (mode Emulation)

  • In the var/run/log/hostd.log file, similar entries are seen:

    2###-0#-0#T10:00:04.273Z info hostd[4277801] [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 350617 : Lost access to volume 5f3a7daf-########-####-0090fada2b68 (Datastore1) due to connectivity issues. Recovery attempt is in progress and outcome will be reported shortly.

Environment

  • VMware ESXi 8.x
  • VMware ESXi 9.x

Cause

The VM hung due to temporary instability in the storage path. This sudden drop in connection triggered a domino effect at the storage layer:

  • Commands started to abort and retry.
  • VMFS heartbeat timeouts occurred, meaning the host temporarily lost access to the datastore.
  • The system logged anomalies like invalid Remote Port Index (RPI) retries and failed command aborts.
  • The Host Bus Adapter (HBA) driver began canceling commands at the link or session level.

Because storage I/O was stalled for an extended period, the delay directly impacted the VM. As a result, the ESXi scheduler forced a CPU hard reset on the VM to attempt recovery.

Validation:

You can confirm this specific issue by checking the ESXi host logs for the patterns below.

1. VMFS Heartbeat Timeouts (/var/log/vobd.log)

Check the vobd.log file to verify if the datastore lost access and then recovered. You should see an esx.problem.vmfs.heartbeat.timedout event followed shortly by a vob.vmfs.heartbeat.recovered event.

2###-0#-0#T10:00:04.272Z: [vmfsCorrelator] 52944128166192us: [esx.problem.vmfs.heartbeat.timedout] 5f3a7daf-########-####-0090fada2b68 Datastore1
2###-0#-0#T10:00:04.278Z: [vmfsCorrelator] 52943286477311us: [vob.vmfs.heartbeat.recovered] Reclaimed heartbeat for volume 5f3a7daf-########-####-0090fada2b68 (Datastore1): [Timeout] [HB state abcdef02 offset 4050944 gen 46153 stampUS 52943286477004 uuid 6596cc43-########-####-0090fada2b28 jrnl <FB 6792000> drv 14.81]

2. SCSI Error Codes and Stuck I/O (/var/log/vmkernel.log)

In the vmkernel.log file, look for read and write commands failing with specific SCSI host error codes (H:0x2, H:0x8, H:0xc, D:0x8). These errors indicate the datastore became unusable due to timeouts and stuck I/O.

2###-0#-0#T10:00:01.902Z cpu41:2098225)WARNING: NMP: nmp_DeviceRequestFastDeviceProbe:237: NMP device "naa.###" state in doubt; requested fast path state update...
2###-0#-0#T10:00:04.272Z cpu67:2098227)NMP: nmp_ThrottleLogForDevice:3867: Cmd 0x89 (0x45d92bc98388, 21894817) to dev "naa.###" on path "vmhba2:C0:T2:L0" Failed:
2###-0#-0#T10:00:04.272Z cpu67:2098227)NMP: nmp_ThrottleLogForDevice:3875: H:0x8 D:0x0 P:0x0 . Act:EVAL. cmdId.initiator=0x430794e64800 CmdSN 0x27d5a72
2###-0#-0#T10:15:53.511Z cpu49:2098225)WARNING: NMP: nmp_DeviceRequestFastDeviceProbe:237: NMP device "naa.###" state in doubt; requested fast path state update...
2###-0#-0#T10:15:53.511Z cpu49:2098225)ScsiDeviceIO: 4124: Cmd(0x45d91b3a9fc8) 0x2a, CmdSN 0x359 from world 21827278 to dev "naa.####" failed H:0x2 D:0x8 P:0x0
2###-0#-0#T10:15:53.525Z cpu49:2098225)NMP: nmp_ThrottleLogForDevice:3867: Cmd 0x2a (0x45d91b3a9fc8, 21827278) to dev "naa.####" on path "vmhba2:C0:T2:L1" Failed:
2###-0#-0#T10:15:53.525Z cpu49:2098225)NMP: nmp_ThrottleLogForDevice:3875: H:0xc D:0x0 P:0x0 . Act:NONE. cmdId.initiator=0x430bf5453780 CmdSN 0x359

3. Fabric Anomalies and Invalid RPI (/var/log/vmkernel.log)

Further down in the vmkernel.log, check for abort failures and Invalid RPI messages. This confirms an HBA driver command failed because the remote port connection state became invalid during the fabric disruption.

2###-0#-0#T10:00:04.480Z cpu29:2097576)brcmfcoe: lpfc_handle_status:5079: 3:(0):3271: FCP cmd x89 failed <2/0> sid x011852, did x012c01, oxid xa13 iotag x54c SCSI Chk Cond - 0xe: Data(x2:xe:x1d:x0)
2###-0#-0#T10:00:08.327Z cpu59:21888071)WARNING: brcmfcoe: lpfc_sli_issue_abort:10358: 1:(0):3169 Abort failed: Abort INP: Data: xb43 x67c x1004 x98
2###-0#-0#T10:00:08.327Z cpu3:2098115)brcmfcoe: lpfc_handle_status:5079: 1:(0):3271: FCP cmd x89 failed <0/0> sid x011801, did x012f01, oxid xb43 iotag x67c Abort Requested Host Abort Req
2###-0#-0#T10:15:53.511Z cpu3:2546250)brcmfcoe: lpfc_handle_status:5079: 1:(0):3271: FCP cmd x2a failed <2/1> sid x011801, did x012c01, oxid x921 iotag x45a Invalid RPI Host Retry
2###-0#-0#T10:48:00.231Z cpu21:2098115)brcmfcoe: lpfc_els_unsol_buffer:7818: 1:(0):3717 LOGO received from NPORT x12c01 state x7 Data: x20 x800220 x11801 x11801
2026-05-16T05:50:34.391Z In(182) vmkernel: cpu1:2097863)ScsiDeviceIO: 13405: Task mgmt request issued to device naa.##### is stuck (WorldID 2097256, Cmd 0xfe, CmdSN 2048d6). Issuing yellow notification to the application

Resolution

To resolve this issue, engage your Storage and Fabric vendors.

Additional Information