ESXi Host Disconnects from vCenter Due to AHCI Controller and SATADOM Communication Failure
search cancel

ESXi Host Disconnects from vCenter Due to AHCI Controller and SATADOM Communication Failure

book

Article ID: 453205

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

 

1. The ESXi host experiences a severe hardware communication failure between the AHCI storage controller and the local boot device.

2. The High-Performance Plug-in (HPP) storage path becomes trapped in a permanent failover state (failoverState=2 blocked=1).

3. HPP repeatedly attempts to reconnect to the SuperMicro SATADOM device, identified in logs as t10.ATA_____SMC_VD....

4. This hardware condition typically triggers an All Paths Down (APD) state for the boot volume.

5. Consequently, the host disconnects from vCenter and virtual machines become unresponsive.

 

Environment

VMware ESXi 8.x

Cause

 

VMkernel.log:

2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)HPP: HppAttemptFailoverRequest:1071: Re-issuing first command for HPP device "t10.ATA_____SMC_VD__________________________________d9a##############000000" (NO_CONNECT_ON_APD = CLEAR)  failoverState=2 blocked=1
2026-08-15T15:08:03.825Z Wa(180) vmkwarning: cpu42:2098695)WARNING: vmw_ahci[0000af000]:<0> IssueCommand:ERROR: Tag 1 SActive already set: SACT:3fe CI:3fe activeTags:0 reissue_flag:0
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)Backtrace for current CPU #42, worldID=2098695, fp=0x4322b02647f0
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bac0:[0x42002ad9bef7]vmk_LogBacktraceMessage@vmkernel#nover+0xb stack: 0x45bad2bffa98, 0x4322b02648b8, 0x4322b02648b8, 0xff
ffffffffffffff, 0x200000000
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bad0:[0x42002bcdc95c]ahciRequestIo@(vmw_ahci)#<None>+0x719 stack: 0x4322b02648b8, 0xffffffffffffffff, 0x200000000, 0x420000000012, 0xa3804322b0262a20
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bb50:[0x42002bce22e6]scsiExecReadWriteCommand@(vmw_ahci)#<None>+0x57 stack: 0x453a6899bb8f, 0x42002bce242e, 0x45bad72bbf80, 0x42002bce2572, 0x45dae3e318c0
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bb70:[0x42002bce242d]ataIssueCommand@(vmw_ahci)#<None>+0x56 stack: 0x45dae3e318c0, 0x10000000000104e, 0x42004a803b40, 0x45bad72bc2d0, 0x430b1e31ad40
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bb80:[0x42002bce2571]scsiQueueCommand@(vmw_ahci)#<None>+0xc2 stack: 0x42004a803b40, 0x45bad72bc2d0, 0x430b1e31ad40, 0x42002b17d458, 0x0
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bbb0:[0x42002b17d457]SCSIIssueCommandDirect@vmkernel#nover+0x214 stack: 0x430b1e31ad40, 0x4322b02647f0, 0x4322b0260036, 0x42002bce24b0, 0x42002b17d440
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bc30:[0x42002b17e8c1]SCSIStartAdapterCommands@vmkernel#nover+0x2c2 stack: 0x0, 0x430b1e31b430, 0x0, 0x43034d68aa90, 0x45dae3e318c0
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bcb0:[0x42002b192ce0]SCSIStartPathCommands@vmkernel#nover+0x735 stack: 0x4322b02647f0, 0x420000000001, 0x20093500000000, 0x430b1e3215e0, 0x0
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bdb0:[0x42002b1994de]SCSIIssueAsyncPathCommandDirect@vmkernel#nover+0x2d7 stack: 0x5f5f5f5f5f5f5f5f, 0x653861356139645f, 0x3130303130393931, 0x3030303030303030, 0x435f4f4e28202230
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bea0:[0x42002b19b009]vmk_ScsiIssueAsyncPathCommandDirect@vmkernel#nover+0x4a stack: 0x0, 0x42002b1b11a1, 0x45dae3e318c0, 0x45dae3e31e00, 0xe3e31a00
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bec0:[0x42002b1b11a0]vmk_PsaStorIssueAsyncPathCommandDirect@vmkernel#nover+0x2d stack: 0xe3e31a00, 0x42002ce5b1b8, 0x42f, 0x430b1e331970, 0x4319e5406500
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bee0:[0x42002ce5b1b7]HppIssueCommand@(hpp)#<None>+0x164 stack: 0x4319e5406500, 0x45dae3e31e00, 0x42002ce6560c, 0x2, 0x1
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bf20:[0x42002ce5b59e]HppAttemptFailoverRequest@(hpp)#<None>+0xfb stack: 0x453a6899f000, 0x430340601220, 0x4319e5407490, 0x4319e54018d0, 0x4319e54074d0
2026-08-15T15:08:03.825Z In(182) vmkernel: cpu42:2098695)0x453a6899bf60:[0x42002ad5bf7f]HelperQueueFunc@vmkernel#nover+0x300 stack: 0x4319e54018e8, 0x453a6899f000, 0x453a46c9006e, 0x42002ce5b4a4, 0x42002ad5bf78
2026-08-15T15:07:57.825Z In(182) vmkernel: cpu26:2098695)HPP: HppAttemptFailoverRequest:1071: Re-issuing first command for HPP device "t10.ATA_____SMC_VD__________________________________d9a##############000000" (NO_CONNECT_ON_APD = CLEAR)  failoverState=2 blocked=1
2026-08-15T15:07:58.825Z In(182) vmkernel: cpu26:2098695)HPP: HppAttemptFailoverRequest:1071: Re-issuing first command for HPP device "t10.ATA_____SMC_VD__________________________________d9a##############000000" (NO_CONNECT_ON_APD = CLEAR)  failoverState=2 blocked=1
2026-08-15T15:07:59.825Z In(182) vmkernel: cpu30:2098695)HPP: HppAttemptFailoverRequest:1071: Re-issuing first command for HPP device "t10.ATA_____SMC_VD__________________________________d9a##############000000" (NO_CONNECT_ON_APD = CLEAR)  failoverState=2 blocked=1
2026-08-15T15:08:00.825Z In(182) vmkernel: cpu23:2098695)HPP: HppAttemptFailoverRequest:1071: Re-issuing first command for HPP device "t10.ATA_____SMC_VD__________________________________d9a##############000000" (NO_CONNECT_ON_APD = CLEAR)  failoverState=2 blocked=1
2026-08-15T15:08:01.825Z In(182) vmkernel: cpu23:2098695)HPP: HppAttemptFailoverRequest:1071: Re-issuing first command for HPP device "t10.ATA_____SMC_VD__________________________________d9a##############000000" (NO_CONNECT_ON_APD = CLEAR)  failoverState=2 blocked=1
2026-08-15T15:08:02.825Z In(182) vmkernel: cpu32:2098695)HPP: HppAttemptFailoverRequest:1071: Re-issuing first command for HPP device "t10.ATA_____SMC_VD__________________________________d9a##############000000" (NO_CONNECT_ON_APD = CLEAR)  failoverState=2 blocked=1


The ESXi host experiences an abnormal storage I/O condition involving the ATA/AHCI device.

The vmw_ahci driver reports an inconsistent AHCI command state, logging errors such as Tag 1 SActive already set: SACT:3fe, CI:3fe.

This indicates that the underlying storage controller and device communication path is unable to successfully complete or recover outstanding I/O commands.

Because the AHCI stack is locked, the HPP multi-pathing module repeatedly fails to route command traffic and enters a continuous recovery loop.

The primary fault domain is likely a controller lockup, a SATA/SATADOM failure, a PCIe communication issue, or a related firmware problem.

A PCIe Master Abort is a possible underlying mechanism, but it requires corroborating hardware or PCIe logs from the BMC to be officially confirmed.

 

Resolution

Phase A: Immediate Recovery

1. After protecting or relocating workloads, perform a full host reboot or cold power cycle to reset the AHCI controller, SATA link, and PCIe device states.

2. Restarting management agents (like hostd or vpxa) or vCenter services will not clear a physical AHCI controller lockup or PCIe Master Abort.


Phase B: Firmware Lifecycle Management

1. Update the server hardware to the latest vendor-supported firmware baseline.

2. Required updates include the System BIOS/UEFI, BMC/IPMI, AHCI/SATA controller, and SATADOM firmware.


Phase C: Hardware Validation

1. Identify the exact physical device mapped to the SMC_VD identifier using ESXi CLI commands such as esxcli storage core device list.

2. Check for accumulated failed commands on the device by running esxcli storage core device stats get -d <device-ID>.

3. Review the BMC/IPMI hardware logs at the exact incident time for hardware faults like PCIe errors, SATA link resets, Master Aborts, or SATADOM errors.


Phase D: Hardware Replacement

1. If the vmw_ahci errors and HPP failover loops recur after updating firmware and completing a cold boot, replace the affected SATADOM storage device first.

2. If the problem persists with a known-good SATADOM installed, escalate the issue to the hardware vendor for a deeper motherboard, PCIe, or AHCI controller investigation.