Virtual Machines experience reboots or unresponsiveness during storage node maintenance or reboot.
search cancel

Virtual Machines experience reboots or unresponsiveness during storage node maintenance or reboot.

book

Article ID: 449540

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

Virtual machines running on Clustering Solutions experience an unexpected reboot during an underlying storage node failover or reboot. The SAN maintenance causes transient storage conditions, and the available paths enter an ALUA transition state.

Symptoms:

  • Cluster VMs reboot unexpectedly during storage node maintenance.
  • In the vmware.log will see , the Guest Operating System issued a hardware reset command to the virtual chipset (Chipset: The guest has requested that the virtual machine be hard reset
    2026-07-07T09:45:58.156Z In(05) vcpu-0 - Chipset: The guest has requested that the virtual machine be hard reset.
  • In the var/run/log/hostd.log will notice that VMware Tools stopped running (toolsStatus=toolsNotRunning), followed shortly by the heartbeat turning red (guestHB=red).
    2026-07-07T09:46:11.013Z In(166) Hostd[7545627]: [Originator@6876 sub=Vmsvc.vm:/vmfs/volumes/##################/abc.vmx] VMware Tools (version=13322) not running; Setting heartbeat to red
    2026-07-07T09:46:11.013Z Db(167) Hostd[7545627]: [Originator@6876 sub=Vmsvc.vm:/vmfs/volumes/#################/abc.vmx] Updating current heartbeatStatus: green -> red

  • In the var/run/log/fdm.log will notice that VMware Tools stopped running (toolsStatus=toolsNotRunning), followed shortly by the heartbeat turning red (guestHB=red). .
    2026-07-07T09:46:07.311Z Db(167) Fdm[7897786]: [Originator@6876 sub=Invt opID=WorkQueue-b543de9] Vm /vmfs/volumes/################/abc/abc.vmx changed toolsStatus=toolsNotRunning
    2026-07-07T09:46:11.015Z Db(167) Fdm[7897793]: [Originator@6876 sub=Invt opID=WorkQueue-63a8467a] Vm /vmfs/volumes/########################/abc/abc.vmx changed  guestHB=red

Environment

  • VMware vSphere 8.x
  • VMware vSphere 9.x

Cause

During a storage controller reboot, ESXi Native Multipathing (NMP) failover experiences the transient storage condition on the other available paths . The host pauses I/O while waiting for the transient error to clear and for paths to fail over to the surviving controller.
Because of the prolonged I/O wait, the clustering software running inside the Guest OS interprets the delay as a cluster node failure.


Cause Validation-
The var/run/log/vmkernel.log onfirms the transient storage condition
VMW_SATP_ALUA: satp_alua_issueCommandOnPath:966: Path (vmhba##:C#:T#:L##) command 0xa3 : Failed with transient error status Transient storage condition, suggest retry. sense data: 0x6 0x2a 0x6.
VMW_SATP_ALUA: satp_alua_issueCommandOnPath:972: Path (vmhba##:C#:T#:L##) command 0xa3 : Waiting for 20 seconds for the transient error to change before marking the path down

Resolution

  1. Engage storage vendor to o investigate why the storage node reboot resulted in an extended I/O outage rather than a seamless path failover.
  2. Verify that all ESXi hosts have active, redundant paths to the alternate storage controller before initiating future storage maintenance.

Additional Information

For further assistance, contact Broadcom Support