An ESXi host experiences an unexpected reboot without throwing a Purple Screen of Death (PSOD). The host abruptly drops offline and automatically power cycles.
Upon reviewing the logs, the following indicators are present:
1. IPMI / BMC System Event Log (SEL)
The hardware log shows a fatal error followed immediately by an ACPI power state transition to soft-off, and then a reboot back to working state:
YYYY-MM-DDTHH:MM:SS Fatal/NonRecoverable System Event Sensor 183 Assert + Module/Board Transition to Critical from less severe
YYYY-MM-DDTHH:MM:SS Unknown System Event Sensor 170 Assert + System ACPI Power State S4/S5: soft-off
YYYY-MM-DDTHH:MM:SS Unknown System Event Sensor 170 Assert + System ACPI Power State S0/G0: working
2.vobd.log
YYYY-MM-DDTHH:MM:SS In(14) vobd[#####]: [UserLevelCorrelator] 166878543us: [esx.audit.host.boot] Host has booted.
YYYY-MM-DDTHH:MM:SS In(14) vobd[#####]: The event ([esx.audit.host.boot] Host has booted.) was sent immediately to hostd;
YYYY-MM-DDTHH:MM:SS In(14) vobd[#####]: [GenericCorrelator] #####us: [vob.user.host.poweroff.reason.unavailable] The host is being powered off. The poweroff was not the result of a kernel error, deliberate reboot, or shut down. This could indicate a hardware issue. Hardware may reboot abruptly due to power outages, faulty components, and heating issues. To investigate further, engage the hardware vendor.
VMware vSphere ESXi 8.0
This issue is caused by a physical hardware failure on the server chassis or motherboard.
The BMC (Baseboard Management Controller) detected a Fatal/Non-Recoverable error on Sensor 183 (associated with a Module/Board transition to a critical state). To protect the server components from catastrophic damage,the hardware/BIOS triggered an immediate hard shutdown (S5 Soft-Off).
Since this is strictly a hardware-level failure, no software changes within ESXi will resolve the underlying issue.
Suggest to follow the steps below to fix:
1. Place the affected ESXi host into Maintenance Mode.
2. Ensure all active Virtual Machines are successfully migrated to healthy nodes in the cluster to prevent further service disruption.
3. Engage the Hardware Vendor