An ESXi host encountering a Purple Screen of Death (PSOD) might display a reference to a "Machine Check Exception".
The ESXi host may enter a "Not Responding" state in vCenter Server.
At the ESXi console, the PSOD screen will have entries similar to the following:
VMware ESXi #.#.# [Releasebuild-######## x86_64]Machine Check Exception on PCPU## in world ######:idle51System has encountered a Hardware Error - Please contact the hardware vendorUncorrectable/recoverable memory error in world ####; unable to recover in kernel contextData Cache DataRead Error
The purple screen backtrace may also include vSAN-related functions such as BucketlistInsertInNode, Bucketlist_Insert, Rangemap_Insert, or PLOGPrepareWrite.
The following entries are observed in the /var/run/log/vmkernel.log file on the ESXi host:
YYYY-MM-DDTHH:MM:SS.FFFZ cpu##:40848027)ALERT: MCA: 200: SRAR Excp G7 B1 ###### Cache Hierarchy: Level 0 Data Cache DataRead Error.YYYY-MM-DDTHH:MM:SS.FFFZ cpu##:40848027)MCAIntel: 1120: Force retiring MPN ###### to recover from MCA error detected by cpu## in ####.Machine Check Exception on PCPU## in world 2099051:VSAN_0x4##18
The error can also be caused by a failing hardware device. In such cases, the PSOD screen may report an error similar to the following:
YYYY-MM-DDTHH:MM:SS.FFFZ cpu##:########)IDT: ### : Uncorrectable/unrecoverable machine check errorYYYY-MM-DDTHH:MM:SS.FFFZ cpu##:########)MCA: ### : UC Excp G4 86 Sbb###########e#b AB M###### P8/8 I/O error reported by PCI ####:##:##.#.
The Machine Check Architecture (MCA) is a CPU feature designed to detect and report hardware anomalies. When the hardware detects a critical or fatal condition, it raises a Machine Check Exception (MCE). These exceptions are considered severe and unrecoverable, resulting in an expected ESXi host crash, often resulting in a Purple Screen of Death (PSOD).
In this scenario, the MCE was categorized as a Software Recoverable Action Required (SRAR), which indicates the following:
The faulty thread was executing within a critical vmkernel context (e.g., a vSAN thread like VSAN_0x45018) and the ESXi host was unable to isolate or terminate it, leaving the hypervisor with no viable recovery path. As a result, the MCE was escalated to a fatal system error, causing the ESXi host to halt to ensure data integrity.
The underlying hardware failure must be addressed. Perform the following steps:
To speak with a customer representative or a Support Engineer see Contact Support. Scroll to the bottom of the page and click on your respective region.