Understand how to identify, analyze, and address generic hardware faults occurring on VMware vSphere ESXi hosts. This article outlines common symptoms, log categories, and the recommended troubleshooting path for hardware-related issues.
An ESXi host might experience the following behavior when a generic hardware fault occurs:
When a hardware error occurs, the host generates an alert and indicates the problem within the hardware monitoring interface. Note that these alerts may be transitory and clear automatically once the fault condition is no longer actively reported, even if the underlying fault persists.
Hardware and CIM diagnostic logs are critical for determining if fault conditions have occurred in the past.
The following categories are the severity of states that indicate required action to resolve with examples of the log entries below.
Processor Errors:
Processor IERRProcessor Thermal TripProcessor Configuration ErrorProcessor Machine Check ExceptionProcessor Correctable Machine CheckMemory Errors:
Memory Configuration ErrorMemory Uncorrectable ECCMemory Transition to CriticalMemory Critical OvertemperatureDisk Errors:
Drive Slot In Critical ArrayDrive Slot In Failed ArrayDrive Bay in Critical ArrayDrive Bay in Failed ArrayDrive Slot Drive FaultBus Errors:
PCI PERRPCI SERRBus Correctable ErrorBus Uncorrectable ErrorBus Fatal ErrorAdd-in Card Install ErrorCable/Interconnect Transition to Critical from less severeSlot/Connector Transition to CriticalSlot/Connector Transition to Non-criticalFan Errors:
Fan Transition to Critical from less severeFan Transition to Off LineTemperature Errors:
Temperature Lower Critical going lowTemperature Transition to Critical from less severeTemperature Transition to Non-recoverable from less severeTemperature Upper Critical going highVoltage Errors:
Voltage Limit ExceededVoltage Transition to Critical from less severeExample:
The following is an example of what the CIM diagnostic log might display:OMC_IpmiLogRecord.CreationClassName="OMC_IpmiLogRecord",LogCreationClassName="OMC_IpmiRecordLog",LogName="IPMI SEL",MessageTimestamp="YYYYMMDDHHMMSS.000000+000",RecordID="1"RecordID = 1MessageTimestamp = (NULL)LogName = IPMI SELLogCreationClassName = OMC_IpmiRecordLogCreationClassName = OMC_IpmiLogRecordRecordFormat = *string CIM_Sensor.DeviceID*uint8[2] IPMI_RecordID*uint8 IPMI_RecordType*uint8[4] IPMI_Timestamp*uint8[2] IPMI_GeneratorID*uint8 IPMI_EvMRev*uint8 IPMI_SensorType*uint8 IPMI_SensorNumber*boolean IPMI_AssertionEvent*uint8 IPMI_EventType*uint8 IPMI_EventData1*uint8 IPMI_EventData2*uint8 IPMI_EventData3*uint32 IANA*RecordData = *###.#.##*# #*#*## ## ### ##*## #*#*##*###*false*###*#*###*###*#*ElementName = IPMI SELDescription = Assert + Voltage Transition to Critical from less severeCaption = Assert + Voltage Transition to Critical from less severePerceivedSeverity = (NULL)Locale = (NULL)InstanceID = (NULL)DataFormat = (NULL)
To resolve hardware faults, follow these steps:
For instructions on how to retrieve log bundles from your ESXi host, refer to the official Broadcom Documentation.
To speak with a customer representative or a Support Engineer see Contact Support. Scroll to the bottom of the page and click on your respective region.