Hardware faults on ESXi hosts
search cancel

Hardware faults on ESXi hosts

book

Article ID: 336323

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

Understand how to identify, analyze, and address generic hardware faults occurring on VMware vSphere ESXi hosts. This article outlines common symptoms, log categories, and the recommended troubleshooting path for hardware-related issues.

An ESXi host might experience the following behavior when a generic hardware fault occurs: 

  • Erratic host behavior
  • Purple screen errors
  • Corrupt disk drives
  • Erratic virtual machine behavior
  • Incorrect CPU usage reporting (e.g., >100%)

When a hardware error occurs, the host generates an alert and indicates the problem within the hardware monitoring interface. Note that these alerts may be transitory and clear automatically once the fault condition is no longer actively reported, even if the underlying fault persists.

Hardware and CIM diagnostic logs are critical for determining if fault conditions have occurred in the past.

The following categories are the severity of states that indicate required action to resolve with examples of the log entries below.

Processor Errors:

  • Processor IERR
  • Processor Thermal Trip
  • Processor Configuration Error
  • Processor Machine Check Exception
  • Processor Correctable Machine Check

Memory Errors:

  • Memory Configuration Error
  • Memory Uncorrectable ECC
  • Memory Transition to Critical
  • Memory Critical Overtemperature

Disk Errors:

  • Drive Slot In Critical Array
  • Drive Slot In Failed Array
  • Drive Bay in Critical Array
  • Drive Bay in Failed Array
  • Drive Slot Drive Fault

Bus Errors:

  • PCI PERR
  • PCI SERR
  • Bus Correctable Error
  • Bus Uncorrectable Error
  • Bus Fatal Error
  • Add-in Card Install Error
  • Cable/Interconnect Transition to Critical from less severe
  • Slot/Connector Transition to Critical
  • Slot/Connector Transition to Non-critical

Fan Errors:

  • Fan Transition to Critical from less severe
  • Fan Transition to Off Line

Temperature Errors:

  • Temperature Lower Critical going low
  • Temperature Transition to Critical from less severe
  • Temperature Transition to Non-recoverable from less severe
  • Temperature Upper Critical going high

Voltage Errors:

  • Voltage Limit Exceeded
  • Voltage Transition to Critical from less severe

Example:

The following is an example of what the CIM diagnostic log might display:

OMC_IpmiLogRecord.CreationClassName="OMC_IpmiLogRecord",LogCreationClassName="OMC_IpmiRecordLog",LogName="IPMI SEL",MessageTimestamp="YYYYMMDDHHMMSS.000000+000",RecordID="1"
RecordID = 1
MessageTimestamp = (NULL)
LogName = IPMI SEL
LogCreationClassName = OMC_IpmiRecordLog
CreationClassName = OMC_IpmiLogRecord
RecordFormat = *string CIM_Sensor.DeviceID*uint8[2] IPMI_RecordID*uint8 IPMI_RecordType*uint8[4] IPMI_Timestamp*uint8[2] IPMI_GeneratorID*uint8 IPMI_EvMRev*uint8 IPMI_SensorType*uint8 IPMI_SensorNumber*boolean IPMI_AssertionEvent*uint8 IPMI_EventType*uint8 IPMI_EventData1*uint8 IPMI_EventData2*uint8 IPMI_EventData3*uint32 IANA*
RecordData = *###.#.##*# #*#*## ## ### ##*## #*#*##*###*false*###*#*###*###*#*
ElementName = IPMI SEL
Description = Assert + Voltage Transition to Critical from less severe
Caption = Assert + Voltage Transition to Critical from less severe
PerceivedSeverity = (NULL)
Locale = (NULL)
InstanceID = (NULL)
DataFormat = (NULL)

Environment

  • VMware vSphere ESXi 6.x
  • VMware vSphere ESXi 7.x
  • VMware vSphere ESXi 8.x

Cause

  • When a hardware error occurs, the host generates an alert and indicates the hardware problem on the hardware monitor tab.
  • However, the alert is only displayed while the hardware error occurs, and the alert sometimes clears.
  • This does not indicate that the hardware fault has stopped occurring, but that the indications of the fault stopped.
  • When a hardware fault occurs, even if it displays in a transitory state, the host will log the fault in the hardware and CIM diagnostics logs.

Resolution

To resolve hardware faults, follow these steps:

  1. Review the hardware monitor tab in the vSphere Client for active alerts.
  2. Examine the IPMI/CIM diagnostic logs for historical error entries.
  3. If logs indicate hardware fault conditions, contact the server hardware vendor support team for detailed diagnostics and component replacement.
  4. For additional information on downloading Broadcom products, patches, or software updates, refer to Download Broadcom Products and Software.

Additional Information

For instructions on how to retrieve log bundles from your ESXi host, refer to the official Broadcom Documentation.

To speak with a customer representative or a Support Engineer see Contact Support. Scroll to the bottom of the page and click on your respective region.