Virtual machine experiences network loss after MCE (Memory Controller Errors)
search cancel

Virtual machine experiences network loss after MCE (Memory Controller Errors)

book

Article ID: 389882

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

  • Virtual machine network connectivity is lost.

  • Virtual machine becomes non-responsive and reboots.

  • Slow response or loss of access to backing storage.

  • The /vmfs/volumes/[Datastore]/[vm folder]/vmware.log contains the following entries:

    YYYY-MM-DDTHH:MM:SS In(05) vcpu-0 - Tools: Tools heartbeat timeout.
    YYYY-MM-DDTHH:MM:SS In(05)+ vcpu-0 - The CPU has been disabled by the guest operating system. Power off or reset the virtual machine.

  • ESXi /var/run/log/vobd.log shows multiple MCE and datastore timeout events:

    YYYY-MM-DDTHH:MM:SS In(14) vobd[####]: [cpuCorrelator] ####us: [vob.cpu.mce.log4] MCE bank 8: status:0x### misc:0x### addr:0x###### cpu:4 physAddr:0x###### physSize:0x40 ceCount:0x1
    YYYY-MM-DDTHH:MM:SS In(14) vobd[####]: [cpuCorrelator] ####us: [vob.cpu.mem.ce.log] Corrected memory error at physAddr:0x######
    YYYY-MM-DDTHH:MM:SS In(14) vobd[####]: [vmfsCorrelator] ####us: [vob.vmfs.heartbeat.timedout] ###-###-###
    YYYY-MM-DDTHH:MM:SS In(14) vobd[####]: [vmfsCorrelator] ####us: [esx.problem.vmfs.heartbeat.timedout] ###-###-###

Environment

  • VMware ESXi 7.x

  • VMware ESXi 8.x

Cause

The issue is caused by physical hardware failure involving the memory controller or individual DIMM modules on the ESXi host. These hardware-level errors trigger memory corruption that impacts storage I/O and virtual machine execution.

Resolution

Contact the physical server hardware vendor to perform hardware diagnostics, focusing on physical memory modules and the CPU memory controller.