Manual crash triggered by NMI on a virtual machine hosted on ESXi
search cancel

Manual crash triggered by NMI on a virtual machine hosted on ESXi

book

Article ID: 389912

calendar_today

Updated On:

Products

VMware vSphere ESX 8.x VMware vSphere ESXi 8.0

Issue/Introduction

  • A manual crash triggers on the system using a Non-Maskable Interrupt (NMI), as dump analysis from the OS vendor indicates.

    NMI_HARDWARE_FAILURE (80)
    This is typically due to a hardware malfunction. The hardware supplier should be called.
    Arguments:
    Arg1: 00000000004f4454, 'TDO'
    Arg2: 0000000000000000, Status Byte
    Arg3: 0000000000000000
    Arg4: 0000000000000000
    Debugging Details:
    VIRTUAL_MACHINE: VMware
    FAULTING_THREAD:  fffff8021b1c5980
    PROCESS_NAME: System
    STACK_TEXT: 
    fffff802`1d02fc08 fffff802`1b65aa4e     : 00000000`00000080 00000000`004f4454 00000000`00000000 00000000`00000000 : nt!KeBugCheckEx [d:\rs1\minkernel\ntos\ke\amd64\procstat.asm @ 127]
    fffff802`1d02fc10 fffff802`1b034920     : ffff810c`6aa82968 fffff802`1b672e30 fffff802`1b672e30 00000000`00000001 : hal!HalBugCheckSystem+0x7e [d:\rs1\minkernel\hals\lib\whea\mca.c @ 3171]
    fffff802`1d02fc50 fffff802`1b65ba1e     : fffff802`000006c0 fffff802`1b1c3630 fffff802`1d02fd30 fffff802`1aec41b8 : nt!WheaReportHwError+0x258 [d:\rs1\minkernel\ntos\whea\whea.c @ 810]
    fffff802`1d02fcb0 fffff802`1aec3ad6     : 00000000`00000001 00000000`00000000 000000a8`1f6b2685 fffff802`1aec3f88 : hal!HalHandleNMI+0xfe [d:\rs1\minkernel\hals\lib\whea\nmi.c @ 396]
    fffff802`1d02fce0 fffff802`1af6e242     : 00000000`00000001 fffff802`1d02fef0 00000000`00000000 00000000`00000000 : nt!KiProcessNMI+0x106 [d:\rs1\minkernel\ntos\ke\amd64\misc.c @ 298]
    fffff802`1d02fd30 fffff802`1af6e043     : 00000000`00000000 00000000`00000000 00000000`00000000 00000000`00000000 : nt!KxNmiInterrupt+0x82 [d:\rs1\minkernel\ntos\ke\amd64\trap.asm @ 697]
    fffff802`1d02fe70 fffff802`1b011cd9     : 00000000`00000046 ffff810c`6ac460f0 00000000`00000000 00000000`00000000 : nt!KiNmiInterrupt+0x1c3 [d:\rs1\minkernel\ntos\ke\amd64\trap.asm @ 657]
    fffff802`1d01c200 fffff802`1ae66c8e     : fffff802`1b011cf0 00000000`00000000 ffff810c`6ac46010 00000000`000000ae : nt!PpmIdleGuestExecute+0x15 [d:\rs1\minkernel\ntos\po\ppmhv.c @ 697]
    fffff802`1d01c240 fffff802`1ae65e1a     : fffff804`0fa65f40 012d92f8`012d92f8 00000000`00000000 00000000`00000000 : nt!PpmIdleExecuteTransition+0xcbe [d:\rs1\minkernel\ntos\po\ppmidle.c @ 3925]
    fffff802`1d01c4c0 fffff802`1af6611c     : 00000000`00000000 fffff802`1b149180 fffff802`1b1c5980 ffff810c`79728040 : nt!PoIdle+0x33a [d:\rs1\minkernel\ntos\po\ppmidle.c @ 1046]
    fffff802`1d01c620 00000000`00000000     : fffff802`1d01d000 fffff802`1d016000 00000000`00000000 00000000`00000000 : nt!KiIdleLoop+0x2c [d:\rs1\minkernel\ntos\ke\amd64\idle.asm @ 110]
  • The VM enters a suspended state, and the vCenter task fails to provide sufficient diagnostic information.
  • Administrators assert that they initiate no suspended tasks for the VM.

Environment

  • VMware vSphere ESX 8.x
  • VMware vSphere ESX 7.x

Cause

  • In this scenario, a VM enters a suspended state, with the vCenter task providing insufficient information. Users assert that they initiate no suspended tasks, suggesting that the VM's suspension is not a result of standard user intervention.
  • Errors reported from the VMware.log and hostd.log show that the VM crashed. The VM's failure to suspend automatically leads to a manual crash via NMI to address operational inefficiencies.

/var/run/log/hostd.log
YYYY-MM-DD TT:37.090Z info hostd[2162299] [Originator@6876 sub=Vimsvc.CgiServiceManager opID=CGI-server-8b44] Validating CGI ticket for URL '/cgi-bin/vm-support.cgi?listmanifests=1'
YYYY-MM-DD TT:51.170Z info hostd[2162204] [Originator@6876 sub=Vimsvc.CgiServiceManager opID=CGI-server-8b5b] Validating CGI ticket for URL '/cgi-bin/vm-support.cgi?manifests=VirtualMachines:CoreDumpHung HungVM:Send_NMI_To_Guest HungVM:Suspend_VM'
YYYY-MM-DD:45.908Z info hostd[2162334] [Originator@6876 sub=Vimsvc.CgiServiceManager opID=CGI-server-9088] Validating CGI ticket for URL '/cgi-bin/vm-support.cgi?manifests=VirtualMachines:CoreDumpHung'

/vmfs/volume/datastore/<VM-Name>/vmware.log
YYYY-MM-DD TT.499Z Wa(03) vcpu-0 - WinBSOD: Synthetic MSR[0x40000100] 0x80
YYYY-MM-DD TT.499Z Wa(03)+ vcpu-0 -
YYYY-MM-DD TT.499Z Wa(03) vcpu-0 - WinBSOD: Synthetic MSR[0x40000101] 0x4f4454
YYYY-MM-DD TT.499Z Wa(03)+ vcpu-0 -
YYYY-MM-DD TT.499Z Wa(03) vcpu-0 - WinBSOD: Synthetic MSR[0x40000102] 0x0
YYYY-MM-DD TT.499Z Wa(03)+ vcpu-0 -
YYYY-MM-DD TT.499Z Wa(03) vcpu-0 - WinBSOD: Synthetic MSR[0x40000103] 0x0
YYYY-MM-DD TT.499Z Wa(03)+ vcpu-0 -
YYYY-MM-DD TT.499Z Wa(03) vcpu-0 - WinBSOD: Synthetic MSR[0x40000104] 0x0
YYYY-MM-DD TT.499Z Wa(03)+ vcpu-0 -YYYY-MM-DD TT.521Z In(05) vmx - SUSPEND: Start suspend (flags=0)
YYYY-MM-DD TT.535Z In(05) vcpu-0 - Progress -1% (msg.checkpoint.saveStatus)
YYYY-MM-DD TT.535Z In(05) vcpu-0 - Checkpointed in VMware ESX, 8.x, build-########, Linux Host
YYYY-MM-DD TT.679Z In(05) vcpu-0 - Progress 0% (none)
YYYY-MM-DD TT.679Z In(05) vcpu-0 - MainMem: Writing full memory image, '/vmfs/volumes/########-########-###-##########/<VM-Name>/VMName-#######.vmem'.
YYYY-MM-DD TT:53.011Z In(05) vcpu-0 - Progress 1% (none)
YYYY-MM-DD TT:54.394Z In(05) vcpu-0 - Progress 2% (none)
YYYY-MM-DD TT:55.768Z In(05) vcpu-0 - Progress 3% (none)
YYYY-MM-DD TT:57.143Z In(05) vcpu-0 - Progress 4% (none)
YYYY-MM-DD TT:14.637Z In(05) vcpu-0 - SUSPEND: Completed suspend: 'Operation completed successfully' (0)
YYYY-MM-DD TT:14.637Z In(05) vmx - Stopping VCPU threads...
YYYY-MM-DD TT:14.637Z In(05) vcpu-0 - VMMon_WaitForExit: vcpu-0: worldID=2103623
YYYY-MM-DD TT:14.637Z In(05) vcpu-1 - VMMon_WaitForExit: vcpu-1: worldID=2103627
YYYY-MM-DD TT:14.637Z In(05) vcpu-3 - VMMon_WaitForExit: vcpu-3: worldID=2103629
YYYY-MM-DD TT:14.637Z In(05) vcpu-2 - VMMon_WaitForExit: vcpu-2: worldID=2103628
YYYY-MM-DD TT:14.637Z In(05) vcpu-4 - VMMon_WaitForExit: vcpu-4: worldID=2103630
YYYY-MM-DD TT:14.637Z In(05) vmx - MonitorVMMCoreRequest: ********************************************
YYYY-MM-DD TT:14.637Z In(05) vmx - MonitorVMMCoreRequest: Sync core dump requested; not a real fault
YYYY-MM-DD TT:14.637Z In(05) vmx - MonitorVMMCoreRequest: ********************************************
YYYY-MM-DD TT:14.637Z In(05) vcpu-5 - VMMon_WaitForExit: vcpu-5: worldID=2103631
YYYY-MM-DD TT:14.637Z Wa(03) vmx - Dumping vmx core per request
YYYY-MM-DD TT:15.658Z Wa(03) vmx - A core file is available in "/vmfs/volumes/########-########-###-##########/<VM-Name>/vmx-zdump.000"

Resolution

To prevent unexpected VM suspension and crashes during log collection, verify the options selected during log bundle generation.

Do not select the HungVM option unless you intentionally want to force a core dump and suspend the virtual machine. Initiating a log bundle generation with the HungVM option selected causes the VM to enter a suspend state. The traces indicate that there are no underlying issues with either ESXi or vCenter.

Additional Information