A NMI IPI PSOD is triggered by a vmnic link failure and storage timeouts on the ESXi host
search cancel

A NMI IPI PSOD is triggered by a vmnic link failure and storage timeouts on the ESXi host

book

Article ID: 444524

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

  • ESXi host shows "Not Responding" in the vSphere client.

  • Host crashes with a Purple Screen of Death (PSOD) showing: "NMI IPI: Panic requested by another PCPU".

    Panic Message: @BlueScreen: NMI IPI: Panic requested by another PCPU. PC 0x#######, SP 0x###### (Src 0x#, CPU##)
    Backtrace:
      0x#########:[0x#########]PanicvPanicInt@vmkernel#nover+0x### stack: 0x####, 0x#########, 0x0, 0x#########, 0x#########
      0x#########:[0x#########]Panic_WithBacktrace@vmkernel#nover+0x## stack: 0x#########, 0x#########, 0x#########, 0x#########, 0x#########
      0x#########:[0x#########]NMI_Interrupt@vmkernel#nover+0x### stack: 0x#########, 0x#########, 0x#########, 0x#########, 0x#########
      0x#########:[0x#########]IDTNMIWork@vmkernel#nover+0x## stack: 0x#, 0x#########, 0x0, 0x#########, 0x###
      0x#########0:[0x#########]Int2_NMI@vmkernel#nover+0x# stack: 0x###, 0x###, 0x#, 0x#, 0x#########
      0x#########:[0x#########]gate_entry@vmkernel#nover+0x## stack: 0x#, 0x#, 0x#########, 0x#, 0x#########
      0x#########:[0x#########]int_13@vmkernel#nover+0x0 stack: 0x###, 0x#####, 0x#########, 0x0, 0x#########
    Saved backtrace from: pcpu 17 Heartbeat NMI
      0x#########:[0x#########]int_12@vmkernel#nover+0x## stack: 0x###

  • Virtual machines restart on other hosts in the cluster following a vSphere HA event.

  • In /var/run/log/vobd.log, physical link down events for "vmnic####" appear with "Failed criteria: 128".

    YYYY-MM-DDTHH:MM:SS In(##) vobd[#######]: [netCorrelator] #############us: [vob.net.dvport.uplink.transition.down] Uplink: vmnic# is down. Affected dvPort: ###/## ## ## ## ## ## ## ##-## ## ## ## ## ## ## ##. # uplinks up. Failed criteria: 128

  • The /var/run/log/vobd.log records vob.vmfs.heartbeat.timedout for the backend storage array or alternate backing LUNs shortly before the crash occurs.

    YYYY-MM-DDTHH:MM:SS In(##) vobd[#######]: [vmfsCorrelator] ###############us: [vob.vmfs.heartbeat.timedout] ########-####-####-############ ####-#######-#####-LUN-###
    YYYY-MM-DDTHH:MM:SS In(##) vobd[#######]: [vmfsCorrelator] ###############us: [esx.problem.vmfs.heartbeat.timedout] #######-########-####-######### ####-####-####-LUN-##
    YYYY-MM-DDTHH:MM:SS In(##) vobd[#######]: The event ([esx.problem.vmfs.heartbeat.timedout] ########-######-####-########### ######-##-######-LUN-##) was sent immediately to hostd;

Environment

VMware vSphere ESXi 8.x

Cause

A physical link down event on a vmnic results in lost network redundancy. Subsequent storage communication timeouts to backing LUNs trigger a cascading failure, culminating in an NMI IPI panic.

Resolution

Identify and resolve the underlying hardware or environmental fault.

  1. Perform hardware inspection on the physical NIC, cables, and upstream switch ports.
  2. Ensure network adapter firmware and drivers match supported versions in the Broadcom Compatibility Guide.  
  3. Consult the hardware vendor for physical layer diagnostics.

Additional Information

For more information on Failed Criteria 128, refer to the following KB: Network adapter (vmnic) is down or fails with a Failed Criteria Code