The ESXi host experiences a Purple Screen of Death (PSOD) with the error message NMI IPI: Panic requested by another PCPU.
While the active backtrace on the panic screen may show threads executing within networking, security, or firewall modules—such as vmware-sfw or vsip (VMware vDefend Firewall)—the underlying failure is triggered by critical physical hardware memory faults.
(Important Note: This underlying hardware issue can disrupt any active thread running on the system. The specific log trace provided below is merely one instance of how the failure manifested while a thread happened to be processing DFW IOCTL functionality.)
vmkernel.log): Contains ACPI Platform Error Interface (APEI) references indicating page retirement actions due to memory array errors: vmkernel: ApeiPageRetire: 730: Processing HEST GESB, severity 0x2, with 3 GEDE record(s)
vmkernel: ApeiPageRetire: 654: Memory error 1: val=#### err=400 adr=#### msk=0 nod=1...
WARNING: Heartbeat: [ID]: PCPU [number] didn't have a heartbeat for [X] seconds, timeout is 10, [X] IPIs sent; *may* be locked up.
WARNING: StorageDeviceIO: 201: Device eui.#### performance has deteriorated. I/O latency increased from average value of 47 microseconds to 38100 microseconds.
Panic_WithBacktrace (sbt=sbt@entry=0x[MASKED], fmt=fmt@entry=0x[MASKED] "NMI IPI: Panic requested by another PCPU. PC %#lx, SP %#lx (Src %#x, CPU%u)") at bora/vmkernel/main/panic.c:141
NMIHandleBtOrHaltRequest (source=NMI_SRC_HEARTBEAT, fullFrame=0x[MASKED]) at bora/vmkernel/main/nmi.c:732
#2 NMIHandleIPISource (fullFrame=0x[MASKED]) at bora/vmkernel/main/nmi.c:551
#3 NMI_Interrupt (fullFrame=fullFrame@entry=0x[MASKED]) at bora/vmkernel/main/nmi.c:781
IDTNMIWork (fullFrame=fullFrame@entry=0x[MASKED]) at bora/vmkernel/main/x86/idt.c:1773
Int2_NMI (fullFrame=0x[MASKED]) at bora/vmkernel/main/x86/idt.c:1015
#6 0x[MASKED] in gate_entry () at bora/vmkernel/main/x86/gates64.S:175
n rn_walktree (h=<optimized out>, f=f@entry=0x[MASKED] <pfr_walktree>, w=w@entry=0x[MASKED]) at datapath/esx/modules/vsip/vsip_pf/pf_vmk/radix.c:1406
pfr_enqueue_addrs (kt=kt@entry=0x[MASKED], workq=workq@entry=0x[MASKED], naddr=naddr@entry=0x[MASKED], sweep=sweep@entry=0) at datapath/esx/modules/vsip/vsip_pf/pf/pf_table.c:1244
pfr_destroy_ktable (kif=kif@entry=0x[MASKED], kt=kt@entry=0x[MASKED], flushflags=flushflags@entry=5, set=set@entry=PFR_SET_INACTIVE) at datapath/esx/modules/vsip/vsip_pf/pf/pf_table.c:4222
pfr_activate_post_task (kif=kif@entry=0x[MASKED], delete_list=delete_list@entry=0x[MASKED]) at datapath/esx/modules/vsip/vsip_pf/pf/pf_table.c:3428
in pfioctl (kif=kif@entry=0x[MASKED], dev=dev@entry=0x[MASKED], cmd=cmd@entry=[MASKED], addr=addr@entry=0x[MASKED] "", flags=flags@entry=2, td=td@entry=0x[MASKED]) at datapath/esx/modules/vsip/vsip_pf/pf/pf_ioctl.c:7544
VSIPConversionActivateAddrSets (kif=0x[MASKED], msg=0x[MASKED], msgLen=<optimized out>, result=0x[MASKED]) at datapath/esx/modules/vsip/vsip_pf/pf_vmk/msg2pf.c:4666
VSIPToPFIoctl (cookie=<optimized out>, cmd=<optimized out>, data=<optimized out>, dataLen=<optimized out>, result=0x[MASKED]) at datapath/esx/modules/vsip/vsip_pf/pf_vmk/msg2pf.c:10185
VSIPFWImplIter (sol=0x[MASKED], filter=filter@entry=0x[MASKED], data=data@entry=0x[MASKED]) at datapath/esx/modules/vsip/vsip_fw.c:183
VSIPDVFConfigFilterByName (fpAgentName=fpAgentName@entry=0x[MASKED] "vmware-sfw", filterName=filterName@entry=0x[MASKED] "shared-addrset-filter-", iter=iter@entry=0x[MASKED] <VSIPFWImplIter>, data=data@entry=0x[MASKED]) at datapath/esx/modules/vsip/vsip_dvfilter.c:6677
VSIPFWImpl (dvsport=dvsport@entry=0x[MASKED], cmdHdr=cmdHdr@entry=0x[MASKED]) at datapath/esx/modules/vsip/vsip_fw.c:219
VSIPIoctlFWCtrl (cmd=<optimized out>, iocData=<optimized out>, fnData=<optimized out>, result=<optimized out>) at datapath/esx/modules/vsip/vsip_fw_ioctl.c:303
VSIPIoctlImpl (cmd=cmd@entry=15, req=req@entry=0x[MASKED], result=result@entry=0x[MASKED]) at datapath/esx/modules/vsip/vsip_ioctl.c:148
VSIPCharDevIoctl (attr=<optimized out>, cmd=15, userData=[MASKED], callerSize=<optimized out>, result=0x[MASKED]) at datapath/esx/modules/vsip/vsip_dev.c:55
VMKAPICharDevIoctl (handle=0x[MASKED], handle=0x[MASKED], ioctlResult=0x[MASKED], userData=[MASKED], cmd=15) at bora/vmkernel/main/vmkapi_char.c:649
#21 VMKAPICharDevDevfsWrapIoctl (handle=0x[MASKED], cmd=15, userData=[MASKED], ioctlResult=0x[MASKED]) at bora/vmkernel/main/vmkapi_char.c:1507
CharDriverIoctl (deviceHandleID=<optimized out>, cmd=<optimized out>, dataInOut=0x[MASKED]) at bora/vmkernel/filesystems/devices/charDriver.c:948
FDS_Ioctl (dataInOut=0x[MASKED], cmd=FDS_IOCTL_PASS_THRU, fdsHandle=<optimized out>) at bora/vmkernel/private/fsDeviceSwitch.h:780
#24 DevFSIoctl (fileDesc=0x[MASKED], fhID=[MASKED], cmd=IOCTLCMD_DEVFS_OPAQUE, dataIn=0x[MASKED], result=0x[MASKED]) at bora/vmkernel/filesystems/devfs/devfs.c:5459
FSSVec_Ioctl (desc=<optimized out>, fhID=<optimized out>, cmd=<optimized out>, dataIn=<optimized out>, result=<optimized out>) at bora/vmkernel/filesystems/fsSwitchVec.c:778
FSSObjectIoctlCommon (fhID=[MASKED], file=0x[MASKED], cmd=<optimized out>, dataIn=0x[MASKED], result=0x[MASKED], ioFlags=<optimized out>) at bora/vmkernel/filesystems/fsSwitch.c:4857
FSS_IoctlByFH (fileHandleID=[MASKED], cmd=cmd@entry=IOCTLCMD_DEVFS_OPAQUE, dataIn=dataIn@entry=0x[MASKED], result=0x[MASKED], ioFlags=ioFlags@entry=FS_INVALID_FLAG) at bora/vmkernel/filesystems/fsSwitch.c:4953
UserFile_PassthroughIoctl (vmfsObj=<optimized out>, cmd=<optimized out>, userData=<optimized out>, result=<optimized out>) at bora/vmkernel/user/userFile.c:1435
#29 0x[MASKED] in UserVmfs_Ioctl (obj=<optimized out>, cmd=<optimized out>, userArg=[MASKED], ioctlReturnCode=<optimized out>) at bora/vmkernel/user/userVmfs.c:1847
LinuxFileDesc_Ioctl (fd=<optimized out>, cmd=15, userData=[MASKED]) at bora/vmkernel/user/linuxFileDesc.c:4720
User_LinuxSyscallHandler (fullFrame=0x[MASKED]) at bora/vmkernel/user/user.c:2109
gate_entry () at bora/vmkernel/main/x86/gates64.S:175ESXi 8.x
This issue is caused by a physical hardware memory defect (RAM fault) and is not a software bug within ESXi or VMware vDefend Firewall.
The failing physical RAM generates excessive Corrected Machine Check errors (indicated by the HEST GESB / GEDE ECC records with severity 0x2 and continuous ApeiPageRetire messages). When the physical server's motherboard detects these faults, it triggers System Management Interrupts (SMIs) to correct the errors and safely retire the bad memory pages. SMIs execute directly in the hardware firmware and temporarily freeze the ESXi hypervisor.
In this scenario, the underlying memory degradation and the extensive hardware-level interrupts corrupted the memory structure of the VSIP module's radix tree. When pfr_enqueue_addrs attempted to walk the tree, the CPU locked up. Because the CPU was unresponsive for an extended period, the ESXi kernel's heartbeat monitor detected the lockup and triggered an NMI IPI PSOD to halt the system and prevent further data corruption.
Additionally, PSOD can also occur when the thread fails to update its heartbeat for an extended period due to taking excessive time while holding spinlocks. Multiple threads run slowly, and a large volume of memory errors are logged concurrently.
Proactively migrate (vMotion) all active virtual machines and place the affected ESXi host into Maintenance Mode.
Check the server's out-of-band management logs (iDRAC/iLO/CIMC) to locate the physical DIMM slot matching the memory addresses reported in the vmkernel.log.
Update the server's BIOS/UEFI firmware to the latest version and contact your OEM hardware vendor to replace the defective physical memory module.
For guidance or if help is needed validating hardware logs, see Contact Support.