Purple Screen of Death (PSOD) occurs on an ESXi host within a vSAN cluster.
Version Details: VMware ESXi 8.0.3 build-24585383Panic Message: @BlueScreen: NMI IPI: Panic requested by another PCPU. PC 0x420017690e24, SP 0x453bc321ba28 (Src 0x4, CPU0)
cpu107:2098596)NVMEDEV:9363 recover controller 256cpu107:2098596)NvmeDiscover: 6804: Scan operation 2 received on adapter vmhba1cpu107:2098596)NvmeDiscover: 4724: Controller nqn.2021-08.com.intel:############## fuseOp 0 oncs 4e cmic 0 nscnt 80cpu92:2097959)NvmeDeviceIO: 1973: cleanup from TM handler worldcpu92:2097959)NvmeDeviceIO: 150: StuckIoCounter for t10.NVMe____####_Ent_NVMe_#####_MU_U.2_1.6TB________################## : 0. Clearing PSA_STOR_DEVICE_FLAG_STUCK_IO_CONDcpu92:2097959)NvmeUtil: 428: Transient status for command 0x1 set to VMK_TIMEOUT because the timeout has expired: cmdId.initiator=0x430b28edf500 cmdId.serialNumber=0xb0828277)
PsaNVMe_AsyncTokenIODone@vmkernel#nover+0x76 PsaNvmeDeviceTimeoutHandlerFn@vmkernel#nover+0x3b2 Lock_CheckSpinCount@vmkernel#nover+0x157
[vob.vsan.lsom.stuckiotimeout] vSAN device #### detected I/O timeout error.
A software interaction between the vSAN kernel and the NVMe storage stack fails to handle stalled I/O commands exceeding 120 seconds.
While attempting to clean up these stalled tasks, the system enters a CPU spinlock deadlock (Lock_CheckSpinCount), triggering a kernel panic to prevent data corruption.
Fixed in ESXi 8.0 Update 3e (build 24674464) and higher.
Schedule a maintenance window for the affected ESXi hosts.
Update the hosts to the fixed version or higher.
See https://knowledge.broadcom.com/external/article/142814/download-broadcom-products-and-software.html for steps to download this release.
This issue is tracked via PR 3473626. For further assistance, speak with a customer representative or a Support Engineer see Contact Support. Scroll to the bottom of the page and click on your respective region.