This article provides diagnostic steps for resolving ESXi host performance degradation and storage failures caused by terminal controller states on Dell BOSS (Boot Optimized Storage Solution) hardware.
vmkernel.log: WARNING: NVMEIO:#### Controller #### in state 9 or in recovery mode WARNING: NvmeDiscover: ####: Admin NVMe Command 0x6 is being aborted CONTROLLER_STATE_FAILED
NvmeDiscover: 8341: subsystem wide controller probe still in progressThis issue typically stems from a hardware-level failure of the storage controller, such as a Dell BOSS (Boot Optimized Storage Solution) M.2 card. When the controller enters a failed state, the ESXi NvmeDiscover process repeatedly fails to initialize the device. This resource contention blocks vMotion and backup operations that require consistent storage access.
Review System Logs: Check the /var/log/vmkernel.log for recurring NVMe aborts and controller state warnings. If these errors are prevalent, the issue is likely hardware-related rather than a configuration error.
Run Hardware Diagnostics:
Reboot the affected ESXi host to verify if the failure persists after a power cycle.
Coordinate with Hardware Vendor: If diagnostics confirm physical failure of the controller or the M.2 drive, contact your hardware vendor (e.g., Dell).
Restore Redundancy: Once the faulty hardware is replaced, confirm that the system initializes correctly and that storage redundancy is fully restored