Troubleshooting CONTROLLER_STATE_FAILED Errors on Dell BOSS Cards (VMware vSphere ESXi)
search cancel

Troubleshooting CONTROLLER_STATE_FAILED Errors on Dell BOSS Cards (VMware vSphere ESXi)

book

Article ID: 452171

calendar_today

Updated On:

Products

VMware vSphere ESXi VMware vCenter Server

Issue/Introduction

This article provides diagnostic steps for resolving ESXi host performance degradation and storage failures caused by terminal controller states on Dell BOSS (Boot Optimized Storage Solution) hardware.

    • ESXi host fails to complete virtual machine backups.
    • Virtual machine migrations (vMotion) fail to initialize or stall indefinitely.
    • The following errors appear in vmkernel.log:

      WARNING: NVMEIO:#### Controller #### in state 9 or in recovery mode WARNING: NvmeDiscover: ####: Admin NVMe Command 0x6 is being aborted CONTROLLER_STATE_FAILED
      NvmeDiscover: 8341: subsystem wide controller probe still in progress

Environment

  • ESX 9.x
  • vCenter Server 9.x

Cause

This issue typically stems from a hardware-level failure of the storage controller, such as a Dell BOSS (Boot Optimized Storage Solution) M.2 card. When the controller enters a failed state, the ESXi NvmeDiscover process repeatedly fails to initialize the device. This resource contention blocks vMotion and backup operations that require consistent storage access.

Resolution

  1. Review System Logs: Check the /var/log/vmkernel.log for recurring NVMe aborts and controller state warnings. If these errors are prevalent, the issue is likely hardware-related rather than a configuration error.

  2. Run Hardware Diagnostics:

    • Reboot the affected ESXi host to verify if the failure persists after a power cycle.

    • Log in to the server management console (e.g., iDRAC) to execute system diagnostics. Run a full hardware health check specifically targeting the storage backplane and the M.2 BOSS controller.
  3. Coordinate with Hardware Vendor: If diagnostics confirm physical failure of the controller or the M.2 drive, contact your hardware vendor (e.g., Dell).

  4. Restore Redundancy: Once the faulty hardware is replaced, confirm that the system initializes correctly and that storage redundancy is fully restored