ESXi host crashes or loses configuration persistence due to local boot storage device connectivity failure
search cancel

ESXi host crashes or loses configuration persistence due to local boot storage device connectivity failure

book

Article ID: 442482

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

An ESXi host experiences an unexpected reboot or kernel panic (PSOD) following a severe local storage connectivity loss. Prior to the crash, management agents raise alerts regarding configuration persistence and local datastores become completely inaccessible.

Symptoms include:

  • Error message: Lost connectivity to the device <REDACTED_MAC_ADDRESS> backing the boot filesystem; /vmfs/devices/disks/<REDACTED_MAC_ADDRESS>. As a result, host configuration changes will not be saved to persistent storage.

  • Local VMFS datastores residing on the boot drive enter an inactive or unreachable state.

  • Storage subsystem logs entry reflecting All Paths Down (APD) or Permanent Device Loss (PDL) for the local NAA boot ID.

  • Intermittent log data gaps occur right before the unexpected system reboot due to the scratch partition becoming unwriteable.

Environment

vSphere ESXi 8.0 Update 3

Cause

A hardware failure within the local storage subsystem controller (RAID card) causes a total loss of connectivity to the physical disk backing both the ESXi boot filesystem and local datastore partitions.

Resolution

To resolve this issue, coordinate a hardware maintenance window to replace the faulty storage components:

  1. Verify Hardware Fault: Review the server out-of-band management console (e.g., HPE iLO, Dell iDRAC) or engage the Original Equipment Manufacturer (OEM) to confirm diagnostic faults targeting the RAID controller or storage backplane.

  2. Evacuate the Host: Place the affected ESXi host into Maintenance Mode. If there are any remaining active virtual machines residing on other accessible datastores, migrate them to an alternate operational host within the cluster.

  3. Power Down Server: Gracefully shut down the ESXi host via the vSphere Client or the command-line interface.

  4. Replace Faulty Component: Work with the OEM hardware technician to replace the defective RAID controller card.

  5. Restore and Validate: Power on the physical server, ensure the disk array profile is imported correctly within the controller BIOS/UEFI, boot into ESXi, and verify that both configuration changes persist and local VMFS datastores auto-mount normally.

Additional Information

For instructions on parsing specific SCSI sense codes within storage logs prior to hardware failure, see "Interpreting SCSI sense codes in VMware ESXi"