NVMe storage devices may report critical S.M.A.R.T. health warnings (specifically regarding power-loss protection) without triggering automatic disk evacuation or isolation by the vSAN health management service. This document explains the behavior and appropriate manual remediation steps.
iDRAC, iLO, XCC) report hardware drive failures.
[vob.vsan.lsom.backupfailednvmediskhealthcriticalwarning] NVMe critical health warning for disk <disk_id> is: The disk's backup device has failed.esxcli nvme device log smart get) indicates: Volatile Memory Backup Device Failure: true[root@ESXi:~] esxcli nvme device log smart get -A vmhba1SMART And Health Info: Available Spare Space Below Threshold: false Temperature Warning: false NVM Subsystem Reliability Degradation: false Read Only Mode: false Volatile Memory Backup Device Failure: true Composite Temperature: 306 K Available Spare: 100 % Available Spare Threshold: 10 % Percentage Used: 0 % Data Units Read: 0x60ea0528 Data Units Written: 0x2f8fbda9 Host Read Commands: 0x27f7a927fb Host Write Commands: 0x12f8084edf Controller Busy Time: 0x13522 Power Cycles: 0x1a Power On Hours: 0x3f91 Unsafe Shutdowns: 0x9 Media Errors: 0x0 Number of Error Info Log Entries: 0x2c Warning Composite Temperature Time: 0 Mins Critical Composite Temperature Time: 0 Mins Temperature Sensor 1: 319 K Temperature Sensor 2: 309 K Temperature Sensor 3: 0 K Temperature Sensor 4: 0 K Temperature Sensor 5: 0 K Temperature Sensor 6: 0 K Temperature Sensor 7: 0 K Temperature Sensor 8: 0 K
vsandevicemonitord.log:2026-02-26T20:00:06Z In(14) vsandevicemonitord[2100825]: [70509974144]: WARNING - NVMe critical health warning for disk t10.NVMe____Dell_NVMe_ISE_PS1030_MU_U.2_6.4TB_______############# is: 'The disk's backup device has failed'.2026-02-26T20:10:07Z In(14) vsandevicemonitord[2100825]: [70509974144]: WARNING - NVMe critical health warning for disk t10.NVMe____Dell_NVMe_ISE_PS1030_MU_U.2_6.4TB_______############# is: 'The disk's backup device has failed'.
esxcli nvme device log smart get -A vmhba1[root@ESXi:~] esxcli storage core device smart get -d t10.NVMe____Dell_NVMe_ISE_PS1030_MU_U.2_6.4TB_______#####################Parameter Value Threshold Worst Raw------------------------ ------- --------- ----- ---Health Status WARNING N/A N/A N/APower-on Hours 16273 N/A N/A N/APower Cycle Count 26 N/A N/A N/AReallocated Sector Count 0 90 N/A N/ADrive Temperature 33 75 N/A N/AvSAN does not automatically isolate NVMe drives based solely on vendor-specific S.M.A.R.T. metrics (such as a volatile memory backup failure). These metrics currently lack industry-wide standardization, and automated responses risk creating false-positive failures that cause unnecessary cluster-wide data rebuilds and performance degradation. vSAN logic specifically monitors for NVM Subsystem Reliability Degradation flags.
esxcli nvme device log smart get -A <vmhba_id> Verify if Volatile Memory Backup Device Failure is set to true.vob.vsan.lsom.backupfailednvmediskhealthcriticalwarningCritical Warning (CWARN): This field indicates critical warnings for the Controller.
The value of this field shall indicate the value of the Critical Warning field in the Controller’s SMART / Health Information log page.
Volatile Memory Backup Failed (VMBF): This bit shall indicate the same value as the Volatile Memory Backup Failed (VMBF) bit (i.e., bit 4) in the Critical Warning field in the Controller’s SMART / Health Information log page.