A vSAN disk enters a permanent error state when cumulative I/O retry timeouts or SMART read errors exceed internal monitoring thresholds. This article provides steps to verify the failure and initiate replacement.
/var/log/vobd.log" of ESXi host will also report the disk is under permanent error:YYYY-MM-DDTHH:MM:SSZ:[vSANCorrelator] 12219239351514u3:[esx.problem.vob.vsan.pdl.offline] vSAN device ########-####-####-####-########### has gone offline.YYYY-MM-DDTHH:MM:SSZ:[vSANCorrelator] 12219239351514u3:[esx.problem.vob.vsan.lsom.devicerepair] Device e########-####-####-####-########### is in offline state and is getting repaired.YYYY-MM-DDTHH:MM:SSZ:[vSANCorrelator] 12219239351514u3:[vob.vsan.pdl.offline] vSAN device e########-####-####-####-########### has gone offline.YYYY-MM-DDTHH:MM:SSZ:[vSANCorrelator] 12219239351514u3:[esx.problem. vob.vsan.pdl.offline] vSAN device e########-####-####-####-########### has gone offline.YYYY-MM-DDTHH:MM:SSZ:[vSANCorrelator] 12219239351514u3:[esx.problem.vob.vsan.lsom.diskerror] vSAN device #### is under permanent error.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2098148]: [vSANCorrelator] 9883129325851us: [vob.vsan.pdl.offline] vSAN device 5####3c-#####5##d-d####3-######1 has gone offline.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2098148]: [APDCorrelator] 9882990619157us: [esx.problem.storage.apd.start] Device or filesystem with identifier [##########] has entered the All Paths Down state.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2098148]: [vSANCorrelator] 9882990619196us: [esx.problem.vob.vsan.pdl.offline] vSAN device 5####c-1626-####-####-f01#####1 has gone offline.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2098148]: [psastorCorrelator] 9882990619971us: [esx.problem.storage.connectivity.lost] Lost connectivity to storage device #############. Path vmhba0:C0:T0:L0 is down. Affected datastores: Unknown.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2098148]: [psastorCorrelator] 9883129325791us: [vob.psastor.device.state.permanentloss] Device :eui.############ has been removed or is permanently inaccessible.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2098148]: [psastorCorrelator] 9882990620278us: [esx.problem.psastor.device.state.permanentloss] Device: eui.############### has been removed or is permanently inaccessible. Affected datastores (if any): Unknown.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2098148]: [scsiCorrelator] 9882990620278us: [esx.problem.scsi.device.state.permanentloss]] Device: eui.############### has been removed or is permanently inaccessible. Affected datastores (if any): Unknown.H:0x1) event logged for the disk:YYYY-MM-DDTHH:MM:SSZ In(182) vmkernel: cpu88:2098876)ScsiDeviceIO: 4605: Cmd(0x45dac1e2ee80) 0x25, CmdSN 0xf5a2b42 from world 0 to dev "naa.###############" failed H:0x1 D:0x0 P:0x0
var/run/log/vobd.log, you will see below entries:YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2097812]: [vSANCorrelator] 690804744us: [esx.problem.vob.vsan.lsom.diskunhealthy] vSAN device 528c1c3d-ddd2-7210-c721-############ is unhealthy.YYYY-MM-DDTHH:MM:SSZ In(14) vobd[2097812]: [vSANCorrelator] 690811954us: [esx.problem.vob.vsan.lsom.diskunhealthy] vSAN device 5207b712-e8a7-2b82-3283-############ is unhealthy.
ocalcli storage core device smart get -d <device> shows a high Read Error Count (e.g., 1,320,182)./var/log/vsandevicemonitord.log indicates the disk resilience state: Device <naa.ID> state is DISK_UNDER_PERM_ERROR.[root@esxi01:~] localcli storage core device smart get -d naa.###################Parameter Value Threshold Worst Raw------------------------ ----------------- --------- ----- ---Health Status IMPENDING FAILURE N/A N/A N/AMedia Wearout Indicator 0 100 N/A N/AWrite Error Count 0 N/A N/A N/ARead Error Count 0 N/A N/A N/APower Cycle Count 0 N/A N/A N/AReallocated Sector Count 0 N/A N/A N/ADrive Temperature 27 N/A N/A N/AWrite Sectors TOT Count 5239103362346 N/A N/A N/ARead Sectors TOT Count 2353051044834 N/A N/A N/AProgram Fail Count 0 N/A N/A N/AErase Fail Count 0 N/A N/A N/A
The vsandevicemonitord service transitions a disk to DISK_UNDER_PERM_ERROR when hardware-level I/O failures exceed threshold limits, necessitating replacement to maintain cluster redundancy
Identify the failing disk UUID via the vCenter UI or by running esxcli vsan storage list.
Validate storage controller driver and firmware versions against the VMware Compatibility Guide.
Check the SMART health status using the command: localcli storage core device smart get -d <naa-id>.
Engage the hardware vendor to perform physical diagnostics and confirm media failure.
Replace the defective hardware using the procedure specific to your architecture:
vSAN OSA: Follow the procedure in Replace a Capacity Device (vSAN OSA)
vSAN ESA: Follow the procedure in Replace a Storage Pool Device (vSAN ESA)
Verify the new disk is added to the disk group and monitor the resync progress in the vCenter UI.
To speak with a customer representative or a Support Engineer see Contact Support. Scroll to the bottom of the page and click on your respective region.