A VMware vSAN cluster reports extreme disk group read latency (diskgroup.lat.avg.read) and Skyline Health "Operational Health" alarms. This technical guide explains how to identify physical hardware failure via SMART data and SCSI sense codes, and provides the standard drive replacement procedure for vSAN OSA environments.
diskgroup.lat.avg.read reported as 45105.00)./var/run/log/vmkernel.log) show read failures and SCSI medium errorsWARNING: PLOG: PLOGProbeDevice:6851: Failed to read the device : Not found
esxcli storage core device smart get -d <device_ID>.VMware vSAN 8.0U3 (OSA)
Physical hardware failure or communication degradation at the physical layer prevents the storage device from reading data from the specified Logical Block Address (LBA). This indicates that the block is unreadable and the error could not be corrected by the device. As SSDs degrade over time, vSAN fails out the disk or disk group when a failure to read is experienced in the metadata region.
Involve the hardware vendor for further investigation and replacement of the failing physical disk drive.