Error: diskgroup.lat.avg.read latency due to physical disk hardware failure
search cancel

Error: diskgroup.lat.avg.read latency due to physical disk hardware failure

book

Article ID: 448309

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

A VMware vSAN cluster reports extreme disk group read latency (diskgroup.lat.avg.read) and Skyline Health "Operational Health" alarms. This technical guide explains how to identify physical hardware failure via SMART data and SCSI sense codes, and provides the standard drive replacement procedure for vSAN OSA environments.

  • Disk group latency metrics show values exceeding 30,000ms (e.g., diskgroup.lat.avg.read reported as 45105.00).
  • System logs (/var/run/log/vmkernel.log) show read failures and SCSI medium errors
     

        WARNING: PLOG: PLOGProbeDevice:6851: Failed to read the device : Not found

  • SMART data for the affected disk indicates a high Read Error Count when checked via esxcli storage core device smart get -d <device_ID>.

Environment

VMware vSAN 8.0U3 (OSA)

Cause

Physical hardware failure or communication degradation at the physical layer prevents the storage device from reading data from the specified Logical Block Address (LBA). This indicates that the block is unreadable and the error could not be corrected by the device. As SSDs degrade over time, vSAN fails out the disk or disk group when a failure to read is experienced in the metadata region.

Resolution

Involve the hardware vendor for further investigation and replacement of the failing physical disk drive.