vSAN device encountered unrecoverable read error causing recurring unhealthy disk group
search cancel

vSAN device encountered unrecoverable read error causing recurring unhealthy disk group

book

Article ID: 448263

calendar_today

Updated On:

Products

VMware vSAN 8.x

Issue/Introduction

  • A vSAN disk group enters an unhealthy state shortly after creation or hardware replacement.

  • The vbd logs show that the specific device has encountered an unrecoverable read error
    .YYYY-MM-DDTHH:MM:SS In(14) vobd[2098628]:  [vSANCorrelator] 250132350271us: [vob.vsan.lsom.metadataURE] vSAN device ########-####-####-####-######## encountered unrecoverable read error. This disk will be evacuated and rebuilt.If the device is part of a dedup disk group, the entire disk group will be evacuated and rebuilt.
    YYYY-MM-DDTHH:MM:SS In(14) vobd[2098628]:  [vSANCorrelator] 250132350303us: [vob.vsan.lsom.diskunhealthy] vSAN device ########-####-####-####-######## is unhealthy.

  • Following each hardware MEDIUM error, vSAN attempts to self-heal by automatically rebuilding the affected disk group. While the software successfully recovers the disk group temporarily, it immediately crashes again when it hits the unreadable physical sectors on the faulty drive.
    YYYY-MM-DDTHH:MM:SS In(14) vobd[2098628]:  [vSANCorrelator] 5007613997us: [esx.audit.vob.vsan.lsom.diskgrouprebuild] Diskgroup naa.################ is rebuilt successfully after MEDIUM error. Old UUID XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX New UUID YYYYYYYY-YYYY-YYYY-YYYY-YYYYYYYYYYYY.
    YYYY-MM-DDTHH:MM:SS In(14) vobd[2098628]:  The event ([esx.audit.vob.vsan.lsom.diskgrouprebuild] Diskgroup naa.################ is rebuilt successfully after MEDIUM error. Old UUID XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX New UUID YYYYYYYY-YYYY-YYYY-YYYY-YYYYYYYYYYYY.) was sent immediately to hostd;

  • The vsandevicemonitor.log excerpts show Latent Sector Error (LSE) detection
    YYYY-MM-DDTHH:MM:SS In(14) vsandevicemonitord[2100330]: [712821068416]: Device naa.################ state is DG_PROPAGATED_UNHEALTHY_BY_LSE
    YYYY-MM-DDTHH:MM:SS In(14) vsandevicemonitord[2100330]: [712821068416]: Device naa.################ state is DISK_UNHEALTHY_BY_LSE
    YYYY-MM-DDTHH:MM:SS In(14) vsandevicemonitord[2100330]: [712821068416]: Device naa.################ state is DG_PROPAGATED_UNHEALTHY_BY_LSE
    YYYY-MM-DDTHH:MM:SS In(14) vsandevicemonitord[2100330]: [712821068416]: Device naa.################ state is DG_PROPAGATED_UNHEALTHY_BY_LSE

Environment

  • VMware vSAN 8.x

Cause

A Latent Sector Error (LSE) or Unrecoverable Read Error (URE) occurs within the vSAN metadata region of a capacity/cache drive. In deduplication-enabled disk groups, the loss of a single metadata sector forces the entire group into an unhealthy state to ensure the data integrity.

Resolution

Recurring failures in the same physical slot despite disk replacement indicate a potential fault in the disk slot, SAS backplane, cables, or controller.

A deep-dive hardware inspection should be performed by the vendor; specifically targeting the disks, SAS backplane, controller cables, and the physical slot associated with the affected diskgroup.