External Hardware Monitoring (Cisco Intersight) reported drive as "inoperable" (Fault Code F0181) while vSAN Disk Health reports same vSAN device as "Healthy"
search cancel

External Hardware Monitoring (Cisco Intersight) reported drive as "inoperable" (Fault Code F0181) while vSAN Disk Health reports same vSAN device as "Healthy"

book

Article ID: 448991

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

Upon bootup of a vSAN host: 

  • Cisco Intersight flags one or more devices as: "Storage Local disk [Id] is inoperable: reseat or replace the storage drive [Id]."
  • At the same time, vSAN Disk Management pane reports same devices as "Mounted" / "Healthy"
  • /var/run/log/vmkernel.log of ESXi host shows one or more capacity devices (MDs) changing states between 'Number ... of unannounced MDs' and 'Successfully announced VSAN MD'
  • This happens within short (couple seconds) time-frame and then stabilizes with 'Successfully announced VSAN MD' seen logged last, example:  

PLOG: PLOGMapDataPartition:2982: VSAN device 52831f27-####-####-####-############ not ready to map. Info: is stale 1, ann. 1 grpHandle 0x4500c002ee88
PLOG: PLOGMapDataPartition:2995: SSD 52fffc89-####-####-####-############ Number (1) of unannounced MDs [1]
PLOG: PLOGMapDataPartition:2982: VSAN device 52db943d-####-####-####-############ not ready to map. Info: is stale 1, ann. 1 grpHandle 0x4500c0023338
PLOG: PLOGMapDataPartition:2995: SSD 52137f20-####-####-####-############ Number (1) of unannounced MDs [1]
PLOG: PLOGInitAndAnnounceMD:10624: Successfully announced VSAN MD with UUID: 52db943d-####-####-####-############. kt 1, en 0, enC 0.
PLOG: PLOGInitAndAnnounceMD:10624: Successfully announced VSAN MD with UUID: 52831f27-####-####-####-############. kt 1, en 0, enC 0.
PLOG: PLOGMapDataPartition:2999: SSD 52fffc89-####-####-####-############ acks (1) 1 healthy MDs
PLOG: PLOGMapDataPartition:2982: VSAN device 52db943d-####-####-####-############ not ready to map. Info: is stale 1, ann. 1 grpHandle 0x4500c0023338
PLOG: PLOGMapDataPartition:2995: SSD 52137f20-####-####-####-############ Number (1) of unannounced MDs [1]
PLOG: PLOGInitAndAnnounceMD:10624: Successfully announced VSAN MD with UUID: 52db943d-####-####-####-############. kt 1, en 0, enC 0.
PLOG: PLOGInitAndAnnounceMD:10624: Successfully announced VSAN MD with UUID: 52831f27-####-####-####-############. kt 1, en 0, enC 0.
PLOG: PLOGMapDataPartition:2982: VSAN device 52831f27-####-####-####-############ not ready to map. Info: is stale 1, ann. 1 grpHandle 0x4500c002ee88
PLOG: PLOGMapDataPartition:2995: SSD 52fffc89-####-####-####-############ Number (1) of unannounced MDs [1]
PLOG: PLOGMapDataPartition:2999: SSD 52137f20-####-####-####-############ acks (1) 1 healthy MDs
PLOG: PLOGInitAndAnnounceMD:10624: Successfully announced VSAN MD with UUID: 52db943d-####-####-####-############. kt 1, en 0, enC 0.
PLOG: PLOGInitAndAnnounceMD:10624: Successfully announced VSAN MD with UUID: 52831f27-####-####-####-############. kt 1, en 0, enC 0.

Environment

vSAN OSA 8.x / 9.x

Cause

This issue is typically caused by a wear of physical disk (device nearing End of Life).
Discrepancy between vendor's monitoring vs ESXi view is caused by differences in 'failure handling', in this case:

  • Cisco monitoring detected a device as 'not ready' and declared it as faulty (Fault Code F0181 "disk ... is inoperable") at first occurrence. 
  • ESXi retried to mount/announce the device and eventually succeeds at N-th retry. 

Resolution

Broadcom recommends engaging the storage vendor to validate the current status of the device and propose appropriate remediation steps following their investigation.

Additional Information

Fault Code F0181, page 9 of: 
Cisco UCS IMC Faults Reference Guide - Storage-Related Faults