Virtual machines residing on specific Dell PowerMax (PMAX) storage devices may become unresponsive or experience extreme performance degradation. This issue is characterized by high storage latency (up to 7000ms) and repeated SCSI command failures.
Symptoms:
vmkernel.log reports frequent SCSI aborts and resets on specific LUNs.Exact Error Messages:
ScsiDeviceIO: ####: Cmd(0x################) 0x## ... failed H:0x8 D:0x0 P:0x0 (H:0x8 indicates a SCSI Abort).ScsiDeviceIO: ####: Cmd(0x################) 0x## ... failed H:0x5 D:0x0 P:0x0 (H:0x5 indicates a SCSI Reset).ScsiCore: 2000: Power-on Reset occurred on naa.################################WARNING: NMP: nmp_DeviceRequestFastDeviceProbe:###: NMP device "naa.####" state in doubtnfnic The issue is triggered by the storage array or the Fibre Channel fabric failing to process I/O requests, leading to a "state in doubt" condition for specific LUNs.
Analysis of the vmkernel.log indicates that the sequence starts with failed Test Unit Ready (TUR) commands (0x0), followed by Power-on Resets (POR) and subsequent I/O Aborts (H:0x8) and Resets (H:0x5). Since other LUNs on the same host and HBA remain unaffected, the cause is localized to the specific LUNs or the storage-side processing of those LUNs.
To resolve this issue, perform the following steps to isolate and address the storage-layer failure:
vmkernel.log or run esxtop to identify the specific naa IDs experiencing high latency and H:0x8 or H:0x5 errors.naa IDs and the PMAX Serial Number identified in the logs.nfnic) and firmware versions are within the supported compatibility range.For details on interpreting specific SCSI sense codes (e.g., 0x6 0x29 0x0), refer to: