PM1735a NVMe SSD internal processor deadlock causing ESXi host unresponsiveness or reboots.
search cancel

PM1735a NVMe SSD internal processor deadlock causing ESXi host unresponsiveness or reboots.

book

Article ID: 443835

calendar_today

Updated On:

Products

VMware vSAN

Issue/Introduction

Dell PowerEdge servers equipped with Samsung PM1735a (Dell Ent NVMe v2 AGN) SSDs may experience internal processor deadlocks. These deadlocks result in storage command timeouts, cascading host unresponsiveness, and unexpected reboots.

Symptoms:

  • ESXi host enters a "Not Responding" state in vCenter Server.
  • Host may spontaneously reboot due to storage heartbeat loss.
  • The following errors are observed in vmkernel.log:
    • WARNING: NVMEIO: ... no queue available, QFULL repeated
    • WARNING: HPP: ... Error status H:0x8 D:0x0 P:0x0
    • WARNING: HBX: ... Failed to cleanup registration key on volume

Environment

  • Dell PowerEdge Servers
  • Samsung PM1735a/1733a NVMe drives with firmware version 1.1.1 and earlier
  • ESXi 7.x, 8.x

Cause

A firmware issue in the PM1735a and PM1733a internal processor causes an internal deadlock under specific I/O workloads. When the drive hangs, the ESXi storage stack (HPP/NMP) attempts to abort commands, leading to a resource exhaustion that stalls the hostd and vpxa management agents.

Resolution

Work with your hardware vendor to update the device firmware to 1.2.0 or higher as recommended by the vendor.

Additional Information

Dell advisory