Severe Disk Latency when Powering on VMs Migrated using Dell SRDF Block Level Replication with Shared VMDKs
search cancel

Severe Disk Latency when Powering on VMs Migrated using Dell SRDF Block Level Replication with Shared VMDKs

book

Article ID: 454163

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

After migrating clustered VMs with shared VMDKs using Dell SRDF to NVME over TCP VMFS datastores, disk latency is severely high upon powering on the VMs, making them unusable.

ESXTOP shows extremely high KAVG, GAVG, QAVG

Both the VM and the host go into a hung state, requiring a reboot of the host to get out of the hung state

Host vmkernel.log contain the below messages:

2026-07-20T18:32:38.897Z -WARNING vmkwarning - [esx@4413] cpu44:2099300) NvmeUtil: 155: Error on Cmd(0x459c12691e40) 0xe, isAdmin:0 CmdSN 0xec from world 2126364 to component "eui.#############"  H:0x0 D:0x371 P:0x0 Cancelled from driver layer
2026-07-20T18:32:38.897Z -INFO vmkernel - [esx@4413] cpu7:2100269)nvmetcp:2234 [ctlr 263, queue 11] data transfer of txPdu 0x432427c5ab80 for rxPdu 0x453add79beb0 not finished. Expected 16448, transferred 1408: Failure
2026-07-20T18:32:38.897Z -INFO vmkernel - [esx@4413] cpu7:2100269)nvmetcp:3559 [ctlr 263, queue 11] failed to process txPdu 0x432427c5ab80: Failure
2026-07-20T18:32:38.897Z -INFO vmkernel - [esx@4413] cpu7:2100269)nvmetcp:3585 [ctlr 263, queue 11] Fatal transport error, rxPdu 0x453add79beb0, fes 2, fei 0: Failure
2026-07-20T18:32:38.897Z -INFO vmkernel - [esx@4413] cpu7:2100269)nvmetcp:7078 [ctlr 263, queue 11] Fatal transport error; fes: 2, fei: 0 

vmkernel.log will also be flooded with the below error:

2026-07-20T18:40:06.756Z -WARNING vmkwarning - [esx@4413] cpu25:2130959) NVMEIO:2082 Transport driver failed to submit cmd 0x457b5dd0ebc0, C: nqn.1988-11.com.dell:PowerMax_8500:00:000120203823#vmhba66#10.##.###.11:4420, Q: 4 <0xbad0001>.

Note: The preceding log excerpts are only examples. Date, time, and environmental variables may vary depending on your environment.

This only impacts migrated clustered VMs via Dell SRDF. All other instances of VMs work without issue

Environment

VMware vSphere ESXi (All Versions)

NVMe over TCP

Dell PowerMax

Cause

Clustered VMDK over NVMe/TCP on PowerMax has not been tested and is not currently a supported configuration. As such the storage target is failing to set the LAST_PDU flag on the C2HData PDU prior to sending the command completion.

This is still under investigation by Dell engineering.

Resolution

Engage Dell support for further assistance.