Shared pRDM disks for Windows Failover Cluster (WSFC) become unwritable with SCSI errors or 'BUSY' status following a successful vMotion
search cancel

Shared pRDM disks for Windows Failover Cluster (WSFC) become unwritable with SCSI errors or 'BUSY' status following a successful vMotion

book

Article ID: 444271

calendar_today

Updated On:

Products

VMware vSphere ESXi

Issue/Introduction

In a VMware vSphere environment hosting a Microsoft Windows Failover Cluster (WSFC) with shared physical Raw Device Mappings (pRDMs), you observe the following behavior:

  • A virtual machine hosting a cluster node undergoes a successful, live vMotion (triggered manually or automatically by vSphere DRS).

  • Immediately after the vMotion completes, the shared pRDM disks inside the Guest OS report a "BUSY" status, and the cluster resources may experience extreme I/O latency, delayed writes, or complete I/O failure.

  • The physical SAN logs a high volume of SCSI error frames and SCSI Reservation Conflicts targeting those specific pRDM LUNs during this timeframe.

  • Performing a manual cluster failover of the affected resources to another cluster node immediately resolves the I/O blockage. Failing the resources back to the original node does not reintroduce the issue.

  • The issue may not be seen in every vMotion operation.

Environment

VMware vSphere All Versions

Cause

This may occur if the devices are not configured with a PSP recommended by the storage vendor.

Resolution

Confirm with the storage vendor what is their recommendation on PSP best practices for shared RDMs.