In a VMware vSphere environment hosting a Microsoft Windows Failover Cluster (WSFC) with shared physical Raw Device Mappings (pRDMs), you observe the following behavior:
A virtual machine hosting a cluster node undergoes a successful, live vMotion (triggered manually or automatically by vSphere DRS).
Immediately after the vMotion completes, the shared pRDM disks inside the Guest OS report a "BUSY" status, and the cluster resources may experience extreme I/O latency, delayed writes, or complete I/O failure.
The physical SAN logs a high volume of SCSI error frames and SCSI Reservation Conflicts targeting those specific pRDM LUNs during this timeframe.
Performing a manual cluster failover of the affected resources to another cluster node immediately resolves the I/O blockage. Failing the resources back to the original node does not reintroduce the issue.
VMware vSphere All Versions
This may occur if the devices are not configured with a PSP recommended by the storage vendor.
Confirm with the storage vendor what is their recommendation on PSP best practices for shared RDMs.