Virtual machines configured with NVIDIA vGPUs may experience unresponsiveness, disconnects, or power-off events during vMotion operations. This behaviour correlates with specific NVIDIA driver versions and vMotion timeouts, typically presenting as migration timeouts during the memory precopy phase.
We see the following error:
/var/run/log/vmkernel.log file of the ESXi. <Time Stamp> In(182) vmkernel: cpu37:2110687)NVRM: Xid (PCI:0000:86:00): 109, pid=2111064, name=, channel 0x0000006a, errorString CTX SWITCH TIMEOUT, Info 0x4c433d
VMotionInitiateSrc: Migration failed after VM memory precopy.
NVIDIA GPU driver version 580.126.08 fails to return control to the VMware stack during the vMotion notification phase. This results in a migration timeout and a failure to send data, potentially rendering the ESXi host unresponsive.
Verify the installed NVIDIA driver version and upgrade the driver to vGPU 20.2 (595.91.02), otherwise Engage your hardware vendor to validate the environment configuration and ensure the driver upgrade is applied correctly.
Workaround: If vMotion issues persist, disable "Passthrough VM DRS Automation" in the Cluster Advanced Settings (Cluster -> Configure -> vSphere DRS) to prevent automatic migrations until the driver issue is addressed.