This article addresses critical guest VM unresponsiveness and ESXi host management failures caused by severe host resource starvation. When an ESXi host experiences excessive CPU and memory overcommitment, the vSphere scheduler may become unable to process vital management tasks, leading to the failure of hostd and vpxa agents and eventual management disconnection. In vSAN environments, this resource contention can trigger metadata deadlocks and vMotion memory transfer timeouts, leaving virtual machines in an unmanageable state.
vMotion memory transfer failed or metadata deadlock in vsanmgmtdESXi (all versions)
vSAN (all versions)
Virtual Machines with High Latency Sensitivity set
Excessive CPU and memory overcommitment on the ESXi host led to a failure in the scheduler's ability to process vSAN metadata updates and vMotion memory transfers in a timely manner.
To resolve this and restore service, please follow these steps:
esxcfg-advcfg -g /Net/TcpipHeapSizeesxcfg-advcfg -s 128 /Net/TcpipHeapSizeesxtop to monitor %RDY and %WAIT metrics post-reboot to ensure resources are balanced.