ESXi hosts may intermittently enter a "Not Responding" state in vCenter Server, characterized by management agent (hostd) failures, storage heartbeats timeouts, and vSAN component congestion. This article outlines the troubleshooting steps when standard host-level fixes do not resolve the issue, focusing on identifying underlying upstream network fabric errors.
VMware vSphere ESXi 8.x
VMware vSAN 8.x
The host disconnections are caused by the hostd service failing due to hanging storage and vSAN component congestion, which are secondary symptoms of severe, underlying network TCP/IP errors (such as packet drops, out-of-order packets, or congestion) affecting the host.
If standard VMware troubleshooting steps—such as disabling Tx Pause Frames, swapping physical cables, and changing switchports—have been performed and the issue persists: