ESXi host disconnects intermittently from vCenter Server
search cancel

ESXi host disconnects intermittently from vCenter Server

book

Article ID: 318647

calendar_today

Updated On:

Products

VMware vCenter Server VMware vSphere ESXi

Issue/Introduction

  • ESXi hosts frequently disconnect from vCenter Server and appear as "Not Responding" on the vSphere Client

  • Reconnecting an affected ESXi host fails with error: Cannot contact the specified host <HostName>. The host may not be available on the network, a network configuration problem may exist, or the management services on this host may not be responding.

  • If affected ESXi host is part of a vSAN cluster, Skyline Health (Cluster > Monitor > Skyline Health) may display the alert 'Hosts disconnected from VC'



  • On the vCenter, /var/log/vmware/vpxd/vpxd.log  shows entries similar to below pointing to Missing Heatbeats : 

    YYYY-MM-DDThh:mm:ss  [VpxdIntHost] Missed #### heartbeats for host <HostName>
    YYYY-MM-DDThh:mm:ss  info vpxd[####] [Originator@6876 sub=HostCnx opID=CheckforMissingHeartbeats-####] [VpxdHostCnx] No heartbeats received from host; cnx: ####-####-####, h: host-####, time since last heartbeat: ####
    YYYY-MM-DDThh:mm:ss  info vpxd[####] [Originator@6876 sub=HostCnx opID=CheckforMissingHeartbeats-####] Marking the connection alive to false: ####-####-####-####
    YYYY-MM-DDThh:mm:ss  info vpxd[####] [Originator@6876 sub=InvtHostCnx opID=CheckforMissingHeartbeats-####] Got lost connection callback for host-####
    YYYY-MM-DDThh:mm:ss  warning vpxd[####] [Originator@6876 sub=InvtHostCnx opID=HostSync-host-####-########] Connection not alive due to missing heartbeats; [vim.HostSystem:host-#####,<HOST FQDN>], cnx: ####-####-####

  • Additionally the vpxd.log may contain SOAP HTTP failure when attempting to communicate with affected host's vpxa similar to below: 

    YYYY-MM-DDThh:mm:ss info vpxd[####] [Originator@6876 sub=vmomi.soapStub[####] opID=HB-host-#####@####-####] SOAP request returned HTTP failure; <<io_obj p:####, h:132, <UNIX ''>, <UNIX '/var/run/envoy-hgw/hgw-pipe'>>, /hgw/host-#####/vpxa>, method: getChanges; code: 503(Service Unavailable); fault: (null)
    YYYY-MM-DDThh:mm:ss error vpxd[####] [Originator@6876 sub=Vmomi opID=HB-host-#####@####-####] Got vmacore exception when invoking VMOMI method; <</hgw/host-#####>, /vpxa>, vpxapi.VpxaService.getChanges, N7Vmacore4Http13HttpExceptionE(HTTP error response: Service Unavailable)

  • This behavior typically indicates intermittent loss of UDP heartbeat traffic between the ESXi hosts and the vCenter Server, often due to network congestion, packet drops, or firewall configuration issues.

Environment

VMware vCenter Server

VMware vSphere ESXi

Cause

  • This issue arises when vCenter Server fails to receive the UDP heartbeat messages sent by an ESXi host. ESXi hosts transmit these heartbeats every 10 seconds, and vCenter expects to receive at least one within a 60-second window. If no heartbeat is received during this period, the host is marked as "not responding." This behavior may indicate network congestion, packet loss, or a misconfigured firewall between the ESXi host and the vCenter Server.

    Note: If the host consistently disconnects at 60-second intervals, it is a strong indication that UDP port 902 traffic from the ESXi host to the vCenter Server is being blocked most commonly by a firewall. Check if firewalls are correctly resolving the FQDN of the vCenter.

  • This issue can also occur if the vCenter is subject to image based backups from tools like Commvault/Veeam/Rubrik. During snapshot creation or deletion, prolonged stun times can cause the virtual machine to freeze, and not acknowledging UDP 902 packets, hosts will report disconnected. This is most common when vCenter is included in a backup job that contains other VM's where it is busy processing multiple workloads.

  • The issue can also occur if there are security or network tools that scan the management network and cause temporary network interruptions over port TCP 443 or UDP 902 between vCenter Server and the ESXi hosts. 

Resolution

  1. Validate ESXi Heartbeat Communication to vCenter

    To confirm that the affected ESXi host is sending heartbeat packets to the vCenter Server every 10 seconds over UDP port 902, perform the following checks. Note that the vCenter is not expected to reply to the heartbeat packets.
    1. Verify Heartbeat Transmission from ESXi Host

      1. SSH to the affected ESXi Host

      2. Get the physical NIC bound to the management vmkernel port by running esxtop and pressing the n key for the network view:


      3. Perform a live packet review with the following command to confirm the ESXi host is actively sending heartbeat packets to the specified vCenter Server:

        pktcap-uw --uplink vmnicX --capture UplinkSndKernel --udpport 902 -o -| tcpdump-uw -enr -

        Replace vmnicX with the vmnic used for Management vmernel port on ESXi host

    2. Verify Heartbeat Reception on vCenter Server

      1. From an SSH session on the vCenter Server Appliance (VCSA), run the following command to verify the vCenter Server is receiving heartbeat packets from the affected ESXi host:

        vcsa# tcpdump src host <esxi_host_ip_address> and udp port 902

      2. If the packets are not seen on the vCenter, then confirm if they are being received on the physical NIC of the host the vCenter VM is running on.
        1. Get the physical NIC which is bound to the vCenter virtual switch port:

          esxcli network vm list
          esxcli network vm port list -w <world_id_of_vcenter_vm_obtained_from_the_above_command>

        2. Run the following packet capture on the physical NIC:

          pktcap-uw --uplink vmnicX --capture UplinkRcvKernel --udpport 902 -o -| tcpdump-uw -enr - | grep <Affected_source_host_IP>

          Replace <Affected_source_host_IP> with IP of the affected Host which is transmitting packet, used in step 1
          Replace vmnicX with the vmnic used for Management vmernel port on ESXi host where the vCenter is hosted

  1. Based on Test Results, troubleshoot further at affected ESX host, Network or vCenter level 
    1. Heartbeats sent from ESXi but not receivedon vCenter:

      1. Investigate the network path between the ESXi host and vCenter for potential issues such as firewalls, ACLs, or other traffic filtering mechanisms blocking UDP port 902.

      2. Ensure that external firewalls are correctly resolving the FQDN of the vCenter.

      3. Ensure that TCP ports 443 and 902 are open for bidirectional communication between ESXi hosts and vCenter.

    2. Heartbeats not sent from ESXi:

      Inspect ESXi host services and relevant log files to identify and resolve the root cause. It is recommended to engage Broadcom VCF Technical Support for further assistance

    3. Heartbeats sent from ESXi and received successfully on vCenter:

      If heartbeat packets are reaching vCenter but hosts are still disconnecting, the issue is likely not network-related. Further investigation into the vCenter Server itself is recommended by engaging Broadcom VCF Technical Support

 

Temporary Workaround : Increasing VPXD Heartbeat Timeout

In case the number of missed heartbeats are less than or equal to 6 for the concerned esxi host(s), as a short-term measure, increase the heartbeat timeout value in vCenter Server to allow more time for ESXi host heartbeat responses.

NOTE: This should be treated as a temporary workaround .It is strongly recommended to engage your Networking Team to identify and resolve the underlying network issue.

  1. Open the vSphere UI in a web browser and log in as administrator

  2. Select the vCenter object from the Hosts and Clusters inventory.

  3. Navigate to the Configure tab.

  4. Under Settings, select Advanced Settings.

  5. Click Edit.

  6. In the Key field, enter:

    config.vpxd.heartbeat.notRespondingTimeout
  7. In the Value field, enter:

    120

    (adjust this value as needed.)

  8. Click Add, then OK to save the change.

  9. Restart the vCenter Server service by running:

    vcsa# service-control --stop vmware-vpxd && service-control --start vmware-vpxd

 

Additional Information

If the issue persists despite making all the above changes or to troubleshoot on the Network PCAP, open a support case with Broadcom Support and refer to this KB article.

For more information, see Creating and managing Broadcom cases