"vCenter Server is unable to find a master vSphere HA agent" for cluster after vCenter upgrade or snapshot revert
search cancel

"vCenter Server is unable to find a master vSphere HA agent" for cluster after vCenter upgrade or snapshot revert

book

Article ID: 449121

calendar_today

Updated On:

Products

VMware vCenter Server

Issue/Introduction

After upgrading vCenter Server or reverting to a vCenter Server snapshot, vSphere High Availability (HA) fails to function properly on the cluster.

The following symptoms are observed:

  • The warning "vCenter Server is unable to find a master vSphere HA agent" or "Cannot find vSphere HA master agent" is reported under the cluster's All Issues tab.

  • One or more ESXi hosts in the cluster show a "Not Responding" status in the vSphere Client inventory.

  • The "Remediate HA" or HA reconfiguration task for the cluster gets stuck or fails to complete.

  • The vSphere HA state for hosts in the cluster remains stuck in an "Uninitialized" status.

Environment

VMware vCenter 8.x

Cause

An ESXi host in a "Not Responding" state blocks the cluster-level vSphere HA remediation process. Because vCenter cannot communicate with the unresponsive host to check or reconfigure its HA agent, the HA master election process cannot complete, leaving the operational hosts stuck in an "Uninitialized" state.

Resolution

Caution: Always take a fresh file-based backup or an snapshot of the vCenter Server Appliance (VCSA) before performing direct modifications to the vCenter database (VCDB).

Step 1: Manually Disconnect the Unresponsive Host in the vCenter Database

     

    Note: Contact Broadcom Technical Support for assistance. See Contact Support.

 

Step 2: Reconfigure vSphere HA

  1. Log in to the vSphere Client.

  2. Confirm that the unresponsive ESXi host now appears as Disconnected.

  3. Monitor and wait the "Remediate HA" task to complete. For any remaining host that displays an HA issue, right-click the host in the inventory and select Reconfigure for vSphere HA.

  4. Troubleshoot and resolve the underlying network or host issue on the disconnected host before reconnecting it to the cluster.