Network Extension HA pair degraded after upgrade
search cancel

Network Extension HA pair degraded after upgrade

book

Article ID: 391187

calendar_today

Updated On:

Products

VMware HCX

Issue/Introduction

  • HCX Network Extension (NE) appliances were upgraded.
  • NE appliance group shows DEGRADED.
  • Primary peer is ACTIVE and Standby peer is UNDECIDED as shown below:

 

  • The following messages are seen in the HCX Manager logs and in UI (common/logs/admin/app.log): Resume operation failed. Failed to exit Maintenance mode for HA Group
    20##-0#-0## ##:##:23.### UTC [InterconnectService_SvcThread-#####, IX:#######################, J:########, , TxId: ####################] ERROR c.v.v.h.s.i.h.HAGroupRedeployApplianceJob- haGroupResume failed, errorCode:null. stacktrace:null, errorMessage:HA Resume operation failed. Failed to exit Maintenance mode for HA Group

Environment

VMware HCX

Cause

The NE appliance failed to exit maintenance mode during or after the upgrade, and cannot resume HA operations.

Resolution

When an HCX HA group falls out of a Healthy state, you can typically resolve the discrepancy between the appliance status and the HCX database using the following methods.

Option 1: Perform a "Recover" Operation
The RECOVER action attempts to re-initialize the HA group and synchronize the environment.

  • Pre-check:
    • Verify Resource Availability: make sure the target ESXi host has sufficient unallocated CPU/memory for the appliance. If power-on fails, temporarily remove reservations in the Compute Profile and perform a Service Mesh Resync, before proceeding with the redeployment
  • Steps:
    • Navigate to Interconnect > Service Mesh > View Appliances.
    • Select the HA Management tab.
    • Click RECOVER. This triggers a synchronization between the physical appliance state and the HCX database.
  • Note: If this process requires redeploying appliances, there will be a brief loss of connectivity on the extended network while the new appliances initialize.

Option 2: Redeploy Appliances or Contact Support

  • Redeploy the HA Group
    Redeploy the Network Extension appliances via the Service Mesh > View Appliances menu.
    Note: This operation will cause a brief loss of connectivity on the extended network data-path while the new appliances are initialized.

    OR

  • If there are no extensions currently configured for the HA Group: Deactivate HA, reduce the NE appliances count in the Service Mesh to force the removal of the stale NE appliance.
    Then edit the Service Mesh again to increase the count back to its previous value, this will deploy a new NE appliance from scratch.
    Finally, activate HA again, which will select the new NE appliance to join the HA Group.

    OR

  • Contact Support: If redeployment is not feasible due to production constraints or 'Recover' & 'Redeploy' did not help to get the state back to normal, please gather the HCX Support bundle from both the source and destination sites along with Database. Then, open a Support Request with Broadcom for manual database correction 

Additional Information

Managing Network Extension High Availability