HCX RAV and Bulk Migrations Fail with HttpHostConnectException on Port 8123
search cancel

HCX RAV and Bulk Migrations Fail with HttpHostConnectException on Port 8123

book

Article ID: 448301

calendar_today

Updated On:

Products

VMware HCX VMware Cloud Foundation

Issue/Introduction

  • Following the deployment of a new VMware HCX Service Mesh, Data Plane diagnostics triggered from the UI indicate that port 8123 is blocked from the HCX Manager to the IX appliance.
  • The below error message is observed
    HCX is unable to reach HCX-WAN-IX-EDGE on the ports 8123. Please ensure firewall is not blocking the ports or routing is correctly configured.

  • Replication Assisted vMotion (RAV) and Bulk migrations consistently fail to execute.
  • At the CLI level on the destination HCX manager, the /common/logs/app.log displays the following exception:
    com.vmware.vchs.hybridity.adapters.unified.hbr.server.fault.HbrConnectionException: Error connecting to HBR server <REDACTED_IP>:8123. Reason: HttpHostConnectException
  • While attempting a Bulk/RAV migration from the HCX manager, the following error message is displayed.

Error connecting to HBR server <REDACTED_IP>:8123. Reason: HttpHostConnectException

Environment

  • VMware HCX 9.1.0
  • VMware Cloud Foundation (VCF) 9.1

Cause

The issue happens due to a RACE condition where If hbrsrv happens to start (at boot, or whenever the hbrsrv service is (re)started) before hostd has applied that tagging / before the tagged vmknic has acquired its  IP, hbrsrv will load whatever hbrsrv-nic.xml contains at that instant.

Resolution

This issue is resolved in VCF 9.1.1, available at Broadcom downloads.

If you are having difficulty finding and downloading software, review the Download Broadcom products and software KB.

Verification Steps
Before applying the workaround, confirm the service status at the CLI of the HCX IX appliance

  1. SSH into the destination HCX Manager
    Run the command ccli
    go <IX-Appliance-ID-number>
    ssh
  2. On the IX appliance, run the following command to check if the HBR service is listening on the expected port:
    ip netns exec plr_sr bash
    netstat -plnt | grep -i hbrsrv
  3. While being on the IX appliance, run the command to check for active connections on port 8123: netstat -an | grep 8123
    Result: No output indicates the service is not listening on the expected port, thus confirming the issue.
  4. Furthermore, reviewing the configuration file in IX appliance located in /etc/vmware/hbrsrv-nic.xmlor /opt/vmware/etc/hbr/hbrsrv-nic.xml shows the ipForHMSbound to an incorrect IP rather than the expected Management IP of the IX appliance..

    root@<Obfuscated>-IXE-R1 :~ # cat /opt/vmware/etc/hbr/hbrsrv-nic.xml
    <config>
    <ipForHMS><Incorrect-IP></ipForHMS>
    <ipForNFC/>
    <ipForFilter>127.0.0.1</ipForFilter>
    </config>

Workaround Procedure 

Apply the following steps to both the source and destination Service Meshes:

  1. SSH into the HCX Manager.
  2. Enter the CCLI (HCX Command Line Interface).
  3. Execute list to view and identify the affected Inter-Connect Appliance (IX).
  4. Execute go <IX-ID> to select the affected IX appliance.
  5. Access the bash shell of the appliance by executing ssh.
  6. Sop the HBRSRV service by executing: systemctl stop hbrsrv
  7. Start the HBRSRV service by executing: systemctl start hbrsrv
  8. Repeat this procedure for the peer IX appliance as well.

After executing the above steps, ensure that on the IX appliance, the service hbrsrv-bin is listening on to the port 8123 and the file /opt/vmware/etc/hbr/hbrsrv-nic.xml or /etc/hbr/hbrsrv-nic.xml has the entry of the Management IP of the IX appliance for ipForHMS parameter.