Symptoms:
During an NSX Global Manager upgrade (typically to version 4.2.x), the upgrade fails on the first node of a standby cluster.
The NSX Manager UI may show a 503 service unavailable or a spinning wheel.
The HTTP service on the affected node is down.
Running get upgrade progress-status from the admin CLI returns:
% Error executing upgrade step 'null'
The upgrade-coordinator.log contains errors similar to:
com.vmware.nsx.management.upgrade.rpcframework.UcRestRpcException: org.springframework.web.client.HttpServerErrorException$InternalServerError: 500 Internal Server Error: [no body]
Upgrade pre-checks fail with repository certificate errors and version sync status mismatches.
VMware NSX
NSX Federation
The upgrade failure occurs due to a corrupted upgrade plan metadata state, often triggered by service restarts or communication failures during the Management Plane (MP) upgrade phase. This leaves the node in a partial upgrade state where binaries are updated but the database and service configurations are mismatched.
This issue is currently under investigation.
Workaround steps: