When attempting an upgrade of the NSX Management Cluster, the upgrade fails on one or more nodes. The NSX Manager User Interface or upgrade logs display the following error message:Unexpected error while upgrading upgrade unit: [MPP] Node upgrade failed : org.springframework.web.client.HttpServerErrorException$InternalServerError: 500 Internal Server Error: "{"error_code": 36580, "error_message": "Error proxying request to: ####-####-####-####.", "module_name": "node-services"}".
VMware NSX
The issue is caused by physical network infrastructure degradation on the ESXi host hosting the failing NSX Manager appliance.
Because the NSX Management Cluster requires an inter-node network latency of under 10ms to maintain distributed cluster state sync, the combination of storage and network drops causes the node-services module to time out, drop its database connection, and fail the upgrade unit transaction.
For detailed documentation on maximum clustering latencies and architecture limits, see the VMware Cloud Foundation Technical Documentation.