Deployment workflow for VCF Management Services runs indefinitely, and never fails, on VCF 9.1.0
book
Article ID: 452260
calendar_today
Updated On:
Products
VMware SDDC Manager / VCF Installer
Issue/Introduction
During a VCF 9.0.2 to 9.1.0 upgrade, post upgrade of the VCF Operations and SDDC Manager components to 9.1.0, the VCF Management Services workflow is started.
It has been confirmed that no uppercase letters are used in the FQDNs entered for the workflow, per this article.
Error logging in SDDC Manager /var/log/vmware/vcf/domainmanager/domainmanager.log matching that in this article has been identified.
VMSP agent is not healthy: failed to connect to health endpoint: Get "https://<Bootstrap IP>:5480/health": dial tcp <Bootstrap IP>:5480: connect: no route to host.
The VMSP bootstrap VM has been inadvertently deleted, and is not recoverable.
On the VCF Operations workflow page SDDC Manager has been successfully upgraded - on the Next steps: section, 3. Install Components shows a spinning icon
Environment
VCF 9.1.0
Cause
Certain ESXi hosts in the environment have underlying network isolation or routing issues.
When the VCF Management Services bootstrap VM is scheduled on these problematic ESXi hosts by DRS, they become unrouteable.
This prevents the bootstrap VM from connecting to the API server for the VCF Management Services cluster.
The bootstrap VM then having been inadvertently deleted potentially lead to the deployment failing to timeout as expected.
Resolution
Check for, and resolve, underlying network isolation or routing issues on the ESX hosts using this article.