Supervisor upgrade to VCF 9.1 is stalled with "Solution(s) apply failed on host"
search cancel

Supervisor upgrade to VCF 9.1 is stalled with "Solution(s) apply failed on host"

book

Article ID: 443649

calendar_today

Updated On:

Products

VMware vSphere Kubernetes Service

Issue/Introduction

  • The Supervisor upgrade in a VMware Cloud Foundation 9.x environment stalls at the host configuration stage.
  • ESXi hosts may remain in an endless "Updating" loop or stay in Maintenance Mode.
  • The vSphere Client reports the following repetitive task:
    Apply Solution,wcp-###########@vsphere.local, A general system error occurred: Solution(s) apply failed on host:##########
  • WCP logs on the vCenter under /var/log/vmware/wcp/wcpsvc.log:
    YYYY-MM-DDTHH:MM:SS ERROR wcp ###### [vc@####] [pman/client.go:726] [opID=vLCM:Upgrade:domain-c####] PMan API: Task Get for Apply Task API: Attempt#[4]:Failed Attempts:[0] of MaxFailedAttempts[1]Error Message: Solution(s) apply failed on host: '###########'
    YYYY-MM-DDTHH:MM:SS ERROR wcp 2710585 [vc@####] [kubelifecycle/pman_client.go:645] [opID=vLCM:Upgrade:domain-c####] PMan API: ApplyUpgradeTask INTERIM FAILURE Error: PMan API: Task Get for Apply Task API: Attempt#[4]:Failed Attempts:[0] of MaxFailedAttempts[1]Error Type: ERROR
    YYYY-MM-DDTHH:MM:SS ERROR wcp 2710585 [vc@####] [pman/client.go:682] [opID=vLCM:Upgrade:domain-c####] PMan API: Apply Task: Attempt#[1 of 1]: Has Failed - Task Get Polling failed!
  • Hostd logs on the ESXi host under /var/run/log/hostd.log:
    YYYY-MM-DDTHH:MM:SS -INFO Hostd ###### [esx@#### sub="Vimsvc.ha-eventmgr"] Event #####: Could not install image profile: VMware_bootbank_spherelet_9.0.1.32.5.0-25065159: VMware_bootbank_spherelet_9.0.1.32.5.0-25065159: Failed to unmount tardisk spherele.v00 of VIB VMware_bootbank_spherelet_9.0.1.32.5.0-25065159: Error in running [/bin/rm /tardisks/spherele.v00]:
    YYYY-MM-DDTHH:MM:SS -INFO Hostd ###### [esx@####] --> Return code: 1
    YYYY-MM-DDTHH:MM:SS -INFO Hostd ###### [esx@####] --> Output: rm: can't remove '/tardisks/spherele.v00': Device or resource busy
  • Hostd.log stuck in loop
    YYYY-MM-DDTHH:MM:SS -INFO Hostd ###### [esx@#### sub="Libs"] FCDLIB: fcd-catalog: Prophylactica exiting - Host in maintenance mode
  • Sphere service status on ESXI host
    #/etc/init.d/spherelet status
    YYYY-MM-DDTHH:MM:SS -INFO init.d/spherelet ###### [esx@####] spherelet is running
    YYYY-MM-DDTHH:MM:SS -INFO init.d/spherelet ###### [esx@####] WCP Mgmt Proxy is not running

Environment

VMware Cloud Foundation 9.x
VMware vCenter Server
VMware vSphere ESXi
VMware Supervisor
VMware vSphere Kubernetes Service

Cause

An incorrect upgrade sequence causes this issue when ESXi hosts upgrade to version 9.1 before the Supervisor upgrade initiates or completes. This sequence results in a version mismatch and an active lock on the Spherelet tardisks, preventing the Supervisor lifecycle manager from updating the host configuration.

Additionally, when the asynchronous ClusterApplySolutionTask fails during the Supervisor upgrade sequence, vLCM's (vSphere Lifecycle Manager)exception handling does not trigger an ExitMaintenanceMode command. On subsequent retries, vLCM treats the apply attempt as a new, independent workflow. Because the host is already in Maintenance Mode, but the internal operationAttempted flag was never successfully set during the initial failure, vLCM skips taking the host out of Maintenance Mode.

Review the VCF 9.0 Update Sequence here: Update Sequence for VCF 9.0 and Compatible VMware Products

Resolution

To resolve this livelock condition, perform the following manual steps on the affected ESXi host:

  1. Establish an SSH session to the affected ESXi host.
  2. Stop the spherelet service manually to release the file lock by running the following command: 
    /etc/init.d/spherelet stop
  3. Monitor the vSphere Client to see if the Supervisor host configuration task auto-completes shortly after the service stops.
  4. If the task does not auto-complete and the host remains stuck in Maintenance Mode, verify that the spherelet VIB version has been successfully upgraded to 9.1 by running: 
    esxcli software vib list | grep -i spherelet
  5. After confirming the VIB version is 9.1, manually exit Maintenance Mode on the ESXi host via the vSphere Client or command line.
  6. Allow the patching operation to resume for the next node in the cluster and repeat these steps for any additional hosts in the cluster experiencing the same failure.

Additional Information

Exiting an ESXi Host Stuck in Maintenance Mode