Remediation of host fails with health check failure - WCP Service Health Alarm
search cancel

Remediation of host fails with health check failure - WCP Service Health Alarm

book

Article ID: 404575

calendar_today

Updated On:

Products

VMware vCenter Server Tanzu Kubernetes Runtime VMware vSphere Kubernetes Service

Issue/Introduction

  • Below error message is seen on vCenter Server "Updates" page when performing remediation: 

    Remediation of cluster failed
    Health Check for '<cluster name>' failed
    Health Check fails to retrieve data about service 'Workload Management Platform' on '<cluster id>'. Verify that the service 'Workload Management Platform' is running and try again. 

  • "WCP Service Health Alarm" is observed on vCenter Server web page.

  • The Workload Management (vCenter 8) or Supervisor Management (vCenter 9) page is unavailable showing that it encountered and error and the service is unable to be reached.

  • In VMware vCenter Server Appliance Management (VAMI) at port 5480, the service "Workload Control Plane" shows as Degraded with "failed to retrieve service health"

  • Below error messages can be found in /var/log/vmware/wcp/wcpsvc.log on vCenter Server:

    Failed to read VCSA proxy settings from file '/etc/sysconfig/proxy' with error: open /etc/sysconfig/proxy: too many open file
    Unable to get the list of clusters: Post "http://localhost:1080/sdk": dial tcp: lookup localhost: device or resource busy
    Unable to collect information for cluster agencies: Post "http://localhost:1080/sdk": dial tcp: lookup localhost: device or resource busy

Environment

vCenter Server

vSphere Supervisor

Cause

The wcp service on vCenter Server is not in a healthy state.

Resolution

Workaround

To workaround the issue, restart wcp service on vCenter Server with below command:

service-control --restart wcp

Alternatively, the service can also be restarted in vCenter Server Appliance Management Interface(VAMI) by selecting "Workload Control Place" service and click "RESTART".

 

Further checks should be made in the environment to understand the cause of the wcp service failure.

If there is a Supervisor cluster in the environment, its status and underlying system components should be checked for any issues which may be contributing to the wcp service's failure.

Additional Information

Related KB: vSphere Kubernetes Supervisor Cluster Error - Unable to connect to the management DNS servers from control plane VM - The connection was attempted over the workload network