PVC or VolumeSnapshot Operations Fail on Supervisor Cluster Due to Expired Storage Quota Certificates
search cancel

PVC or VolumeSnapshot Operations Fail on Supervisor Cluster Due to Expired Storage Quota Certificates

book

Article ID: 424055

calendar_today

Updated On:

Products

VMware vCenter Server VMware vSphere Kubernetes Service

Issue/Introduction

On a Supervisor cluster or a Guest cluster, creating or updating PVC or VolumeSnapshot objects may occasionally fail during storage quota validation.

You may observe an error message similar to one of the following error messages in your environment:

  • failed to create volume : admission webhook "validate-quota-on-create.k8s.io" denied the request: Operation denied, Post "https://cns-vsphere-vmware-com-service.kube-system.svc.cluster.local:443/getrequestedcapacityforpersistentvolumeclaim": tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 20##-08-27T16:58:17Z is after 20##-08-23T02:15:00Z

  • failed to create volume : admission webhook "validate-quota-on-create.k8s.io" denied the request: Operation denied, Post "https://cns-vsphere-vmware-com-service.kube-system.svc.cluster.local:443/getrequestedcapacityforpersistentvolumeclaim": remote error: tls: expired certificate

PVC Status and Event Logs: Running kubectl describe pvc <example_pvc> -n <example_namespace> reveals the PVC stuck in Pending status with the following event:

Error: admission webhook "validate-quota-on-create.k8s.io" denied the request: Operation denied, Post "https://cns-vsphere-vmware-com-service.kube-system.svc.cluster.local:443/getrequestedcapacityforpersistentvolumeclaim": tls: failed to verify certificate: x509: certificate signed by unknown authority (possibly because of "x509: ECDSA verification failure" while trying to verify candidate authority certificate "storage-quota-selfsigned-issuer-cert")

Environment

  • vCenter Server: 9.0, 9.0.1, 9.0.2
  • VMware vSphere Kubernetes Service (VKS)

Cause

During storage quota validation for CREATE/UPDATE requests received for PVC/VolumeSnapshots on a Supervisor cluster, the storage quota webhook communicates with the CNS extension service via an mTLS connection. This connection relies on client-server certificates managed by cert-manager associated with the respective pods.

When CA certificates are auto-renewed by cert-manager upon their expiry, the respective child certificates signed using the old CA certificate are not automatically refreshed. Furthermore, the new certificate data is not dynamically reloaded internally into the storage quota webhook and CNS extension service pods. Consequently, the services continue to utilize the stale, expired certificate data, resulting in TLS handshake failures and connection loss during the storage quota validation process.

Resolution

To work around this issue, restart the storage quota webhook and CNS extension service pods to force them to reload the new certificate data.

Note: These steps require root access to the vCenter Server and the Supervisor control plane. By design, a standard Supervisor Administrator cannot modify the kube-system namespace. You must have local system-level privileges to perform these platform-level maintenance tasks.

  1. Log in to the vCenter Server appliance via SSH as the root user: ssh root@##.##.##.##

  2. Retrieve the credentials for the Supervisor control plane: /usr/lib/vmware-wcp/decryptK8Pwd.py

  3. Log in to the Supervisor control plane via SSH using the IP address and credentials obtained in the previous step: ssh root@##.##.##.##

  4. Scale down the storage-quota-webhook pods to zero: kubectl -n kube-system scale deploy storage-quota-webhook --replicas=0

  5. Verify the storage-quota-webhook scales down successfully (The READY column shows 0/0): kubectl -n kube-system get deploy storage-quota-webhook

  6. Scale the storage-quota-webhook pods back up (the default replica count is 3): kubectl -n kube-system scale deploy storage-quota-webhook --replicas=3

  7. Verify the storage-quota-webhook scales up successfully (The READY column shows 3/3): kubectl -n kube-system rollout status deploy/storage-quota-webhook

  8. Scale down the cns-storage-quota-extension pods to zero: kubectl -n kube-system scale deploy cns-storage-quota-extension --replicas=0

  9. Verify the cns-storage-quota-extension scales down successfully (The READY column shows 0/0): kubectl -n kube-system get deploy cns-storage-quota-extension

  10. Scale the cns-storage-quota-extension pods back up (the default replica count is 1): kubectl -n kube-system scale deploy cns-storage-quota-extension --replicas=1

  11. Verify the cns-storage-quota-extension scales up successfully (The READY column shows 1/1): kubectl -n kube-system rollout status deploy/cns-storage-quota-extension