A model deployed on PAIS responds to the first few completions normally, then becomes nearly non-responsive, with requests eventually getting aborted
[engine.py:285] Aborted request chatcmpl-..
This is commonly caused by the GPU node failing to acquire a license from the NVIDIA vGPU License Server - the license pool is exhausted, or no licenses are available. Without a license, the vGPU falls back to a restricted/degraded mode.
Prerequisites:
kubectl get secret $(kubectl get secret | grep kubeconfig | awk '{print $1}') -o jsonpath='{.data.value}' | base64 -d > vks-kubeconfig.ymlexport KUBECONFIG=./vks-kubeconfig.ymlkubectl -n gpu-operator get pods -o widenvidia-smi: kubectl -n gpu-operator exec -it -c nvidia-device-plugin -- nvidia-smi -q | grep -i licensevGPU Software Licensed Product
License Status : Licensed (Expiry: <Date>)vGPU Software Licensed Product License Status : Unlicensed (Restricted)Obtaining, generating, or validating NVIDIA NGC API keys and license tokens is managed entirely within the customer's NVIDIA NGC portal account and is outside the scope of Broadcom Support. Please work with your organization's NVIDIA account administrator to acquire valid credentials.