"Add Platform Node(s)" Scale-Out Fails with Auth Timeout/502 for VCF Operations for Networks
search cancel

"Add Platform Node(s)" Scale-Out Fails with Auth Timeout/502 for VCF Operations for Networks

book

Article ID: 454016

calendar_today

Updated On:

Products

VCF Operations for Networks

Issue/Introduction

  • VCF Operations GUI shows error as seen in below screenshot:



  • Unable to add platform nodes  to scale out is failing with below error seen on GUI

    An unexpected error occurred in step expand_vrni_platform_cluster_task_ref. Reference Code: <hex>. Detail: I/O error on POST request for ".../api/auth/login": Read timed out Reference Code: <hex>. Detail: 502 Bad Gateway on POST request for ".../api/auth/login"

  • Adding platform nodes to VCF Operations for Networks (vRNI) via scale-out fails at the "Scale out component" subtask, specifically at step expand_vrni_platform_cluster_task_ref

  • This failure is typically seen shortly after a scale-up of the vRNI platform node, or if the system takes an extended amount of time to boot up. The task reports as "Failed" even though the vRNI cluster may later show all nodes as deployed and "Running" in the inventory.

Environment

  • VCF Operations for Networks 9.1.1

Cause

VCF Operations for Network services take longer to become fully ready after a scale-up or certificate rotation than the current LCM retry logic tolerates.
Fleet Lifecycle executes short internal retries instead of waiting for a longer duration before marking the operation as succeeded
.

Resolution

To resolve this perform below:

  1. Wait to ensure the VCF Operations for Networks (vRNI) platform has had sufficient time to fully boot up and all services are ready.
  2. Manually retry the failed scale-out task. A manual retry typically succeeds once the services have stabilized.