Resolving Persistent VCF Operations Alert: " One or more VMware Cloud Foundation Operations services on a node are down"
search cancel

Resolving Persistent VCF Operations Alert: " One or more VMware Cloud Foundation Operations services on a node are down"

book

Article ID: 448816

calendar_today

Updated On:

Products

VCF Operations

Issue/Introduction

In a VMware Cloud Foundation (VCF) Operations 9.1.x environment, persistent alerts occur stating: "One or more VMware Cloud Foundation Operations services on a node are down."

    • The alert triggers on multiple analytics nodes simultaneously.
    • Standard cluster restart procedures (offline/online) do not clear the alert.
    • The alert persists even when services appear healthy in the UI.
    • In specific scenarios, the alert triggers based on an OR condition between multiple symptoms.
    • The primary cause is a "collector is down" event symptom that remains active in the PostgreSQL database.

Environment

VCF Operations 9.1.x

Cause

The symptom does not automatically cancel after 45 days, causing the alert to remain active indefinitely.

Resolution

To clear the persistent alert, briefly place the affected collector objects into maintenance mode to reset their state.

  1. Log in to the VCF Operations primary node UI.
  2. Navigate to Environment > Object Browser.
  3. For each affected analytics node, select the corresponding object: VCF-Operations Collector-####.
  4. Place the object into Maintenance Mode.
  5. Wait for 1 minute.
  6. Take the object out of Maintenance Mode.

The alert should stop triggering immediately after these steps are completed.

Additional Information

If the issue persists, ensure the VCF Instance Cloud Proxy registration is configured correctly. Refer to KB 443106 for troubleshooting VCF Instance Cloud Proxy registration validation failures.

If you require further assistance, please contact Broadcom Support: Contact Broadcom Support.