Error: One or more essential services are not healthy or running in VCF Operations for Networks cluster
search cancel

Error: One or more essential services are not healthy or running in VCF Operations for Networks cluster

book

Article ID: 444697

calendar_today

Updated On:

Products

VCF Operations for Networks

Issue/Introduction

Symptoms

In a VCF Operations for Networks (formerly Aria Operations for Networks) clustered deployment, the following symptoms are observed:

  • The GUI reports "One or more essential services are not healthy/running".

  • Platform nodes report "Grid processing stopped since kafka cluster is not available".

  • Service status indicates "Launcher Service Not Running" or "No activity detected".

  • Java-based services (FlinkContainer, ElasticSearch, Kafka, HRegionServer) fail to initialize.

Environment

  • VCF Operations for Networks 6.x

  • Clustered Deployments (3-node or higher)

Cause

  • Ungraceful node reboots or power outages can lead to service dependency failures and database locking (Kafka/HBase/Zookeeper), preventing services from starting automatically.

Resolution

Perform a sequential cold shutdown and restart of the entire cluster to clear stale locks and initialize services in the correct order.

  1. Ensure access to vCenter with permissions for Power and Snapshot operations.

  2. Follow the scripted shutdown procedure in KB 314428 - Best practices to shutdown VCF Operations for Networks Clustered deployments

  3. Shutdown all Collector nodes first via vCenter (Power > Shut Down Guest OS).

  4. Log into Platform Node 1 via SSH as support and switch to ubuntu user:
    ub
  5. Execute the cluster shutdown script per the KB 314428 - Best practices to shutdown VCF Operations for Networks Clustered deployments
    ./vrni-cluster-shutdown-script.sh shutdown 127.0.0.1 "/home/ubuntu/shutdown_$(date +%s).log"
  6. Once all nodes are off, take snapshots of all Platform and Collector nodes.

  7. Power on Platform nodes sequentially (Node 1, then Node 2, then Node 3).

  8. On Platform Node 1, start the services after starting a new SSH session, logging in with support account and switching to the ub user:  ./vrni-cluster-shutdown-script.sh start-services 127.0.0.1 "/home/ubuntu/start_$(date +%s).log"

  9. Wait 15 minutes, then power on Collector nodes one by one.

  10. Verify service health in the GUI under Settings > Infrastructure and Support.  NOTE:  It may take up to 20 minutes for Data Collection errors to self-resolve and those errors to disappear.

  11. Delete all snapshots for each of the Platform and Collector nodes.  This should be done while the nodes are running.  It can be done one at a time to minimize performance impact on the running VMs. 

 

To speak with a customer representative or a Support Engineer see Contact Support. Scroll to the bottom of the page and click on your respective region.