In a 15-node VCF Operations for Networks cluster, the following symptoms occur simultaneously:
No route to host).One or more essential services are not running or not healthy.TSDB Server failed to flush data to HBase or Launcher Service Not Running.Data Retention (Metric Store Maintenance) service is unhealthy on Platform Node 1.Healthy (Repartitioning) with several GBs of Moving Data.
VCF Operations for Networks
A node-level hang disrupts the cluster’s distributed data pipeline. Upon recovery via hard reset, FoundationDB (FDB) initiates a high-priority repartitioning process. This intensive activity causes secondary GUI synchronization issues and prevents automated maintenance cron jobs (like Metric Store Retention) from completing on schedule.
To restore service health immediately, follow these procedural steps:
./run_all.sh uptime from a healthy peer node../check-service-health.sh -p -d on all nodes once the affected node is online.fdbcli --exec "status detail". Ensure Moving Data reaches 0 GB before performing further reboots or maintenance.