Management and Workload clusters spontaneously and intermittently enter a disconnected state in the TCA CaaS UI.
Simultaneously, the connected endpoint UI displays the certificates for these clusters as untrusted.
The clusters typically recover automatically and return to a normal status after a period of time.
Additionally, the TCA-CP 9443 appliance management interface becomes intermittently inaccessible, and the tca_cert_obs pod is observed in a constant failed state.
Terminal connections to cluster, from within the TCA UI , continue to function properly even when clusters show as disconnected.
Environment
TCA 3.4.0
TCP 5.1
Cause
The root cause is currently unknown. The issue is intermittent and cannot be reliably reproduced on demand.
Resolution
As the issue is not readily reproducible, further diagnostic data is required for analysis.
Perform the following data collection steps:
Download and run the diagnostic script named pg-diagnose.sh attached to this KB article on the affected TCA Control Plane nodes while the issue is actively occurring. Save the output of the script in a file and attach it with the support case
Collect complete support log bundles for all involved TCA Manager and TCA-CP nodes. When generating these bundles ensure that the DB dump and all core logs are selected.
Capture full screenshots/snapshots from the TCA UI clearly showing the disconnected cluster status and the untrusted certificate state while the issue is actively occurring.
Attach all collected logs, script outputs, and screenshots to your active support case for review.