Pod failures via Network Service Mesh (NSM) and VPP Forwarder Context Timeouts
search cancel

Pod failures via Network Service Mesh (NSM) and VPP Forwarder Context Timeouts

book

Article ID: 453321

calendar_today

Updated On:

Products

VMware Telco Cloud Automation VMware Telco Cloud Platform

Issue/Introduction

  • Huge pod failures occur within a short window across application workload namespaces.
  • The underlying Cloud Infrastructure (vCenter / ESXi / TCA) shows no active alarms, hardware events, or host node issues.
  • Restarting the Network Service Mesh (NSM) VPP forwarder pods temporarily resolves the connectivity issue.

    Error Message / Log Pattern Application and NSM proxy pods log gRPC connection failures similar to:

    rpc error: code = Unknown desc = 0. An error during select forwarder pod-forwarder-vpp-jc9r6 --> 
    failed to dial unix:///proc/37/fd/2877: context deadline exceeded: 
    all forwarders have failed: cannot support any of the requested mechanism

Environment

TCA 3.3
TCP 5.0

Cause

Communication timeouts within the Network Service Mesh (NSM) layer. This is typically isolated to the application network layer rather than the underlying physical or virtual infrastructure.

Resolution

Follow these steps to identify and resolve NSM communication timeouts:

  1. Analyze TCA and node support bundles to confirm the incident window.
  2. Verify vCenter health metrics to rule out virtualization platform failures.

    •  If vCenter is healthy, focus troubleshooting on the NSM/Application network layer.

  3. Examine service mesh component logs for specific communication timeout errors during the event window.
  4. Capture network communication patterns and traffic flows during the failure window.
  5. Engage the internal application team to analyze the captured traffic and identify triggers for NSM timeouts (e.g., traffic volume spikes or configuration issues).
  6. Adjust NSM configuration or traffic management policies based on the application team's findings.