Intermittent ICMP Packet Loss and LCM Table Update Failures in NSX 4.1.1.0
search cancel

Intermittent ICMP Packet Loss and LCM Table Update Failures in NSX 4.1.1.0

book

Article ID: 452903

calendar_today

Updated On:

Products

VMware NSX

Issue/Introduction

  • Intermittent ICMP packet loss and production datapath alarms on NSX-T Edge nodes.
  • Asymmetric traffic flow observed across host transport nodes and Edges. 
  • Packet captures confirm ICMP Echo Requests are forwarded, but specific sequence numbers drop due to unexpected Tunnel Endpoint (TEP) routing.

Environment

VMware NSX 4.1.1.0

Cause

Prolonged system uptime (>300 days) leads to stale transient memory states, causing microservice desynchronization within the NSX Manager control plane. During LCM table updates, this desynchronization causes host traffic to route to an incorrect Destination TEP IP address. The resulting asymmetric routing path triggers packet drops and active datapath alarms.

Resolution

   Perform a Sequential Rolling Reboot

    1. Initiate a graceful reboot of the first NSX Manager node.
    2. Wait for the node to return to a fully synchronized and Healthy cluster state before proceeding to the next node.
    3. Sequentially reboot all remaining NSX Manager nodes in the cluster to clear stale memory states and re-initialize control plane microservices.