NSX Edge BGP Sessions Down During Stretched vSAN Cluster Connectivity Loss
search cancel

NSX Edge BGP Sessions Down During Stretched vSAN Cluster Connectivity Loss

book

Article ID: 452153

calendar_today

Updated On:

Products

VMware NSX

Issue/Introduction

  • NSX Edge Nodes experience BGP and BFD session drops, leading to north/south routing failures and loss of connectivity for virtual machines on NSX segments.

  • This is observed during an underlying physical network maintenance, such as a core switch reboot. 

  • NSX Edge logs indicate the following status transition during the event: 

    2026-07-21T11:50:23.117Z [NSX Edge hostname] NSX 1 FABRIC [nsx@6876 comp="nsx-edge" subcomp="nsxa" s2comp="ha-cluster" level="WARN"]Self Node [UUID] status changed from Up (Routing Down) to Down (VTEP tunnels down)

  • Evidence of a storage liveness loss is seen in the ESXi hostd logs:

    2026-07-21T11:49:07.262Z In(182) vmkernel: cpu57:2098860)DOM: DOMOwner_SetLivenessState:11608: Object [UUID] lost liveness [0x45d####ee2c0]

Environment

  • VMware NSX

  • Configuration: Stretched vSAN cluster environments with dependency on physical network topology for storage connectivity.

Cause

  • The issue is caused by a loss of liveness for stretched vSAN datastore objects during physical network maintenance.
    • When the vSAN VLAN and BGP routes are not fully advertised across all core switches in the fabric, rebooting one switch causes an interruption in the vSAN datastore access.
    • This storage disruption propagates to the NSX Edge VM (a guest on the ESXi host), causing it to become unresponsive or hang. Consequently, the datapathd service fails, leading to the termination of all TEP tunnels and the resulting BGP adjacency drops.

Resolution

To resolve this issue, the physical network configuration must ensure consistent and redundant advertisement of the vSAN VLAN and BGP routes across all physical core switches.

  1. Identify the physical core switches involved in the stretched vSAN fabric.

  2. Verify that the vSAN VLAN is trunked and active on all core switches.

  3. Ensure that BGP routing is correctly configured on both core switches in both datacenters to handle redundant traffic.

  4. Confirm that BGP peering is established on all physical switches in the path.

  5. Perform a maintenance reboot of one core switch to verify that storage liveness is maintained via the redundant path.

Additional Information

Multiple objects are in an inaccessible state due to network issues on vSAN cluster