/var/run/log/hostd.log] indicate a correlation with storage connectivity events. Lost access to volume" regarding the datastore hosting the Edge node. <timestamps> In(166) Hostd[2099671]: [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 66186 : Lost access to volume <Datastore-UUID> (<datastore-name>) due to connectivity issues. Recovery attempt is in progress and outcome will be reported shortly.
<timestamps> In(166) Hostd[2099625]: [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 66187 : Successfully restored access to volume <Datastore-UUID> (<datastore-name>) following connectivity issues.
<timestamps> In(166) Hostd[2099675]: [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 66190 : Lost access to volume <Datastore-UUID> (<datastore-name>) due to connectivity issues. Recovery attempt is in progress and outcome will be reported shortly.
<timestamps> In(166) Hostd[2099225]: [Originator@6876 sub=Vimsvc.ha-eventmgr] Event 66191 : Successfully restored access to volume <Datastore-UUID> (<datastore-name>) following connectivity issues.[/var/log/frr/frr.log] logs shows BGP flaps.less frr.log | grep "BFD status for peer"
<timestamps> BGP: BFD status for peer <BGP-neighbor-ip> changed from Down -> Up
<timestamps> BGP: BFD status for peer <BGP-neighbor-ip> changed from Up -> Down
<timestamps> BGP: BFD status for peer <BGP-neighbor-ip> changed from Down -> Up
<timestamps> BGP: BFD status for peer <BGP-neighbor-ip> changed from Up -> Down
[/var/log/syslog] logs shows BGP state changes.less syslog | grep "state=BGP"
<timestamps> edge-node NSX 9755 FABRIC [nsx@6876 comp="nsx-edge" subcomp="rcpm" s2comp="routing-service-realization" level="INFO"] Alarm for BGP <BGP-neighbor-ip>, peer_uuid: <####-uuid-####> in SR: <####-uuid-####>, state=BGP_DOWN
<timestamps> edge-node NSX 9755 FABRIC [nsx@6876 comp="nsx-edge" subcomp="rcpm" s2comp="routing-service-realization" level="INFO"] Alarm for BGP <BGP-neighbor-ip>, peer_uuid: <####-uuid-####> in SR: <####-uuid-####>, state=BGP_UP
<timestamps> edge-node NSX 9755 FABRIC [nsx@6876 comp="nsx-edge" subcomp="rcpm" s2comp="routing-service-realization" level="INFO"] Alarm for BGP <BGP-neighbor-ip>, peer_uuid: <####-uuid-####> in SR: <####-uuid-####>, state=BGP_DOWN
<timestamps> edge-node NSX 9755 FABRIC [nsx@6876 comp="nsx-edge" subcomp="rcpm" s2comp="routing-service-realization" level="INFO"] Alarm for BGP <BGP-neighbor-ip>, peer_uuid: <####-uuid-####> in SR: <####-uuid-####>, state=BGP_UPVMware NSX
VMware ESXi
The root cause is underlying storage instability or intermittent loss of connectivity to the datastore hosting the NSX Edge VM. When the ESXi host loses access to the datastore (the backing volume), I/O operations are paused or queued.
This storage deadlock prevents the guest OS of the Edge VM from processing time-sensitive BGP and BFD keepalive packets. Once the BGP hold timer expires, the BGP session resets, leading to the observed flapping behavior.
To resolve this issue, perform the following steps to verify storage stability and relocate the affected Edge node.
`/var/log/hostd.log` and `/var/log/vmkernel.log`) for storage access errors corresponding to the timestamps of the BGP flaps.`Lost access to volume <datastore_naa_id> (<datastore_name>)`NSX Manager UI under Networking > Tier-0 Gateways > BGP.