Day 5 — BGP Finite-State Machine: States, Transitions and Session Troubleshooting

Learning objective

Use the BGP finite-state machine as a troubleshooting model: identify what a current state proves, what it does not prove, and which layer of evidence should be collected next.


BGP Finite-State Machine



1. Opening — a BGP state is evidence, not a diagnosis

BGP state = Active is one of the most misunderstood lines in routing operations.

The word *Active* sounds healthy. In BGP it is not the desired steady state. The desired operational state is Established. Active is part of the connection-establishment process and often appears when the speaker is repeatedly trying to create the underlying TCP session.

Likewise, Idle does not automatically mean “interface down,” and OpenSent does not mean “routes are being sent.” Each state represents a particular stage in the BGP finite-state machine (FSM). The state narrows the investigation, but the state alone is not the root cause.

The best BGP troubleshooters therefore ask two questions:

  1. What has already succeeded if the session reached this state?
  2. What is the next protocol event that must succeed before the state can advance?

That turns the FSM into a diagnostic map instead of a list to memorize.


2. Why BGP has a finite-state machine

BGP session establishment includes several independent mechanisms:

  • local configuration and administrative enablement;
  • TCP connection establishment;
  • BGP OPEN exchange;
  • version/ASN/hold-time/identifier/optional-parameter processing;
  • capability negotiation;
  • KEEPALIVE confirmation;
  • and, finally, ongoing UPDATE exchange.

An FSM gives the implementation a deterministic way to react to events such as:

  • start or stop;
  • successful or failed TCP connection;
  • TCP connection loss;
  • OPEN received;
  • KEEPALIVE received;
  • NOTIFICATION received;
  • timer expiry;
  • and protocol errors.

RFC 4271 defines the core BGP FSM. RFC 9687 later defines the SendHoldTimer and an associated FSM event to address cases where a BGP speaker is unable to send messages for too long even though the remote side has not closed the connection. That extension is useful standards context, but platform implementation/support must be verified before assuming a specific Cisco release exposes or implements it in a particular way.


3. The six core states

3.1 Idle

Idle is the starting/reset state of the BGP FSM.

Conceptually, in Idle the speaker is not in a usable BGP relationship with the peer. Depending on events and configuration, the implementation initializes resources and waits for a start event before attempting transport establishment.

Operational interpretation:

  • The session is not established.
  • Do not assume the cause is physical connectivity.
  • Check whether the neighbor is administratively configured/activated, whether the process is running, whether a reset just occurred, and what the last reset reason says.

3.2 Connect

In Connect, BGP is waiting for the TCP connection to complete.

What reaching Connect suggests:

  • the BGP process is attempting session establishment;
  • the transport connection is the immediate dependency.

What it does not prove:

  • that IP reachability is bidirectional;
  • that TCP/179 is permitted;
  • that the peer address is correct;
  • that the remote ASN matches;
  • or that BGP OPEN negotiation will succeed.

3.3 Active

In Active, the speaker is also working on transport establishment, generally after a connection attempt has failed or while retry behavior continues.

The crucial operational rule:

> Active does not mean Established and it does not mean routes are being exchanged.

If a session repeatedly appears in Connect/Active, focus first on transport and peer reachability rather than route policy.

Typical areas to check:

  • routing to the peer address;
  • return routing;
  • source address/update-source;
  • TCP/179 ACL/firewall handling;
  • multihop/TTL design;
  • peer availability;
  • duplicate or incorrect neighbor configuration.

3.4 OpenSent

OpenSent means the TCP transport has progressed far enough for BGP OPEN processing to be underway.

This is a major troubleshooting boundary. If you reliably reach OpenSent, the problem is no longer simply “TCP cannot connect.” Now inspect BGP session parameters and OPEN validation:

  • BGP version compatibility;
  • configured/received ASN;
  • BGP Identifier validity;
  • hold-time negotiation constraints;
  • optional parameters;
  • capability requirements;
  • and NOTIFICATION information.

An ASN mismatch frequently manifests around OPEN processing, although the exact visible state may be brief because the peer can immediately send a NOTIFICATION and reset.

3.5 OpenConfirm

OpenConfirm means OPEN negotiation has succeeded sufficiently for the speaker to wait for confirmation through KEEPALIVE processing.

At this point:

  • TCP has succeeded;
  • OPEN has been accepted;
  • the session is close to Established.

If a session repeatedly reaches OpenConfirm and fails, investigate KEEPALIVE/timer behavior, transport stability, protocol errors and any reset reason rather than basic neighbor IP configuration alone.

3.6 Established

Established is the operational state in which BGP peers can exchange UPDATE, KEEPALIVE and other negotiated/defined BGP messages as applicable.

But Established proves only that the session is operational.

It does not prove:

  • required prefixes were received;
  • received prefixes passed import policy;
  • the expected path became best;
  • the route entered the RIB;
  • the route entered the FIB;
  • or end-to-end forwarding works.

This distinction will appear throughout the entire series.


4. Troubleshooting map by state

| State | What to investigate first | Common evidence |
|---|---|---|
| Idle | configuration/start/reset cause | running config, logs, neighbor detail |
| Connect | TCP establishment | routing, ping, TCP capture, ACLs |
| Active | failed/retrying transport | route to peer, source address, TCP/179, TTL/multihop |
| OpenSent | OPEN parameters | remote AS, BGP Identifier, NOTIFICATION, capabilities |
| OpenConfirm | KEEPALIVE/transport stability | message counters, timers, resets |
| Established | route exchange and forwarding | received/accepted/best routes, RIB/FIB, traffic tests |

This table is a starting point, not a substitute for evidence. A fast state transition may never be visible in a periodic CLI poll. Logs or packet capture can reveal events that show ip bgp summary misses.


5. Scenario — RETICUX Lab FND-005

Two directly connected eBGP routers form a stable session. We then apply an interface ACL that blocks BGP TCP traffic between them while leaving ICMP permitted.

Topology


BGP Finite-State Machine: States, Transitions and Session Troubleshooting
BGP Finite-State Machine: States, Transitions and Session Troubleshooting


+----------------+      203.0.113.0/30      +----------------+
| EDGE-A AS64512 |---------------------------| EDGE-B AS64513 |
| 203.0.113.1    |          eBGP             | 203.0.113.2    |
+----------------+                           +----------------+

Success criteria before fault

  • Ping works both ways.
  • BGP is Established.
  • Neighbor detail shows normal message activity.

Fault objective

Demonstrate that IP reachability can remain healthy while TCP/BGP establishment fails, causing the FSM to remain outside Established.


6. Baseline configuration

EDGE-A

hostname EDGE-A
!
interface GigabitEthernet0/0
 ip address 203.0.113.1 255.255.255.252
 no shutdown
!
router bgp 64512
 bgp log-neighbor-changes
 neighbor 203.0.113.2 remote-as 64513
 address-family ipv4
  neighbor 203.0.113.2 activate
 exit-address-family

EDGE-B

hostname EDGE-B
!
interface GigabitEthernet0/0
 ip address 203.0.113.2 255.255.255.252
 no shutdown
!
router bgp 64513
 bgp log-neighbor-changes
 neighbor 203.0.113.1 remote-as 64512
 address-family ipv4
  neighbor 203.0.113.1 activate
 exit-address-family

7. Verification before modification

On both routers:

show ip bgp summary
show ip bgp neighbors

Expected relevant evidence:

  • state Established;
  • remote AS correct;
  • connection uptime increasing;
  • KEEPALIVE/OPEN message counters present;
  • no recent unexpected reset.

Verify IP connectivity:

ping 203.0.113.2

from EDGE-A and the reverse from EDGE-B.


8. Controlled modification — block only BGP TCP on EDGE-A inbound

Apply this ACL on EDGE-A:

configure terminal
ip access-list extended BLOCK-BGP-FSM-LAB
 deny tcp host 203.0.113.2 host 203.0.113.1 eq bgp
 deny tcp host 203.0.113.2 eq bgp host 203.0.113.1
 permit ip any any
!
interface GigabitEthernet0/0
 ip access-group BLOCK-BGP-FSM-LAB in
end

Then allow the existing session to fail naturally or, in a controlled lab only, perform a BGP reset to force immediate re-establishment testing.

Do not casually clear production BGP sessions merely to make a lab symptom appear faster.

Expected control-plane result

The session cannot complete/re-establish the TCP connection. The exact instantaneous FSM state can vary as the implementation retries, but the session should not remain Established while the ACL blocks the required TCP exchange.

Expected data-plane observation

ICMP can still pass because the ACL permits other IP traffic. This creates the important troubleshooting clue: ping succeeds but BGP transport fails.


9. Fault injection

Illustrative lab

BGP Finite-State Machine: States, Transitions and Session Troubleshooting
BGP Finite-State Machine: States, Transitions and Session Troubleshooting


Fault

An inbound ACL on EDGE-A blocks TCP traffic associated with BGP between the two peer addresses while permitting other IP traffic.

Symptoms

  • Ping succeeds.
  • BGP is not Established.
  • FSM may cycle through connection-establishment states.
  • Route exchange stops.
  • Logs can show neighbor down/reset events.

Why this fault is useful

It proves that “I can ping the neighbor” is necessary evidence for IP reachability but is not sufficient evidence for BGP transport.


10. Evidence-based troubleshooting

Step 1 — scope

show ip bgp summary

Confirm that the problem is the session, not merely a missing prefix.

Step 2 — confirm Layer 3

ping 203.0.113.2
show ip route 203.0.113.2

If ping succeeds and the route is connected, do not stop troubleshooting.

Step 3 — inspect FSM and reset reason

show ip bgp neighbors 203.0.113.2
show logging

Record the state and recent connection/reset information.

Step 4 — inspect interface policy

show ip interface GigabitEthernet0/0
show access-lists BLOCK-BGP-FSM-LAB

ACL counters can provide direct evidence that the BGP TCP packets are hitting the deny statements.

Step 5 — packet capture where available

A controlled capture can show repeated TCP SYN attempts without successful completion. Do not fabricate packet details if no capture was taken.

Step 6 — correct the proven layer

Remove the ACL or modify it to permit the authorized BGP transport. Do not change ASNs, timers or BGP path policy when the failure is clearly at transport filtering.


11. Root cause and correction

Root cause

The interface ACL blocked the TCP exchange required for BGP while permitting ICMP. IP reachability therefore looked healthy even though BGP could not complete transport/session establishment.

Correction

configure terminal
interface GigabitEthernet0/0
 no ip access-group BLOCK-BGP-FSM-LAB in
end

Or, in a production design, replace the test ACL with a deliberate control-plane policy that permits only authorized BGP peers rather than removing protection entirely.


12. Post-fix verification

show ip bgp summary
show ip bgp neighbors 203.0.113.2
show logging

Confirm:

  • state returns to Established;
  • uptime begins increasing;
  • message counters increment;
  • no recurring reset reason appears.

Then verify any expected routes and forwarding separately.


13. Rollback

The rollback for the fault is the ACL removal shown above.

If the ACL itself is part of a broader security policy, rollback should instead restore the previous known-good ACL version. Preserve:

  • the ACL before/after;
  • hit counters;
  • BGP neighbor detail;
  • syslog timestamps;
  • packet capture if taken.

Independent management access is mandatory when testing control-plane ACLs remotely.


14. Production lessons

Troubleshooting lesson

Use the FSM to identify the next dependency, not to guess the root cause.

Transport lesson

Ping success does not prove TCP/179 success.

Observability lesson

Fast FSM states may be missed by CLI polling. Logs, neighbor reset reasons and packet captures can show the actual sequence.

Change-management lesson

Control-plane ACL changes deserve the same rollback discipline as BGP policy changes because they can drop every route from a peer at once.

Standards lesson

The core FSM comes from RFC 4271, while later standards such as RFC 9687 can extend behavior. Always separate baseline protocol rules from implementation/version support.


15. Knowledge check

Q1

A BGP peer is repeatedly in Active. Should you first inspect LOCAL_PREF or TCP reachability?

Answer: TCP/peer reachability. LOCAL_PREF affects path selection after routes are exchanged; it does not establish the BGP transport session.

Q2

A peer reaches OpenSent and then resets. What category of configuration becomes more likely than a simple missing route to the peer?

Answer: OPEN/session-parameter problems such as ASN mismatch, BGP Identifier/OPEN validation, capability requirements or a received NOTIFICATION.

Q3

A peer is Established but receives zero expected routes. Is the FSM still the primary troubleshooting focus?

Answer: No. The session is operational. Move to route origination, received/accepted route state, policy, best path, next hop, RIB/FIB and forwarding.


16. Sources

  1. RFC 4271 — *A Border Gateway Protocol 4 (BGP-4)*, Section 8 FSM and related message processing.
  2. RFC 9687 — *Border Gateway Protocol 4 (BGP-4) Send Hold Timer*.
  3. Cisco — *BGP Command Reference: show ip bgp neighbors*.
  4. Cisco — *IP Routing Configuration Guide, Cisco IOS XE 17.x*.

Accessed: 2026-08-11.


Featured Post

Day 41 — BGP Confederations: Sub-AS Design, External View and Migration

1. Opening Confederations are another way to scale BGP inside a large administrative domain. They divide the domain into member autonomous systems while presenting a single confederation identifier to external peers. They are powerful, but their operational model is more complex than simply 'using private ASNs inside.' The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture . 2. Concept and standards behavior RFC 5065 defines AS_CONFED_SEQUENCE and AS_CONFED_SET and how member-AS relationships are represented. Confederation external sessions have eBGP-like properties inside the confederation, while the confederation is presented externally as one AS. Modern guidance must also account for the fact that RFC 9774 prohibits new origination of AS_SET/AS_CONFED_SET in ordinary aggregation c...