Day 37 — Why iBGP Full Mesh Does Not Scale
1. Opening
iBGP does not automatically relay every route learned from another iBGP peer. That behavior prevents simple internal loops but creates a scaling consequence: without another mechanism, every iBGP speaker that must share routes with every other speaker needs direct iBGP adjacency.
The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.
| Why iBGP Full Mesh Does Not Scale |
2. Concept and standards behavior
In a plain iBGP design, routes learned from one iBGP peer are not advertised to another iBGP peer. RFC 4456 was created specifically because the resulting full mesh becomes operationally expensive as the number of speakers grows. With n routers, a full mesh requires n(n-1)/2 sessions. Ten speakers require 45 sessions; 50 require 1,225. Session count is not the only cost: every new router also increases configuration, policy and troubleshooting surfaces.
A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.
3. Scenario
Build six routers inside AS65010. R1 originates 203.0.113.0/24. Initially configure only a chain of iBGP sessions R1-R2-R3-R4-R5-R6. The expected failure is that the route does not simply propagate across the chain. Then compare that with a true full mesh.
Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.
Success criteria
- Every speaker learns the routes it is supposed to learn.
- The propagation rule can be explained before commands are applied.
- Loop-prevention attributes are visible where the feature uses them.
- A negative test proves that an invalid or unintended propagation does not occur.
- Rollback returns the topology to a known-good control-plane state.
4. Topology diagram
| Why iBGP Full Mesh Does Not Scale |
5. Prerequisites
- IOS XE 17.18.x documentation baseline.
- Stable IGP reachability among BGP loopbacks.
- Explicit
update-sourceand appropriate neighbor reachability for loopback-based sessions. - Authentication only if the lab image and design have been validated for it.
- Independent management access before changing route-reflection/confederation policy.
- Pre-change capture of BGP summary, selected paths and relevant neighbor state.
6. Baseline configuration
Topic-specific configuration excerpt — not a complete device configuration.
router bgp 65010
bgp log-neighbor-changes
neighbor 192.0.2.2 remote-as 65010
neighbor 192.0.2.2 update-source Loopback0
!
address-family ipv4
neighbor 192.0.2.2 activate
exit-address-family
Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.
7. Verification before modification
Use a combination of topology, session and route evidence:
show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route
For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.
8. Controlled modification
Add the missing direct iBGP adjacencies until the six-node topology forms a full mesh. Predict the required session count before configuring it and verify that R6 now learns R1's prefix.
Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.
9. Fault injection
Illustrative lab — not a real incident.
Remove one direct session that is the only way a particular speaker receives R1's path while keeping IGP reachability intact. The session topology, not IP reachability, becomes the fault.
Capture the state before and after the fault. A useful fault must change one causal variable only.
10. Step-by-step troubleshooting
- Confirm IGP reachability among BGP endpoints.
- Confirm TCP/BGP sessions are Established.
- Identify where the route originated.
- Trace the route one BGP hop at a time.
- Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
- Inspect best-path state before assuming propagation is broken.
- Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
- Inspect policy and next-hop reachability.
- Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
- Apply the smallest proven correction.
- Re-run the same positive and negative tests.
- Observe stability before closing the change.
11. Root cause and correction
The proven cause is incomplete iBGP topology combined with the rule that iBGP-learned routes are not normally re-advertised to other iBGP peers. Correct by building the intended full mesh or introducing an explicit scaling architecture such as route reflection/confederations.
The correction must address the architectural cause rather than adding random neighbor statements until the route appears.
12. Post-fix verification
Verify:
- all intended sessions remain Established;
- the expected prefix is learned by the intended speakers;
- reflected attributes are correct where applicable;
- next hop remains reachable;
- no routing loop is created;
- the negative propagation test still passes.
13. Rollback
Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.
14. Production lessons
Use the full mesh as a conceptual baseline, not as an automatic production recommendation. Count sessions, policy attachment points, change burden and failure domains before deciding how to scale.
Deep-dive engineering notes
Why the session formula matters operationally
The mathematical growth of a full mesh is easy to state but the operational consequence is more important. Every new iBGP speaker can require another neighbor relationship on every existing speaker. That means another authentication relationship if authentication is used, another set of address-family activation decisions, another point where inbound/outbound policy can diverge, another state machine to monitor, and another adjacency to account for during maintenance. The design burden therefore grows faster than the router count.
A full mesh can still be reasonable in a small, stable control-plane core. The lesson is not “full mesh is bad”; the lesson is that it is a scaling baseline whose costs must be consciously accepted. A five-router control plane and a fifty-router control plane are different engineering problems even if both fit in memory.
Why an iBGP chain fails even when every IP hop works
An engineer can ping R1 from R6 and still have missing BGP routes. That distinction is important: IGP reachability proves the TCP endpoints can potentially communicate; it does not override the iBGP advertisement rule. In the chain lab, R2 learning R1's prefix over iBGP does not grant R2 permission to advertise that same path onward to R3 as ordinary iBGP. This is why adding static routes, changing interface costs, or repeatedly clearing sessions does not solve the architectural fault.
A strong troubleshooting workflow therefore records the route at each speaker. If R2 has the prefix and R3 does not, the next question is not “is R3 reachable?” but “what relationship allows R2 to propagate this path to R3?” That question leads naturally to route reflection or confederations.
Scaling metrics worth recording
For a production design review, record at least: number of BGP speakers, expected sessions, number of address families, number of policy variants, number of route sources, route churn, expected convergence objective, and operational ownership. Session count alone can underestimate complexity if each session carries several AFI/SAFI families and unique policy.
Failure-domain implication
A full mesh distributes dependency: there is no single route reflector whose loss removes every reflected path. But it also distributes change surface across every speaker. Route reflection centralizes some control-plane functions and therefore trades adjacency scale for new architectural dependencies. That trade-off is the bridge to Day 38.
15. Knowledge check
- Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
- What evidence distinguishes “route was never learned” from “route was learned but not selected”?
- What negative test proves the scaling feature has not introduced unintended propagation?
Answers
- The relationship among BGP speakers and the rules governing propagation.
- Per-prefix BGP path detail and neighbor route evidence.
- Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.
16. Sources
- RFC 4271 — BGP-4 base specification and iBGP behavior
- RFC 4456 — Route Reflection, motivation for avoiding iBGP full mesh
- Cisco IOS XE 17.18.x BGP configuration guide