Day 40 — Route-Reflector Path Hiding and Optimal Route Reflection

1. Opening

Route reflection can hide paths. A reflector generally selects a best path and reflects that view; clients may therefore never learn alternatives that would have been available in a full mesh. This can create suboptimal egress and can interact with convergence.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 4456's reduction of routing information creates the path-hiding problem. RFC 7911 ADD-PATH allows multiple paths for the same NLRI to be advertised when negotiated. RFC 9107 defines Optimal Route Reflection, in which the reflector can calculate client-appropriate optimal paths using client location/IGP perspective. These mechanisms solve different problems and must not be conflated.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Use two exits for 203.0.113.0/24. RR1 has lower IGP cost to Exit-A, while Client-B is physically closer to Exit-B. Without additional mechanisms, RR1's own best-path view can cause Client-B to receive a path that is not locally optimal.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Route-Reflector Path Hiding and Optimal Route Reflection
Route-Reflector Path Hiding and Optimal Route Reflection

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 ! Route-reflector baseline omitted for brevity.
 !
 address-family ipv4
  ! Verify platform support and exact ADD-PATH/ORR syntax
  ! before enabling either mechanism.
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

First reproduce path hiding with ordinary reflection. Then, in a capability-appropriate lab, compare the information made available by ADD-PATH with the client-specific selection objective of ORR.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Manipulate IGP cost so the reflector and a distant client have different closest exits. If the client never receives the alternative, troubleshooting the client's local BGP policy alone cannot fix the missing information.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The cause is information reduction at the reflector, not necessarily an incorrect client policy. Correct by redesigning reflector placement or using a supported path-diversity/optimal-reflection mechanism after validating platform behavior.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Path visibility is a design input. ADD-PATH increases path diversity; ORR changes which path is optimal for a client perspective. More paths can improve decisions but also increase control-plane state.

Deep-dive engineering notes

Path hiding begins with information reduction

In a full mesh, a speaker can potentially learn alternate paths directly from their originators. With route reflection, the reflector's decision can become the information boundary. If RR1 selects Exit-A and never advertises Exit-B's alternative to a client, the client cannot select Exit-B regardless of its own lower IGP cost to that exit.

This is why local troubleshooting at the client can be misleading. The client's BGP table may be internally consistent; it simply never received the alternative. The correct diagnostic question becomes “which paths were available at the reflector, which one did it select, and which paths did it advertise?”

ADD-PATH and ORR solve different dimensions

RFC 7911 ADD-PATH allows multiple paths for the same NLRI to be advertised with path identifiers when capability negotiation permits it. The client gains more path diversity and can make a richer local decision. The cost is additional control-plane state and update volume.

RFC 9107 Optimal Route Reflection aims at a different problem: the RR can select paths using a perspective appropriate to the client, such as the client's location in the IGP topology. ORR is therefore about selecting a client-optimal view; ADD-PATH is about advertising multiple paths. A design may use one, the other, neither, or platform-specific combinations.

Reflector placement can mitigate or amplify hiding

If the reflector's network location and IGP perspective closely match its clients, its selected path may already be acceptable for those clients. Centralizing a single RR far from diverse client populations increases the chance that the RR's hot-potato choice differs from the client's best exit.

Scale trade-off

Path diversity is not free. More paths can increase Adj-RIB-Out state, memory, update processing, and downstream selection work. Before enabling a feature globally, identify the prefixes or address families that benefit and establish measurable convergence or optimality objectives.

Verification strategy

Capture the same prefix at Exit-A, Exit-B, RR, and client. Record all available paths at the RR, the selected path, and the client's received paths. That four-point evidence proves whether the issue is path hiding, client policy, or next-hop resolution.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — route reflection
  • RFC 7911 — ADD-PATH
  • RFC 9107 — BGP Optimal Route Reflection (ORR)

RETICUX BGP Mastery — Day 39— Redundant and Hierarchical Route Reflectors

1. Opening

One route reflector solves session scale but can become a control-plane single point of failure. Adding a second reflector improves resilience only if clients actually peer to it, policies are consistent, and the reflectors do not share the same hidden dependency.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 4456 permits multiple route reflectors and cluster designs. Redundancy must be evaluated end to end: reflector processes, IGP reachability, physical failure domains, client adjacency, policy distribution and management. Hierarchical reflection can reduce session fan-out further, but it also increases path-selection layers and troubleshooting complexity.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Build RR1 and RR2 in AS65010. Each client peers to both. Place the two RRs on separate simulated failure domains. Compare this with a false-redundant design where both RRs depend on the same transit interface or where half the clients peer to only one.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Redundant and Hierarchical Route Reflectors
Redundant and Hierarchical Route Reflectors


5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 neighbor 192.0.2.101 remote-as 65010
 neighbor 192.0.2.102 remote-as 65010
 !
 address-family ipv4
  neighbor 192.0.2.101 activate
  neighbor 192.0.2.102 activate
 exit-address-family
!
! On each RR, mark the client as route-reflector-client
! according to that reflector's configuration.

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Dual-home every client to RR1 and RR2, then fail RR1. The route should remain available through RR2 if the redundancy is real.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Keep RR2 operational but remove the client's session to it. Then fail RR1. This demonstrates that two reflector devices do not create redundancy unless the client relationship is also redundant.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The failure is architectural: the supposedly redundant path is not end-to-end. Correct the client peering and validate independent IGP/transport reachability to both reflectors.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Redundancy should survive a realistic shared-risk failure, not just a process restart. Hierarchy is justified by scale and geography—not by a desire to draw more layers in the diagram.

Deep-dive engineering notes

Two devices are not automatically redundant

A recurring design error is to count boxes instead of dependencies. RR1 and RR2 can be separate routers yet share the same rack power, aggregation link, IGP adjacency, management plane, configuration pipeline, or software defect. Redundancy should be assessed using shared-risk groups rather than device count.

Clients also need redundant adjacency. If Client-A peers only to RR1 and Client-B peers only to RR2, each client still has a single reflector dependency. A properly dual-homed client design lets either reflector continue distributing reachable paths after the other fails.

Cluster design deserves explicit intent

Cluster IDs participate in reflection loop prevention. Whether redundant RRs share a cluster ID or use distinct clusters depends on the intended reflection design and path behavior. Do not copy a cluster-id convention from another network without understanding what paths may be reflected between the RRs and how CLUSTER_LIST processing will behave.

Hierarchical reflection is a scale tool with a cost

Hierarchical RRs can reduce the number of sessions and align control-plane structure with geography or network tiers. But every reflection layer can further reduce available path information and adds another place where policy or best-path choice can become suboptimal. A hierarchy should therefore have a measurable purpose: peer-scale reduction, regional autonomy, failure containment, or platform limit management.

Failure testing for real redundancy

Test at least four events separately: loss of an RR BGP process, loss of the RR's IGP reachability, loss of a client-to-RR session, and loss of the shared transport beneath both RRs. A design that survives only the first event is not convincingly redundant.

Consistency versus independence

Redundant reflectors generally need consistent baseline policy, but complete operational coupling can create correlated failure. Configuration automation should enforce intended equivalence while allowing controlled staggered deployment, canary validation, and independent rollback. Reliability comes from both similarity of intent and separation of failure.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — route-reflector clusters and loop prevention
  • Cisco IOS XE 17.x — BGP route-reflector configuration
  • RFC 7964 — route oscillation considerations with route reflection/confederations

Day 38 — Route Reflectors: Clients, Non-Clients and Reflection Rules

1. Opening

Route reflection changes the iBGP propagation model so selected speakers can re-advertise iBGP-learned routes. That removes the universal full-mesh requirement, but it also means the engineer must understand client and non-client rules rather than treating the reflector as a transparent relay.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.



Route Reflectors: Clients, Non-Clients and Reflection Rules
Route Reflectors: Clients, Non-Clients and Reflection Rules


2. Concept and standards behavior

RFC 4456 defines route reflectors, clients, ORIGINATOR_ID and CLUSTER_LIST. In simplified terms, a route learned from an RR client may be reflected to other clients and non-clients; a route learned from a non-client is reflected to clients, but not to other non-clients. The RR still runs the BGP decision process—it does not blindly flood every path.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Use RR1 as route reflector with R1, R2 and R3 as clients, plus R4 as a non-client iBGP peer. Originate one prefix at R1 and another at R4, then trace which speakers receive each path.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Route Reflectors: Clients, Non-Clients and Reflection Rules
Route Reflectors: Clients, Non-Clients and Reflection Rules

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 neighbor 192.0.2.11 remote-as 65010
 neighbor 192.0.2.12 remote-as 65010
 neighbor 192.0.2.13 remote-as 65010
 neighbor 192.0.2.14 remote-as 65010
 !
 address-family ipv4
  neighbor 192.0.2.11 route-reflector-client
  neighbor 192.0.2.12 route-reflector-client
  neighbor 192.0.2.13 route-reflector-client
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Convert R3 from ordinary non-client to RR client and predict which reflected routes become newly visible to it. Verify path detail and reflected attributes.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Misclassify two edge speakers as non-clients and assume the RR will reflect routes between them. Both sessions stay Established, yet route visibility is incomplete.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The root cause is a wrong RR relationship, not a failed TCP/BGP session. Correct the intended client/non-client role and verify reflected route attributes.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Route reflectors reduce adjacency requirements by changing propagation semantics. Document client groups explicitly; an Established session alone does not prove the route will be reflected.

Deep-dive engineering notes

Reflection is selective propagation, not flooding

A route reflector still chooses BGP paths. It does not automatically copy every received path to every client. This distinction becomes crucial when several exits advertise the same NLRI. A client may receive only the reflector's selected view unless a separate path-diversity mechanism is used.

The client/non-client rules should be written into the design documentation. A route learned from a client can be reflected to clients and non-clients. A route learned from a non-client is reflected to clients but not to other non-clients. An engineer who treats all iBGP sessions attached to an RR as equivalent can create a topology where all sessions are Established yet some internal destinations disappear.

ORIGINATOR_ID and CLUSTER_LIST are diagnostic evidence

ORIGINATOR_ID identifies the BGP speaker that originally advertised a reflected route inside the AS. CLUSTER_LIST records clusters traversed by the reflected path and supports loop prevention. These fields are not decoration; they help prove that a route has been reflected and can explain why an RR rejects a path that appears to have returned through its own cluster.

When troubleshooting, compare two paths for the same prefix and explicitly inspect these attributes. If a path disappears after a cluster design change, the CLUSTER_LIST is often more useful than repeated show bgp summary output.

Next-hop behavior remains a separate problem

Route reflection changes iBGP advertisement rules but does not magically fix next-hop reachability. If a client receives a reflected route whose NEXT_HOP is unreachable, the control plane may show the route while the forwarding result remains broken or the path may not be usable. Maintain IGP reachability to relevant next hops and apply next-hop-self only where the design actually requires it.

Placement principle

Place reflectors for control-plane stability and topology awareness, not merely where configuration is convenient. A route reflector should not be forced through a fragile access failure domain just because it is physically near clients. Dedicated control-plane nodes, redundant transport, and consistent policy often matter more than raw hop count.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — BGP Route Reflection
  • RFC 4271 — base BGP behavior
  • Cisco IOS XE 17.18.x BGP configuration guide

Day 37 — Why iBGP Full Mesh Does Not Scale

1. Opening

iBGP does not automatically relay every route learned from another iBGP peer. That behavior prevents simple internal loops but creates a scaling consequence: without another mechanism, every iBGP speaker that must share routes with every other speaker needs direct iBGP adjacency.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.


Why iBGP Full Mesh Does Not Scale
Why iBGP Full Mesh Does Not Scale


2. Concept and standards behavior

In a plain iBGP design, routes learned from one iBGP peer are not advertised to another iBGP peer. RFC 4456 was created specifically because the resulting full mesh becomes operationally expensive as the number of speakers grows. With n routers, a full mesh requires n(n-1)/2 sessions. Ten speakers require 45 sessions; 50 require 1,225. Session count is not the only cost: every new router also increases configuration, policy and troubleshooting surfaces.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Build six routers inside AS65010. R1 originates 203.0.113.0/24. Initially configure only a chain of iBGP sessions R1-R2-R3-R4-R5-R6. The expected failure is that the route does not simply propagate across the chain. Then compare that with a true full mesh.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Why iBGP Full Mesh Does Not Scale
Why iBGP Full Mesh Does Not Scale

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 bgp log-neighbor-changes
 neighbor 192.0.2.2 remote-as 65010
 neighbor 192.0.2.2 update-source Loopback0
 !
 address-family ipv4
  neighbor 192.0.2.2 activate
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Add the missing direct iBGP adjacencies until the six-node topology forms a full mesh. Predict the required session count before configuring it and verify that R6 now learns R1's prefix.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Remove one direct session that is the only way a particular speaker receives R1's path while keeping IGP reachability intact. The session topology, not IP reachability, becomes the fault.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The proven cause is incomplete iBGP topology combined with the rule that iBGP-learned routes are not normally re-advertised to other iBGP peers. Correct by building the intended full mesh or introducing an explicit scaling architecture such as route reflection/confederations.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Use the full mesh as a conceptual baseline, not as an automatic production recommendation. Count sessions, policy attachment points, change burden and failure domains before deciding how to scale.

Deep-dive engineering notes

Why the session formula matters operationally

The mathematical growth of a full mesh is easy to state but the operational consequence is more important. Every new iBGP speaker can require another neighbor relationship on every existing speaker. That means another authentication relationship if authentication is used, another set of address-family activation decisions, another point where inbound/outbound policy can diverge, another state machine to monitor, and another adjacency to account for during maintenance. The design burden therefore grows faster than the router count.

A full mesh can still be reasonable in a small, stable control-plane core. The lesson is not “full mesh is bad”; the lesson is that it is a scaling baseline whose costs must be consciously accepted. A five-router control plane and a fifty-router control plane are different engineering problems even if both fit in memory.

Why an iBGP chain fails even when every IP hop works

An engineer can ping R1 from R6 and still have missing BGP routes. That distinction is important: IGP reachability proves the TCP endpoints can potentially communicate; it does not override the iBGP advertisement rule. In the chain lab, R2 learning R1's prefix over iBGP does not grant R2 permission to advertise that same path onward to R3 as ordinary iBGP. This is why adding static routes, changing interface costs, or repeatedly clearing sessions does not solve the architectural fault.

A strong troubleshooting workflow therefore records the route at each speaker. If R2 has the prefix and R3 does not, the next question is not “is R3 reachable?” but “what relationship allows R2 to propagate this path to R3?” That question leads naturally to route reflection or confederations.

Scaling metrics worth recording

For a production design review, record at least: number of BGP speakers, expected sessions, number of address families, number of policy variants, number of route sources, route churn, expected convergence objective, and operational ownership. Session count alone can underestimate complexity if each session carries several AFI/SAFI families and unique policy.

Failure-domain implication

A full mesh distributes dependency: there is no single route reflector whose loss removes every reflected path. But it also distributes change surface across every speaker. Route reflection centralizes some control-plane functions and therefore trades adjacency scale for new architectural dependencies. That trade-off is the bridge to Day 38.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4271 — BGP-4 base specification and iBGP behavior
  • RFC 4456 — Route Reflection, motivation for avoiding iBGP full mesh
  • Cisco IOS XE 17.18.x BGP configuration guide

RETICUX BGP Mastery — Day 36 --- Dual-Provider Traffic-Engineering Capstone

1. Opening

A production multihoming design must work as a system. Outbound and inbound engineering are different problems, and a design that optimizes both during steady state can fail badly when one provider, one prefix or one policy condition disappears.

The operational objective is not simply to make BGP choose a route. It is to make the intended behavior predictable during the normal state, degraded state, and rollback state.


Dual-Provider Traffic-Engineering Capstone
Dual-Provider Traffic-Engineering Capstone


2. Concept and standards behavior

The capstone combines previously isolated mechanisms. LOCAL_PREF controls enterprise outbound choice. Export policy controls what each provider can learn. A provider community is treated as a provider-specific signal. Conditional advertisement supplies a backup route only when a defined BGP condition changes. RFC 8212's explicit-policy principle is used as the safety mindset at every external boundary.

BGP remains a policy protocol. RFC 4271 provides the protocol framework, while several practical traffic-engineering controls are Cisco or operator-policy mechanisms. Evaluate each design at three layers: protocol behavior, IOS XE implementation, and neighboring-AS policy.

Separate route eligibility, route selection, and route export. Eligibility asks whether a path is usable. Selection determines the local best path. Export policy determines what a neighbor is permitted to learn.

3. Scenario

AS 65010 is dual-homed to ISP-A AS 65020 and ISP-B AS 65030. The enterprise service prefix is 203.0.113.0/24. Transit links use 192.0.2.0/30 and 198.51.100.0/30.

Success criteria 1. Primary and backup behavior is explicit. 2. No route is exported merely because it exists locally. 3. Failure behavior is observable in BGP and advertised-route evidence. 4. Rollback is independent of session redesign. 5. A negative test proves unintended advertisement is absent.

4. Topology

Dual-Provider Traffic-Engineering Capstone
Dual-Provider Traffic-Engineering Capstone

ISP-A AS65020 --- 192.0.2.0/30 --- EDGE1/AS65010
                                      |
                               203.0.113.0/24
                                      |
ISP-B AS65030 --- 198.51.100.0/30 --- EDGE2/AS65010

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • IPv4 unicast activated for both eBGP peers.
  • Service prefix valid for origination.
  • Independent management access.
  • Pre-change capture of BGP state and advertised routes.

6. Baseline configuration

Topic-specific configuration excerpt --- not a complete device configuration.

ip prefix-list PL-SERVICE permit 203.0.113.0/24

route-map FROM-ISP-A permit 10
 set local-preference 200
route-map FROM-ISP-B permit 10
 set local-preference 150

route-map EXPORT-SERVICE permit 10
 match ip address prefix-list PL-SERVICE

router bgp 65010
 address-family ipv4
  neighbor 192.0.2.1 route-map FROM-ISP-A in
  neighbor 198.51.100.1 route-map FROM-ISP-B in
  neighbor 192.0.2.1 route-map EXPORT-SERVICE out
  ! Add provider-specific community and conditional backup policy
  ! only after validating the provider contract and exact lab syntax.
 exit-address-family

7. Verification before modification

show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show bgp ipv4 unicast neighbors 192.0.2.1 advertised-routes
show bgp ipv4 unicast neighbors 198.51.100.1 advertised-routes
show route-map
show ip prefix-list

The question is not whether the policy object exists; it is whether the intended route is selected, matched and actually exported to the intended neighbor.

8. Controlled modification

Stage the change: first explicit import/export filters, then outbound LOCAL_PREF, then inbound TE signal, then conditional backup behavior. Validate each stage before adding the next.

Predict Adj-RIB-Out behavior before applying the change. Re-evaluate policy using the least disruptive supported mechanism.

9. Fault injection

Illustrative lab — not a real incident.

Inject three faults separately: ISP-A session loss, loss of only the primary condition route while the session stays up, and removal of community transmission. Each should produce a different observable symptom.

Change one variable only, capture the changed state, and compare it with the baseline.

10. Troubleshooting

  1. Confirm affected prefix and traffic direction.
  2. Confirm eBGP sessions are Established.
  3. Confirm local route eligibility/origination.
  4. Inspect selected BGP route.
  5. Inspect match objects.
  6. Inspect route-map sequence and implicit deny.
  7. Inspect advertised routes per provider.
  8. Confirm policy direction.
  9. Determine whether route refresh is required.
  10. Run positive traffic test.
  11. Run negative/containment test.
  12. Correct the smallest proven cause and repeat the same evidence set.

11. Root cause and correction

A capstone failure is usually caused by interacting controls being changed together. Restore the last known-good stage, prove baseline behavior, then reintroduce one control at a time.

12. Post-fix verification

Verify session state, route acceptance, selected path, exported attributes, advertised routes, forwarding and the negative test. Local advertisement does not prove remote selection.

13. Rollback

Revert only the new policy action, re-evaluate policy with the least disruptive supported method, and confirm both providers return to baseline. Roll back immediately for loss of all reachability, unintended transit, unapproved deaggregation or export outside the approved prefix set.

14. Production lessons

The strongest multihoming design is not the one with the most attributes. It is the one whose normal, partial-failure, total-failure and rollback states are all explicit and testable.

15. Knowledge check

map?

policy?

leaking?

  1. Why is advertised-routes stronger evidence than displaying a route
  2. Which parts are controlled locally and which depend on upstream
  3. What negative test proves the restricted/backup advertisement is not

Answers

inbound selection are remote policy.

routes in the healthy state.

  1. It validates resulting per-neighbor export state.
  2. Local selection/export are local; remote LOCAL_PREF, propagation and
  3. Verify the protected prefix is absent from the neighbor's advertised

16. Sources

  • RFC 4271
  • RFC 1997
  • RFC 8212
  • Cisco IOS XE 17.x --- BGP policy and conditional advertisement

Day 35 --- Hot-Potato vs Cold-Potato Routing

1. Opening

Hot-potato routing hands traffic to another AS at the nearest acceptable exit. Cold-potato routing intentionally carries traffic farther inside the local network before handoff. Neither is a BGP message type; both emerge from policy, topology and the interaction between BGP and the IGP.

The operational objective is not simply to make BGP choose a route. It is to make the intended behavior predictable during the normal state, degraded state, and rollback state.

#ccna #ccnp #cisco #network #engineer #BGP

Hot-Potato vs Cold-Potato Routing
Hot-Potato vs Cold-Potato Routing


2. Concept and standards behavior

When earlier BGP attributes tie, IGP cost to the BGP NEXT_HOP can influence which exit is selected. Operators can override that outcome using LOCAL_PREF or other policy. The key design question is whether the AS wants to minimize its own transport cost or control the egress location for performance, economics or service reasons.

BGP remains a policy protocol. RFC 4271 provides the protocol framework, while several practical traffic-engineering controls are Cisco or operator-policy mechanisms. Evaluate each design at three layers: protocol behavior, IOS XE implementation, and neighboring-AS policy.

Separate route eligibility, route selection, and route export. Eligibility asks whether a path is usable. Selection determines the local best path. Export policy determines what a neighbor is permitted to learn.

3. Scenario

AS 65010 is dual-homed to ISP-A AS 65020 and ISP-B AS 65030. The enterprise service prefix is 203.0.113.0/24. Transit links use 192.0.2.0/30 and 198.51.100.0/30.

Success criteria 1. Primary and backup behavior is explicit. 2. No route is exported merely because it exists locally. 3. Failure behavior is observable in BGP and advertised-route evidence. 4. Rollback is independent of session redesign. 5. A negative test proves unintended advertisement is absent.

4. Topology

Hot-Potato vs Cold-Potato Routing
Hot-Potato vs Cold-Potato Routing

ISP-A AS65020 --- 192.0.2.0/30 --- EDGE1/AS65010
                                      |
                               203.0.113.0/24
                                      |
ISP-B AS65030 --- 198.51.100.0/30 --- EDGE2/AS65010

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • IPv4 unicast activated for both eBGP peers.
  • Service prefix valid for origination.
  • Independent management access.
  • Pre-change capture of BGP state and advertised routes.

6. Baseline configuration

Topic-specific configuration excerpt --- not a complete device configuration.

router bgp 65010
 address-family ipv4
  neighbor 192.0.2.1 remote-as 65020
  neighbor 198.51.100.1 remote-as 65030
 exit-address-family

! IGP configuration is topology-specific.
! Verify recursive cost to each BGP NEXT_HOP before changing BGP policy.

7. Verification before modification

show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show bgp ipv4 unicast neighbors 192.0.2.1 advertised-routes
show bgp ipv4 unicast neighbors 198.51.100.1 advertised-routes
show route-map
show ip prefix-list

The question is not whether the policy object exists; it is whether the intended route is selected, matched and actually exported to the intended neighbor.

8. Controlled modification

First allow equal higher-priority BGP attributes so the lower IGP cost wins. Then set a higher LOCAL_PREF on the farther exit and verify that policy overrides hot-potato behavior.

Predict Adj-RIB-Out behavior before applying the change. Re-evaluate policy using the least disruptive supported mechanism.

9. Fault injection

Illustrative lab — not a real incident.

Increase the IGP cost to the currently selected next hop without changing LOCAL_PREF. If higher BGP policy already fixes the exit, the traffic should not move merely because the IGP cost changed.

Change one variable only, capture the changed state, and compare it with the baseline.

10. Troubleshooting

  1. Confirm affected prefix and traffic direction.
  2. Confirm eBGP sessions are Established.
  3. Confirm local route eligibility/origination.
  4. Inspect selected BGP route.
  5. Inspect match objects.
  6. Inspect route-map sequence and implicit deny.
  7. Inspect advertised routes per provider.
  8. Confirm policy direction.
  9. Determine whether route refresh is required.
  10. Run positive traffic test.
  11. Run negative/containment test.
  12. Correct the smallest proven cause and repeat the same evidence set.

11. Root cause and correction

The failure is troubleshooting only the IGP when a higher-priority BGP attribute already determines the exit. Correct by walking the best-path decision in order.

12. Post-fix verification

Verify session state, route acceptance, selected path, exported attributes, advertised routes, forwarding and the negative test. Local advertisement does not prove remote selection.

13. Rollback

Revert only the new policy action, re-evaluate policy with the least disruptive supported method, and confirm both providers return to baseline. Roll back immediately for loss of all reachability, unintended transit, unapproved deaggregation or export outside the approved prefix set.

14. Production lessons

Hot versus cold potato is an architecture decision with cost, latency and failure-domain consequences. Document which layer is intended to own the exit choice.

15. Knowledge check

map?

policy?

leaking?

  1. Why is advertised-routes stronger evidence than displaying a route
  2. Which parts are controlled locally and which depend on upstream
  3. What negative test proves the restricted/backup advertisement is not

Answers

inbound selection are remote policy.

routes in the healthy state.

  1. It validates resulting per-neighbor export state.
  2. Local selection/export are local; remote LOCAL_PREF, propagation and
  3. Verify the protected prefix is absent from the neighbor's advertised

16. Sources

  • RFC 4271 --- BGP decision process
  • Cisco IOS XE 17.x --- BGP best-path/next-hop behavior

Day 34 --- Provider Communities for Inbound Traffic Engineering



1. Opening

Many providers publish communities that customers can attach to routes to request specific upstream actions. These can be more precise than blind prepending because the provider maps the community to its own policy. The critical rule is that the meaning is provider-specific unless a community is standardized.

The operational objective is not simply to make BGP choose a route. It is to make the intended behavior predictable during the normal state, degraded state, and rollback state.



Provider Communities for Inbound Traffic Engineering
Provider Communities for Inbound Traffic Engineering

2. Concept and standards behavior

Standard communities such as NO_EXPORT have defined semantics, but a provider's traffic-engineering communities are contractual/operator policy. Never invent a community value. A production post must cite the provider's current documentation before showing an actual value.

BGP remains a policy protocol. RFC 4271 provides the protocol framework, while several practical traffic-engineering controls are Cisco or operator-policy mechanisms. Evaluate each design at three layers: protocol behavior, IOS XE implementation, and neighboring-AS policy.

Separate route eligibility, route selection, and route export. Eligibility asks whether a path is usable. Selection determines the local best path. Export policy determines what a neighbor is permitted to learn.

3. Scenario

AS 65010 is dual-homed to ISP-A AS 65020 and ISP-B AS 65030. The enterprise service prefix is 203.0.113.0/24. Transit links use 192.0.2.0/30 and 198.51.100.0/30.

Success criteria 1. Primary and backup behavior is explicit. 2. No route is exported merely because it exists locally. 3. Failure behavior is observable in BGP and advertised-route evidence. 4. Rollback is independent of session redesign. 5. A negative test proves unintended advertisement is absent.

4. Topology


Hot-Potato vs Cold-Potato Routing
Hot-Potato vs Cold-Potato Routing

ISP-A AS65020 --- 192.0.2.0/30 --- EDGE1/AS65010
                                      |
                               203.0.113.0/24
                                      |
ISP-B AS65030 --- 198.51.100.0/30 --- EDGE2/AS65010

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • IPv4 unicast activated for both eBGP peers.
  • Service prefix valid for origination.
  • Independent management access.
  • Pre-change capture of BGP state and advertised routes.

6. Baseline configuration

Topic-specific configuration excerpt --- not a complete device configuration.

ip prefix-list PL-SERVICE permit 203.0.113.0/24
route-map TO-ISP-B permit 10
 match ip address prefix-list PL-SERVICE
 ! Set only a community value documented by ISP-B.
 ! Example intentionally omitted from canonical config.

router bgp 65010
 address-family ipv4
  neighbor 198.51.100.1 send-community
  neighbor 198.51.100.1 route-map TO-ISP-B out
 exit-address-family

7. Verification before modification

show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show bgp ipv4 unicast neighbors 192.0.2.1 advertised-routes
show bgp ipv4 unicast neighbors 198.51.100.1 advertised-routes
show route-map
show ip prefix-list

The question is not whether the policy object exists; it is whether the intended route is selected, matched and actually exported to the intended neighbor.

8. Controlled modification

In the lab, emulate ISP-B policy with a documentation-only private community such as 65030:90 and map it to lower LOCAL_PREF inside the simulated provider. Clearly label that value as lab-only, not a real provider convention.

Predict Adj-RIB-Out behavior before applying the change. Re-evaluate policy using the least disruptive supported mechanism.

9. Fault injection

Illustrative lab — not a real incident.

Remove send-community while leaving the route map in place. The route remains advertised but the provider no longer receives the policy signal.

Change one variable only, capture the changed state, and compare it with the baseline.

10. Troubleshooting

  1. Confirm affected prefix and traffic direction.
  2. Confirm eBGP sessions are Established.
  3. Confirm local route eligibility/origination.
  4. Inspect selected BGP route.
  5. Inspect match objects.
  6. Inspect route-map sequence and implicit deny.
  7. Inspect advertised routes per provider.
  8. Confirm policy direction.
  9. Determine whether route refresh is required.
  10. Run positive traffic test.
  11. Run negative/containment test.
  12. Correct the smallest proven cause and repeat the same evidence set.

11. Root cause and correction

The failure is assuming that setting a community locally proves it was transmitted or honored. Correct by enabling the required community exchange, verifying the outgoing attribute, and validating the provider-side policy.

12. Post-fix verification

Verify session state, route acceptance, selected path, exported attributes, advertised routes, forwarding and the negative test. Local advertisement does not prove remote selection.

13. Rollback

Revert only the new policy action, re-evaluate policy with the least disruptive supported method, and confirm both providers return to baseline. Roll back immediately for loss of all reachability, unintended transit, unapproved deaggregation or export outside the approved prefix set.

14. Production lessons

Provider communities can create precise inbound TE, blackholing or propagation controls, but only documented meanings are safe. Maintain a provider-community registry with owner, source URL, date checked and rollback behavior.

15. Knowledge check


  1. Why is advertised-routes stronger evidence than displaying a route
  2. Which parts are controlled locally and which depend on upstream
  3. What negative test proves the restricted/backup advertisement is not

Answers

inbound selection are remote policy.

routes in the healthy state.

  1. It validates resulting per-neighbor export state.
  2. Local selection/export are local; remote LOCAL_PREF, propagation and
  3. Verify the protected prefix is absent from the neighbor's advertised

16. Sources

real deployments

  • RFC 1997 --- Communities
  • RFC 8092 --- Large Communities
  • Provider-specific community documentation must be authoritative for

Featured Post

Day 41 — BGP Confederations: Sub-AS Design, External View and Migration

1. Opening Confederations are another way to scale BGP inside a large administrative domain. They divide the domain into member autonomous systems while presenting a single confederation identifier to external peers. They are powerful, but their operational model is more complex than simply 'using private ASNs inside.' The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture . 2. Concept and standards behavior RFC 5065 defines AS_CONFED_SEQUENCE and AS_CONFED_SET and how member-AS relationships are represented. Confederation external sessions have eBGP-like properties inside the confederation, while the confederation is presented externally as one AS. Modern guidance must also account for the fact that RFC 9774 prohibits new origination of AS_SET/AS_CONFED_SET in ordinary aggregation c...