Day 41 — BGP Confederations: Sub-AS Design, External View and Migration

1. Opening

Confederations are another way to scale BGP inside a large administrative domain. They divide the domain into member autonomous systems while presenting a single confederation identifier to external peers. They are powerful, but their operational model is more complex than simply 'using private ASNs inside.'

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 5065 defines AS_CONFED_SEQUENCE and AS_CONFED_SET and how member-AS relationships are represented. Confederation external sessions have eBGP-like properties inside the confederation, while the confederation is presented externally as one AS. Modern guidance must also account for the fact that RFC 9774 prohibits new origination of AS_SET/AS_CONFED_SET in ordinary aggregation contexts; that does not erase RFC 5065's confederation architecture, but it reinforces the need to distinguish attributes and mechanisms carefully.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Create confederation identifier 65000 with member AS65010 and AS65020. Establish a session between members and an external peer to AS65100. Verify the path representation internally and the single confederation identity seen externally.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

A matching SVG/PNG diagram is stored in media/diagrams/.

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 bgp confederation identifier 65000
 bgp confederation peers 65020
 neighbor 192.0.2.20 remote-as 65020
!
! Member AS65020 uses the same confederation identifier
! and declares AS65010 as a confederation peer.

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Add a second member AS and trace a route across two member-AS boundaries to the external peer. Verify that external policy still sees the intended confederation AS.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Misconfigure a member relationship as ordinary eBGP by omitting the confederation peer declaration. Observe changes in path handling and session semantics.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The cause is inconsistent confederation membership configuration. Correct the member-AS and confederation-peer declarations on both sides, then re-check path attributes and external view.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Confederations are a control-plane architecture, not merely an ASN trick. Prefer the simplest design that meets scale, policy and failure-domain requirements; route reflection is more common and often easier to operate.

Deep-dive engineering notes

Confederation identity has two scopes

A confederation divides a large routing domain into member ASes for internal scaling and policy, while external peers see the confederation identifier rather than the internal member structure. This dual view is the core architectural idea. Engineers must know whether a policy is matching the member-AS representation used inside the confederation or the public/external AS representation seen outside.

Member-AS boundaries are neither ordinary iBGP nor ordinary Internet eBGP

Sessions between confederation members borrow eBGP-like behavior for scaling while preserving the confederation's single external identity. Treating them exactly like normal eBGP can lead to incorrect assumptions about AS-path representation and policy. Treating them like ordinary iBGP loses the reason confederations exist.

Route reflection versus confederations

Both reduce the universal iBGP full-mesh requirement, but their operational models differ. Route reflection changes propagation relationships inside one AS. Confederations introduce member-AS structure and explicit boundaries. A network with strong administrative regions and highly structured internal policy may benefit from confederation semantics, while many deployments favor route reflection because it is simpler and more common operationally.

Migration risk

A confederation migration changes session roles and path representation. Build coexistence carefully, define which routers move first, and test route policy that matches AS paths. A regex written for the pre-migration AS_PATH may behave differently when member-AS segments appear internally.

Oscillation awareness

RFC 7964 documents that route reflection or confederations can interact with certain topologies and policies to create persistent route oscillation. That is not a reason to avoid both mechanisms; it is a reason to validate policy ordering, MED behavior, topology and path visibility rather than assuming the scaling architecture is neutral.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 5065 — Autonomous System Confederations for BGP
  • RFC 4271 — BGP-4
  • RFC 7964 — persistent route oscillation considerations

Day 40 — Route-Reflector Path Hiding and Optimal Route Reflection

1. Opening

Route reflection can hide paths. A reflector generally selects a best path and reflects that view; clients may therefore never learn alternatives that would have been available in a full mesh. This can create suboptimal egress and can interact with convergence.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 4456's reduction of routing information creates the path-hiding problem. RFC 7911 ADD-PATH allows multiple paths for the same NLRI to be advertised when negotiated. RFC 9107 defines Optimal Route Reflection, in which the reflector can calculate client-appropriate optimal paths using client location/IGP perspective. These mechanisms solve different problems and must not be conflated.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Use two exits for 203.0.113.0/24. RR1 has lower IGP cost to Exit-A, while Client-B is physically closer to Exit-B. Without additional mechanisms, RR1's own best-path view can cause Client-B to receive a path that is not locally optimal.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Route-Reflector Path Hiding and Optimal Route Reflection
Route-Reflector Path Hiding and Optimal Route Reflection

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 ! Route-reflector baseline omitted for brevity.
 !
 address-family ipv4
  ! Verify platform support and exact ADD-PATH/ORR syntax
  ! before enabling either mechanism.
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

First reproduce path hiding with ordinary reflection. Then, in a capability-appropriate lab, compare the information made available by ADD-PATH with the client-specific selection objective of ORR.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Manipulate IGP cost so the reflector and a distant client have different closest exits. If the client never receives the alternative, troubleshooting the client's local BGP policy alone cannot fix the missing information.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The cause is information reduction at the reflector, not necessarily an incorrect client policy. Correct by redesigning reflector placement or using a supported path-diversity/optimal-reflection mechanism after validating platform behavior.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Path visibility is a design input. ADD-PATH increases path diversity; ORR changes which path is optimal for a client perspective. More paths can improve decisions but also increase control-plane state.

Deep-dive engineering notes

Path hiding begins with information reduction

In a full mesh, a speaker can potentially learn alternate paths directly from their originators. With route reflection, the reflector's decision can become the information boundary. If RR1 selects Exit-A and never advertises Exit-B's alternative to a client, the client cannot select Exit-B regardless of its own lower IGP cost to that exit.

This is why local troubleshooting at the client can be misleading. The client's BGP table may be internally consistent; it simply never received the alternative. The correct diagnostic question becomes “which paths were available at the reflector, which one did it select, and which paths did it advertise?”

ADD-PATH and ORR solve different dimensions

RFC 7911 ADD-PATH allows multiple paths for the same NLRI to be advertised with path identifiers when capability negotiation permits it. The client gains more path diversity and can make a richer local decision. The cost is additional control-plane state and update volume.

RFC 9107 Optimal Route Reflection aims at a different problem: the RR can select paths using a perspective appropriate to the client, such as the client's location in the IGP topology. ORR is therefore about selecting a client-optimal view; ADD-PATH is about advertising multiple paths. A design may use one, the other, neither, or platform-specific combinations.

Reflector placement can mitigate or amplify hiding

If the reflector's network location and IGP perspective closely match its clients, its selected path may already be acceptable for those clients. Centralizing a single RR far from diverse client populations increases the chance that the RR's hot-potato choice differs from the client's best exit.

Scale trade-off

Path diversity is not free. More paths can increase Adj-RIB-Out state, memory, update processing, and downstream selection work. Before enabling a feature globally, identify the prefixes or address families that benefit and establish measurable convergence or optimality objectives.

Verification strategy

Capture the same prefix at Exit-A, Exit-B, RR, and client. Record all available paths at the RR, the selected path, and the client's received paths. That four-point evidence proves whether the issue is path hiding, client policy, or next-hop resolution.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — route reflection
  • RFC 7911 — ADD-PATH
  • RFC 9107 — BGP Optimal Route Reflection (ORR)

RETICUX BGP Mastery — Day 39— Redundant and Hierarchical Route Reflectors

1. Opening

One route reflector solves session scale but can become a control-plane single point of failure. Adding a second reflector improves resilience only if clients actually peer to it, policies are consistent, and the reflectors do not share the same hidden dependency.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 4456 permits multiple route reflectors and cluster designs. Redundancy must be evaluated end to end: reflector processes, IGP reachability, physical failure domains, client adjacency, policy distribution and management. Hierarchical reflection can reduce session fan-out further, but it also increases path-selection layers and troubleshooting complexity.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Build RR1 and RR2 in AS65010. Each client peers to both. Place the two RRs on separate simulated failure domains. Compare this with a false-redundant design where both RRs depend on the same transit interface or where half the clients peer to only one.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Redundant and Hierarchical Route Reflectors
Redundant and Hierarchical Route Reflectors


5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 neighbor 192.0.2.101 remote-as 65010
 neighbor 192.0.2.102 remote-as 65010
 !
 address-family ipv4
  neighbor 192.0.2.101 activate
  neighbor 192.0.2.102 activate
 exit-address-family
!
! On each RR, mark the client as route-reflector-client
! according to that reflector's configuration.

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Dual-home every client to RR1 and RR2, then fail RR1. The route should remain available through RR2 if the redundancy is real.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Keep RR2 operational but remove the client's session to it. Then fail RR1. This demonstrates that two reflector devices do not create redundancy unless the client relationship is also redundant.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The failure is architectural: the supposedly redundant path is not end-to-end. Correct the client peering and validate independent IGP/transport reachability to both reflectors.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Redundancy should survive a realistic shared-risk failure, not just a process restart. Hierarchy is justified by scale and geography—not by a desire to draw more layers in the diagram.

Deep-dive engineering notes

Two devices are not automatically redundant

A recurring design error is to count boxes instead of dependencies. RR1 and RR2 can be separate routers yet share the same rack power, aggregation link, IGP adjacency, management plane, configuration pipeline, or software defect. Redundancy should be assessed using shared-risk groups rather than device count.

Clients also need redundant adjacency. If Client-A peers only to RR1 and Client-B peers only to RR2, each client still has a single reflector dependency. A properly dual-homed client design lets either reflector continue distributing reachable paths after the other fails.

Cluster design deserves explicit intent

Cluster IDs participate in reflection loop prevention. Whether redundant RRs share a cluster ID or use distinct clusters depends on the intended reflection design and path behavior. Do not copy a cluster-id convention from another network without understanding what paths may be reflected between the RRs and how CLUSTER_LIST processing will behave.

Hierarchical reflection is a scale tool with a cost

Hierarchical RRs can reduce the number of sessions and align control-plane structure with geography or network tiers. But every reflection layer can further reduce available path information and adds another place where policy or best-path choice can become suboptimal. A hierarchy should therefore have a measurable purpose: peer-scale reduction, regional autonomy, failure containment, or platform limit management.

Failure testing for real redundancy

Test at least four events separately: loss of an RR BGP process, loss of the RR's IGP reachability, loss of a client-to-RR session, and loss of the shared transport beneath both RRs. A design that survives only the first event is not convincingly redundant.

Consistency versus independence

Redundant reflectors generally need consistent baseline policy, but complete operational coupling can create correlated failure. Configuration automation should enforce intended equivalence while allowing controlled staggered deployment, canary validation, and independent rollback. Reliability comes from both similarity of intent and separation of failure.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — route-reflector clusters and loop prevention
  • Cisco IOS XE 17.x — BGP route-reflector configuration
  • RFC 7964 — route oscillation considerations with route reflection/confederations

Day 38 — Route Reflectors: Clients, Non-Clients and Reflection Rules

1. Opening

Route reflection changes the iBGP propagation model so selected speakers can re-advertise iBGP-learned routes. That removes the universal full-mesh requirement, but it also means the engineer must understand client and non-client rules rather than treating the reflector as a transparent relay.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.



Route Reflectors: Clients, Non-Clients and Reflection Rules
Route Reflectors: Clients, Non-Clients and Reflection Rules


2. Concept and standards behavior

RFC 4456 defines route reflectors, clients, ORIGINATOR_ID and CLUSTER_LIST. In simplified terms, a route learned from an RR client may be reflected to other clients and non-clients; a route learned from a non-client is reflected to clients, but not to other non-clients. The RR still runs the BGP decision process—it does not blindly flood every path.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Use RR1 as route reflector with R1, R2 and R3 as clients, plus R4 as a non-client iBGP peer. Originate one prefix at R1 and another at R4, then trace which speakers receive each path.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Route Reflectors: Clients, Non-Clients and Reflection Rules
Route Reflectors: Clients, Non-Clients and Reflection Rules

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 neighbor 192.0.2.11 remote-as 65010
 neighbor 192.0.2.12 remote-as 65010
 neighbor 192.0.2.13 remote-as 65010
 neighbor 192.0.2.14 remote-as 65010
 !
 address-family ipv4
  neighbor 192.0.2.11 route-reflector-client
  neighbor 192.0.2.12 route-reflector-client
  neighbor 192.0.2.13 route-reflector-client
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Convert R3 from ordinary non-client to RR client and predict which reflected routes become newly visible to it. Verify path detail and reflected attributes.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Misclassify two edge speakers as non-clients and assume the RR will reflect routes between them. Both sessions stay Established, yet route visibility is incomplete.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The root cause is a wrong RR relationship, not a failed TCP/BGP session. Correct the intended client/non-client role and verify reflected route attributes.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Route reflectors reduce adjacency requirements by changing propagation semantics. Document client groups explicitly; an Established session alone does not prove the route will be reflected.

Deep-dive engineering notes

Reflection is selective propagation, not flooding

A route reflector still chooses BGP paths. It does not automatically copy every received path to every client. This distinction becomes crucial when several exits advertise the same NLRI. A client may receive only the reflector's selected view unless a separate path-diversity mechanism is used.

The client/non-client rules should be written into the design documentation. A route learned from a client can be reflected to clients and non-clients. A route learned from a non-client is reflected to clients but not to other non-clients. An engineer who treats all iBGP sessions attached to an RR as equivalent can create a topology where all sessions are Established yet some internal destinations disappear.

ORIGINATOR_ID and CLUSTER_LIST are diagnostic evidence

ORIGINATOR_ID identifies the BGP speaker that originally advertised a reflected route inside the AS. CLUSTER_LIST records clusters traversed by the reflected path and supports loop prevention. These fields are not decoration; they help prove that a route has been reflected and can explain why an RR rejects a path that appears to have returned through its own cluster.

When troubleshooting, compare two paths for the same prefix and explicitly inspect these attributes. If a path disappears after a cluster design change, the CLUSTER_LIST is often more useful than repeated show bgp summary output.

Next-hop behavior remains a separate problem

Route reflection changes iBGP advertisement rules but does not magically fix next-hop reachability. If a client receives a reflected route whose NEXT_HOP is unreachable, the control plane may show the route while the forwarding result remains broken or the path may not be usable. Maintain IGP reachability to relevant next hops and apply next-hop-self only where the design actually requires it.

Placement principle

Place reflectors for control-plane stability and topology awareness, not merely where configuration is convenient. A route reflector should not be forced through a fragile access failure domain just because it is physically near clients. Dedicated control-plane nodes, redundant transport, and consistent policy often matter more than raw hop count.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — BGP Route Reflection
  • RFC 4271 — base BGP behavior
  • Cisco IOS XE 17.18.x BGP configuration guide

Featured Post

Day 41 — BGP Confederations: Sub-AS Design, External View and Migration

1. Opening Confederations are another way to scale BGP inside a large administrative domain. They divide the domain into member autonomous systems while presenting a single confederation identifier to external peers. They are powerful, but their operational model is more complex than simply 'using private ASNs inside.' The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture . 2. Concept and standards behavior RFC 5065 defines AS_CONFED_SEQUENCE and AS_CONFED_SET and how member-AS relationships are represented. Confederation external sessions have eBGP-like properties inside the confederation, while the confederation is presented externally as one AS. Modern guidance must also account for the fact that RFC 9774 prohibits new origination of AS_SET/AS_CONFED_SET in ordinary aggregation c...