Day 41 — BGP Confederations: Sub-AS Design, External View and Migration

1. Opening

Confederations are another way to scale BGP inside a large administrative domain. They divide the domain into member autonomous systems while presenting a single confederation identifier to external peers. They are powerful, but their operational model is more complex than simply 'using private ASNs inside.'

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 5065 defines AS_CONFED_SEQUENCE and AS_CONFED_SET and how member-AS relationships are represented. Confederation external sessions have eBGP-like properties inside the confederation, while the confederation is presented externally as one AS. Modern guidance must also account for the fact that RFC 9774 prohibits new origination of AS_SET/AS_CONFED_SET in ordinary aggregation contexts; that does not erase RFC 5065's confederation architecture, but it reinforces the need to distinguish attributes and mechanisms carefully.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Create confederation identifier 65000 with member AS65010 and AS65020. Establish a session between members and an external peer to AS65100. Verify the path representation internally and the single confederation identity seen externally.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

A matching SVG/PNG diagram is stored in media/diagrams/.

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 bgp confederation identifier 65000
 bgp confederation peers 65020
 neighbor 192.0.2.20 remote-as 65020
!
! Member AS65020 uses the same confederation identifier
! and declares AS65010 as a confederation peer.

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Add a second member AS and trace a route across two member-AS boundaries to the external peer. Verify that external policy still sees the intended confederation AS.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Misconfigure a member relationship as ordinary eBGP by omitting the confederation peer declaration. Observe changes in path handling and session semantics.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The cause is inconsistent confederation membership configuration. Correct the member-AS and confederation-peer declarations on both sides, then re-check path attributes and external view.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Confederations are a control-plane architecture, not merely an ASN trick. Prefer the simplest design that meets scale, policy and failure-domain requirements; route reflection is more common and often easier to operate.

Deep-dive engineering notes

Confederation identity has two scopes

A confederation divides a large routing domain into member ASes for internal scaling and policy, while external peers see the confederation identifier rather than the internal member structure. This dual view is the core architectural idea. Engineers must know whether a policy is matching the member-AS representation used inside the confederation or the public/external AS representation seen outside.

Member-AS boundaries are neither ordinary iBGP nor ordinary Internet eBGP

Sessions between confederation members borrow eBGP-like behavior for scaling while preserving the confederation's single external identity. Treating them exactly like normal eBGP can lead to incorrect assumptions about AS-path representation and policy. Treating them like ordinary iBGP loses the reason confederations exist.

Route reflection versus confederations

Both reduce the universal iBGP full-mesh requirement, but their operational models differ. Route reflection changes propagation relationships inside one AS. Confederations introduce member-AS structure and explicit boundaries. A network with strong administrative regions and highly structured internal policy may benefit from confederation semantics, while many deployments favor route reflection because it is simpler and more common operationally.

Migration risk

A confederation migration changes session roles and path representation. Build coexistence carefully, define which routers move first, and test route policy that matches AS paths. A regex written for the pre-migration AS_PATH may behave differently when member-AS segments appear internally.

Oscillation awareness

RFC 7964 documents that route reflection or confederations can interact with certain topologies and policies to create persistent route oscillation. That is not a reason to avoid both mechanisms; it is a reason to validate policy ordering, MED behavior, topology and path visibility rather than assuming the scaling architecture is neutral.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 5065 — Autonomous System Confederations for BGP
  • RFC 4271 — BGP-4
  • RFC 7964 — persistent route oscillation considerations

Day 40 — Route-Reflector Path Hiding and Optimal Route Reflection

1. Opening

Route reflection can hide paths. A reflector generally selects a best path and reflects that view; clients may therefore never learn alternatives that would have been available in a full mesh. This can create suboptimal egress and can interact with convergence.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 4456's reduction of routing information creates the path-hiding problem. RFC 7911 ADD-PATH allows multiple paths for the same NLRI to be advertised when negotiated. RFC 9107 defines Optimal Route Reflection, in which the reflector can calculate client-appropriate optimal paths using client location/IGP perspective. These mechanisms solve different problems and must not be conflated.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Use two exits for 203.0.113.0/24. RR1 has lower IGP cost to Exit-A, while Client-B is physically closer to Exit-B. Without additional mechanisms, RR1's own best-path view can cause Client-B to receive a path that is not locally optimal.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Route-Reflector Path Hiding and Optimal Route Reflection
Route-Reflector Path Hiding and Optimal Route Reflection

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 ! Route-reflector baseline omitted for brevity.
 !
 address-family ipv4
  ! Verify platform support and exact ADD-PATH/ORR syntax
  ! before enabling either mechanism.
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

First reproduce path hiding with ordinary reflection. Then, in a capability-appropriate lab, compare the information made available by ADD-PATH with the client-specific selection objective of ORR.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Manipulate IGP cost so the reflector and a distant client have different closest exits. If the client never receives the alternative, troubleshooting the client's local BGP policy alone cannot fix the missing information.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The cause is information reduction at the reflector, not necessarily an incorrect client policy. Correct by redesigning reflector placement or using a supported path-diversity/optimal-reflection mechanism after validating platform behavior.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Path visibility is a design input. ADD-PATH increases path diversity; ORR changes which path is optimal for a client perspective. More paths can improve decisions but also increase control-plane state.

Deep-dive engineering notes

Path hiding begins with information reduction

In a full mesh, a speaker can potentially learn alternate paths directly from their originators. With route reflection, the reflector's decision can become the information boundary. If RR1 selects Exit-A and never advertises Exit-B's alternative to a client, the client cannot select Exit-B regardless of its own lower IGP cost to that exit.

This is why local troubleshooting at the client can be misleading. The client's BGP table may be internally consistent; it simply never received the alternative. The correct diagnostic question becomes “which paths were available at the reflector, which one did it select, and which paths did it advertise?”

ADD-PATH and ORR solve different dimensions

RFC 7911 ADD-PATH allows multiple paths for the same NLRI to be advertised with path identifiers when capability negotiation permits it. The client gains more path diversity and can make a richer local decision. The cost is additional control-plane state and update volume.

RFC 9107 Optimal Route Reflection aims at a different problem: the RR can select paths using a perspective appropriate to the client, such as the client's location in the IGP topology. ORR is therefore about selecting a client-optimal view; ADD-PATH is about advertising multiple paths. A design may use one, the other, neither, or platform-specific combinations.

Reflector placement can mitigate or amplify hiding

If the reflector's network location and IGP perspective closely match its clients, its selected path may already be acceptable for those clients. Centralizing a single RR far from diverse client populations increases the chance that the RR's hot-potato choice differs from the client's best exit.

Scale trade-off

Path diversity is not free. More paths can increase Adj-RIB-Out state, memory, update processing, and downstream selection work. Before enabling a feature globally, identify the prefixes or address families that benefit and establish measurable convergence or optimality objectives.

Verification strategy

Capture the same prefix at Exit-A, Exit-B, RR, and client. Record all available paths at the RR, the selected path, and the client's received paths. That four-point evidence proves whether the issue is path hiding, client policy, or next-hop resolution.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — route reflection
  • RFC 7911 — ADD-PATH
  • RFC 9107 — BGP Optimal Route Reflection (ORR)

RETICUX BGP Mastery — Day 39— Redundant and Hierarchical Route Reflectors

1. Opening

One route reflector solves session scale but can become a control-plane single point of failure. Adding a second reflector improves resilience only if clients actually peer to it, policies are consistent, and the reflectors do not share the same hidden dependency.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.

2. Concept and standards behavior

RFC 4456 permits multiple route reflectors and cluster designs. Redundancy must be evaluated end to end: reflector processes, IGP reachability, physical failure domains, client adjacency, policy distribution and management. Hierarchical reflection can reduce session fan-out further, but it also increases path-selection layers and troubleshooting complexity.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Build RR1 and RR2 in AS65010. Each client peers to both. Place the two RRs on separate simulated failure domains. Compare this with a false-redundant design where both RRs depend on the same transit interface or where half the clients peer to only one.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Redundant and Hierarchical Route Reflectors
Redundant and Hierarchical Route Reflectors


5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 neighbor 192.0.2.101 remote-as 65010
 neighbor 192.0.2.102 remote-as 65010
 !
 address-family ipv4
  neighbor 192.0.2.101 activate
  neighbor 192.0.2.102 activate
 exit-address-family
!
! On each RR, mark the client as route-reflector-client
! according to that reflector's configuration.

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Dual-home every client to RR1 and RR2, then fail RR1. The route should remain available through RR2 if the redundancy is real.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Keep RR2 operational but remove the client's session to it. Then fail RR1. This demonstrates that two reflector devices do not create redundancy unless the client relationship is also redundant.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The failure is architectural: the supposedly redundant path is not end-to-end. Correct the client peering and validate independent IGP/transport reachability to both reflectors.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Redundancy should survive a realistic shared-risk failure, not just a process restart. Hierarchy is justified by scale and geography—not by a desire to draw more layers in the diagram.

Deep-dive engineering notes

Two devices are not automatically redundant

A recurring design error is to count boxes instead of dependencies. RR1 and RR2 can be separate routers yet share the same rack power, aggregation link, IGP adjacency, management plane, configuration pipeline, or software defect. Redundancy should be assessed using shared-risk groups rather than device count.

Clients also need redundant adjacency. If Client-A peers only to RR1 and Client-B peers only to RR2, each client still has a single reflector dependency. A properly dual-homed client design lets either reflector continue distributing reachable paths after the other fails.

Cluster design deserves explicit intent

Cluster IDs participate in reflection loop prevention. Whether redundant RRs share a cluster ID or use distinct clusters depends on the intended reflection design and path behavior. Do not copy a cluster-id convention from another network without understanding what paths may be reflected between the RRs and how CLUSTER_LIST processing will behave.

Hierarchical reflection is a scale tool with a cost

Hierarchical RRs can reduce the number of sessions and align control-plane structure with geography or network tiers. But every reflection layer can further reduce available path information and adds another place where policy or best-path choice can become suboptimal. A hierarchy should therefore have a measurable purpose: peer-scale reduction, regional autonomy, failure containment, or platform limit management.

Failure testing for real redundancy

Test at least four events separately: loss of an RR BGP process, loss of the RR's IGP reachability, loss of a client-to-RR session, and loss of the shared transport beneath both RRs. A design that survives only the first event is not convincingly redundant.

Consistency versus independence

Redundant reflectors generally need consistent baseline policy, but complete operational coupling can create correlated failure. Configuration automation should enforce intended equivalence while allowing controlled staggered deployment, canary validation, and independent rollback. Reliability comes from both similarity of intent and separation of failure.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — route-reflector clusters and loop prevention
  • Cisco IOS XE 17.x — BGP route-reflector configuration
  • RFC 7964 — route oscillation considerations with route reflection/confederations

Day 38 — Route Reflectors: Clients, Non-Clients and Reflection Rules

1. Opening

Route reflection changes the iBGP propagation model so selected speakers can re-advertise iBGP-learned routes. That removes the universal full-mesh requirement, but it also means the engineer must understand client and non-client rules rather than treating the reflector as a transparent relay.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.



Route Reflectors: Clients, Non-Clients and Reflection Rules
Route Reflectors: Clients, Non-Clients and Reflection Rules


2. Concept and standards behavior

RFC 4456 defines route reflectors, clients, ORIGINATOR_ID and CLUSTER_LIST. In simplified terms, a route learned from an RR client may be reflected to other clients and non-clients; a route learned from a non-client is reflected to clients, but not to other non-clients. The RR still runs the BGP decision process—it does not blindly flood every path.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Use RR1 as route reflector with R1, R2 and R3 as clients, plus R4 as a non-client iBGP peer. Originate one prefix at R1 and another at R4, then trace which speakers receive each path.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Route Reflectors: Clients, Non-Clients and Reflection Rules
Route Reflectors: Clients, Non-Clients and Reflection Rules

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 neighbor 192.0.2.11 remote-as 65010
 neighbor 192.0.2.12 remote-as 65010
 neighbor 192.0.2.13 remote-as 65010
 neighbor 192.0.2.14 remote-as 65010
 !
 address-family ipv4
  neighbor 192.0.2.11 route-reflector-client
  neighbor 192.0.2.12 route-reflector-client
  neighbor 192.0.2.13 route-reflector-client
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Convert R3 from ordinary non-client to RR client and predict which reflected routes become newly visible to it. Verify path detail and reflected attributes.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Misclassify two edge speakers as non-clients and assume the RR will reflect routes between them. Both sessions stay Established, yet route visibility is incomplete.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The root cause is a wrong RR relationship, not a failed TCP/BGP session. Correct the intended client/non-client role and verify reflected route attributes.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Route reflectors reduce adjacency requirements by changing propagation semantics. Document client groups explicitly; an Established session alone does not prove the route will be reflected.

Deep-dive engineering notes

Reflection is selective propagation, not flooding

A route reflector still chooses BGP paths. It does not automatically copy every received path to every client. This distinction becomes crucial when several exits advertise the same NLRI. A client may receive only the reflector's selected view unless a separate path-diversity mechanism is used.

The client/non-client rules should be written into the design documentation. A route learned from a client can be reflected to clients and non-clients. A route learned from a non-client is reflected to clients but not to other non-clients. An engineer who treats all iBGP sessions attached to an RR as equivalent can create a topology where all sessions are Established yet some internal destinations disappear.

ORIGINATOR_ID and CLUSTER_LIST are diagnostic evidence

ORIGINATOR_ID identifies the BGP speaker that originally advertised a reflected route inside the AS. CLUSTER_LIST records clusters traversed by the reflected path and supports loop prevention. These fields are not decoration; they help prove that a route has been reflected and can explain why an RR rejects a path that appears to have returned through its own cluster.

When troubleshooting, compare two paths for the same prefix and explicitly inspect these attributes. If a path disappears after a cluster design change, the CLUSTER_LIST is often more useful than repeated show bgp summary output.

Next-hop behavior remains a separate problem

Route reflection changes iBGP advertisement rules but does not magically fix next-hop reachability. If a client receives a reflected route whose NEXT_HOP is unreachable, the control plane may show the route while the forwarding result remains broken or the path may not be usable. Maintain IGP reachability to relevant next hops and apply next-hop-self only where the design actually requires it.

Placement principle

Place reflectors for control-plane stability and topology awareness, not merely where configuration is convenient. A route reflector should not be forced through a fragile access failure domain just because it is physically near clients. Dedicated control-plane nodes, redundant transport, and consistent policy often matter more than raw hop count.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4456 — BGP Route Reflection
  • RFC 4271 — base BGP behavior
  • Cisco IOS XE 17.18.x BGP configuration guide

Day 37 — Why iBGP Full Mesh Does Not Scale

1. Opening

iBGP does not automatically relay every route learned from another iBGP peer. That behavior prevents simple internal loops but creates a scaling consequence: without another mechanism, every iBGP speaker that must share routes with every other speaker needs direct iBGP adjacency.

The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture.


Why iBGP Full Mesh Does Not Scale
Why iBGP Full Mesh Does Not Scale


2. Concept and standards behavior

In a plain iBGP design, routes learned from one iBGP peer are not advertised to another iBGP peer. RFC 4456 was created specifically because the resulting full mesh becomes operationally expensive as the number of speakers grows. With n routers, a full mesh requires n(n-1)/2 sessions. Ten speakers require 45 sessions; 50 require 1,225. Session count is not the only cost: every new router also increases configuration, policy and troubleshooting surfaces.

A recurring rule in this module is that configuration scale and routing-information scale are different problems. A design can be easy to configure but still create poor path visibility, slow convergence, or an oversized blast radius. Conversely, a topology can carry complete information but be operationally expensive to maintain.

3. Scenario

Build six routers inside AS65010. R1 originates 203.0.113.0/24. Initially configure only a chain of iBGP sessions R1-R2-R3-R4-R5-R6. The expected failure is that the route does not simply propagate across the chain. Then compare that with a true full mesh.

Documentation-safe addressing is used. The common lab AS is 65010. Internal loopbacks use 192.0.2.0/24 host routes; external test prefixes use 203.0.113.0/24 and 198.51.100.0/24 where needed. All fictional failures are lab-only.

Success criteria

  1. Every speaker learns the routes it is supposed to learn.
  2. The propagation rule can be explained before commands are applied.
  3. Loop-prevention attributes are visible where the feature uses them.
  4. A negative test proves that an invalid or unintended propagation does not occur.
  5. Rollback returns the topology to a known-good control-plane state.

4. Topology diagram

Why iBGP Full Mesh Does Not Scale
Why iBGP Full Mesh Does Not Scale

5. Prerequisites

  • IOS XE 17.18.x documentation baseline.
  • Stable IGP reachability among BGP loopbacks.
  • Explicit update-source and appropriate neighbor reachability for loopback-based sessions.
  • Authentication only if the lab image and design have been validated for it.
  • Independent management access before changing route-reflection/confederation policy.
  • Pre-change capture of BGP summary, selected paths and relevant neighbor state.

6. Baseline configuration

Topic-specific configuration excerpt — not a complete device configuration.

router bgp 65010
 bgp log-neighbor-changes
 neighbor 192.0.2.2 remote-as 65010
 neighbor 192.0.2.2 update-source Loopback0
 !
 address-family ipv4
  neighbor 192.0.2.2 activate
 exit-address-family

Do not silently transfer this syntax to IOS XR or NX-OS. Those platforms must use their own policy/configuration model.

7. Verification before modification

Use a combination of topology, session and route evidence:

show bgp ipv4 unicast summary
show bgp ipv4 unicast
show bgp ipv4 unicast 203.0.113.0/24
show bgp ipv4 unicast neighbors
show ip route

For route-reflection topics also inspect the path detail for ORIGINATOR_ID and CLUSTER_LIST where present. The absence or presence of a route must be explained by the propagation rule, not guessed from neighbor state alone.

8. Controlled modification

Add the missing direct iBGP adjacencies until the six-node topology forms a full mesh. Predict the required session count before configuring it and verify that R6 now learns R1's prefix.

Predict the exact peers whose Adj-RIB-In/Loc-RIB should change. Then apply one control-plane change, re-check path detail, and verify that no unrelated peer loses reachability.

9. Fault injection

Illustrative lab — not a real incident.

Remove one direct session that is the only way a particular speaker receives R1's path while keeping IGP reachability intact. The session topology, not IP reachability, becomes the fault.

Capture the state before and after the fault. A useful fault must change one causal variable only.

10. Step-by-step troubleshooting

  1. Confirm IGP reachability among BGP endpoints.
  2. Confirm TCP/BGP sessions are Established.
  3. Identify where the route originated.
  4. Trace the route one BGP hop at a time.
  5. Determine whether the receiving neighbor is eBGP, ordinary iBGP, RR client, RR non-client, or confederation peer.
  6. Inspect best-path state before assuming propagation is broken.
  7. Inspect ORIGINATOR_ID/CLUSTER_LIST for reflected routes.
  8. Inspect policy and next-hop reachability.
  9. Verify whether an alternative path was never learned, learned but not selected, or selected but not advertised.
  10. Apply the smallest proven correction.
  11. Re-run the same positive and negative tests.
  12. Observe stability before closing the change.

11. Root cause and correction

The proven cause is incomplete iBGP topology combined with the rule that iBGP-learned routes are not normally re-advertised to other iBGP peers. Correct by building the intended full mesh or introducing an explicit scaling architecture such as route reflection/confederations.

The correction must address the architectural cause rather than adding random neighbor statements until the route appears.

12. Post-fix verification

Verify:

  • all intended sessions remain Established;
  • the expected prefix is learned by the intended speakers;
  • reflected attributes are correct where applicable;
  • next hop remains reachable;
  • no routing loop is created;
  • the negative propagation test still passes.

13. Rollback

Restore the previous neighbor relationship or policy, not an ad-hoc alternative. If the change affects route reflection, preserve enough connectivity that clients do not become isolated during rollback. Roll back for unexpected loss of route visibility, oscillation, unintended path change, or evidence that the design created a larger failure domain than approved.

14. Production lessons

Use the full mesh as a conceptual baseline, not as an automatic production recommendation. Count sessions, policy attachment points, change burden and failure domains before deciding how to scale.

Deep-dive engineering notes

Why the session formula matters operationally

The mathematical growth of a full mesh is easy to state but the operational consequence is more important. Every new iBGP speaker can require another neighbor relationship on every existing speaker. That means another authentication relationship if authentication is used, another set of address-family activation decisions, another point where inbound/outbound policy can diverge, another state machine to monitor, and another adjacency to account for during maintenance. The design burden therefore grows faster than the router count.

A full mesh can still be reasonable in a small, stable control-plane core. The lesson is not “full mesh is bad”; the lesson is that it is a scaling baseline whose costs must be consciously accepted. A five-router control plane and a fifty-router control plane are different engineering problems even if both fit in memory.

Why an iBGP chain fails even when every IP hop works

An engineer can ping R1 from R6 and still have missing BGP routes. That distinction is important: IGP reachability proves the TCP endpoints can potentially communicate; it does not override the iBGP advertisement rule. In the chain lab, R2 learning R1's prefix over iBGP does not grant R2 permission to advertise that same path onward to R3 as ordinary iBGP. This is why adding static routes, changing interface costs, or repeatedly clearing sessions does not solve the architectural fault.

A strong troubleshooting workflow therefore records the route at each speaker. If R2 has the prefix and R3 does not, the next question is not “is R3 reachable?” but “what relationship allows R2 to propagate this path to R3?” That question leads naturally to route reflection or confederations.

Scaling metrics worth recording

For a production design review, record at least: number of BGP speakers, expected sessions, number of address families, number of policy variants, number of route sources, route churn, expected convergence objective, and operational ownership. Session count alone can underestimate complexity if each session carries several AFI/SAFI families and unique policy.

Failure-domain implication

A full mesh distributes dependency: there is no single route reflector whose loss removes every reflected path. But it also distributes change surface across every speaker. Route reflection centralizes some control-plane functions and therefore trades adjacency scale for new architectural dependencies. That trade-off is the bridge to Day 38.

15. Knowledge check

  1. Which problem in this post is a control-plane topology problem rather than a command-syntax problem?
  2. What evidence distinguishes “route was never learned” from “route was learned but not selected”?
  3. What negative test proves the scaling feature has not introduced unintended propagation?

Answers

  1. The relationship among BGP speakers and the rules governing propagation.
  2. Per-prefix BGP path detail and neighbor route evidence.
  3. Verify a route that should remain hidden/unadvertised is absent from the relevant neighbor's learned/advertised state.

16. Sources

  • RFC 4271 — BGP-4 base specification and iBGP behavior
  • RFC 4456 — Route Reflection, motivation for avoiding iBGP full mesh
  • Cisco IOS XE 17.18.x BGP configuration guide

Day 17 — BGP ORIGIN: IGP, EGP and Incomplete Without the Common Misconceptions

Learning objective

Interpret the ORIGIN attribute correctly, explain the preference order IGP < EGP < incomplete in Cisco best-path selection, and troubleshoot a policy that changes ORIGIN without confusing it with route ownership or the IGP protocol running in the network.


BGP ORIGIN: IGP, EGP and Incomplete Without the Common Misconceptions
BGP ORIGIN: IGP, EGP and Incomplete Without the Common Misconceptions



1. Opening — the word “origin” causes more confusion than the attribute

BGP's ORIGIN attribute has three values:

  • IGP
  • EGP
  • INCOMPLETE

Cisco commonly displays them as:

  • i
  • e
  • ?

The names are historical. They do not mean “this prefix is currently carried by OSPF,” “this prefix belongs to the AS shown here,” or “BGP does not know the AS that owns it.”

ORIGIN indicates how the route information was introduced into BGP according to the attribute semantics. In Cisco configurations, a prefix originated with a BGP network statement commonly appears with ORIGIN IGP, while redistribution commonly creates ORIGIN incomplete unless policy changes it.


2. ORIGIN classification

RFC 4271 defines ORIGIN as a well-known mandatory attribute with three values.

IGP

Indicates the NLRI was learned through an interior routing mechanism in the original BGP semantics. In modern Cisco operations, network-originated routes are typically displayed with i.

EGP

Represents the historical Exterior Gateway Protocol origin value. The old EGP protocol is obsolete in modern networks, but the ORIGIN value remains defined.

INCOMPLETE

Indicates the origin cannot be represented as IGP or EGP under the attribute semantics. Redistributed routes commonly appear as ? on Cisco.


3. Best-path preference

After earlier Cisco selection criteria such as Weight, LOCAL_PREF, local origination and AS_PATH length, Cisco prefers the route with the lowest ORIGIN type in this order:

IGP (i)  preferred over  EGP (e)  preferred over  incomplete (?)

Do not translate that into “IGP routes are always preferred over redistributed routes.” The comparison applies at this stage only if all higher-priority criteria have tied and the routes are eligible.


4. ORIGIN is not the same as ORIGINATOR_ID

These are completely different attributes.

  • ORIGIN is the base BGP attribute described here.
  • ORIGINATOR_ID is a route-reflection attribute used to prevent reflection loops.

The route-reflector module will treat ORIGINATOR_ID separately.


5. Scenario — Delta Origin-Code Tie Break

Illustrative lab — not a real incident.

Topology

           Service S AS65100
              /          \
             /            \
       P1 AS65010      P2 AS65020
             \            /
              \          /
             Enterprise E
                AS65050

S originates 192.0.2.128/25 and both providers carry it to E.

P1 advertises the path without changing ORIGIN, so E sees an IGP-origin code.

P2 uses an outbound route map toward E:

set origin incomplete

The AS_PATH lengths are intentionally equal:

  • P1: 65010 65100 i
  • P2: 65020 65100 ?

If earlier attributes tie, E should prefer the IGP-origin path via P1.


6. Prerequisites

  • Four-router topology similar to Day 16.
  • Equal Weight and LOCAL_PREF at E.
  • Equal AS_PATH length.
  • Reachable next hops.
  • No MED difference used to decide earlier/later unexpectedly.

The lab isolates ORIGIN by controlling other variables.


7. Topic-specific configuration excerpt

P2 policy toward E

route-map MARK-INCOMPLETE permit 10
 set origin incomplete
!
router bgp 65020
 address-family ipv4 unicast
  neighbor 203.0.113.6 route-map MARK-INCOMPLETE out

P1 has no ORIGIN-changing policy toward E.

The full topology configuration is stored in the companion lab directory.


8. Verification before modification

At E:

show ip bgp 192.0.2.128
show ip route 192.0.2.128 255.255.255.128

Identify the Path field and origin code at the end of each path.

Expected conceptual view:

65010 65100 i
65020 65100 ?

The exact formatting depends on the IOS XE image; do not fabricate a full command output.


9. Controlled modification

Change P2's route map from incomplete to IGP for the lab comparison:

route-map MARK-INCOMPLETE permit 10
 set origin igp

After a safe outbound refresh, both paths should have equal ORIGIN. E then proceeds to later best-path criteria.

This demonstrates that ORIGIN is one comparison step, not an absolute winner independent of the rest of the algorithm.


10. Fault injection

Illustrative lab — not a real incident.

Lab name

Delta Origin Rewrite

Fault

A broad outbound route map on P1 accidentally marks all advertisements as incomplete:

route-map BROAD-ORIGIN permit 10
 set origin incomplete

If P2 continues advertising IGP origin and earlier attributes tie, E can switch to P2.

Why this is subtle

The prefix, AS_PATH and next hop all look legitimate. The only policy difference is the one-character origin code at the end of the path display.


11. Troubleshooting sequence

  1. Confirm the candidate paths and best path.
  2. Check Weight and LOCAL_PREF.
  3. Compare AS_PATH length.
  4. Read the ORIGIN code at the end of each path.
  5. Determine whether the value is natural from route origination or modified by policy.
  6. Inspect set origin route-map clauses.
  7. Confirm the route-map direction and scope.
  8. Do not confuse ORIGIN with route ownership or ORIGINATOR_ID.
  9. Correct the smallest unintended rewrite.
  10. Refresh safely and confirm the path-selection result.

12. Root cause and correction

Root cause

An outbound policy rewrote ORIGIN to incomplete on the path that was intended to remain preferred.

Correction

Remove the unnecessary set origin statement or constrain it to the prefixes that genuinely require that policy.

Why it works

With the artificial ORIGIN disadvantage removed, the route can tie or win according to the remaining best-path criteria.


13. Post-fix verification

show ip bgp 192.0.2.128
show route-map
show running-config | section route-map

If the best path changes, verify the RIB and data plane as well:

show ip route 192.0.2.128 255.255.255.128

14. Rollback

Remove the route-map from the neighbor or restore the previous set origin value, depending on the intended policy.

Example:

router bgp 65010
 address-family ipv4 unicast
  no neighbor <E-peer> route-map BROAD-ORIGIN out

Preserve the original route display so the origin-code difference remains documented.


15. Production lessons

Terminology lesson

ORIGIN is historical BGP metadata; it is not a statement of legal prefix ownership or current IGP reachability.

Selection lesson

IGP origin is preferred over EGP, which is preferred over incomplete—but only after earlier criteria tie.

Troubleshooting lesson

The tiny i, e or ? at the end of a path can explain a best-path decision when more obvious attributes are equal.

Policy lesson

Avoid rewriting ORIGIN without a clear, documented reason. LOCAL_PREF and communities are often clearer tools for policy expression.

Documentation lesson

Never confuse ORIGIN with ORIGINATOR_ID; their names overlap but their purposes do not.


15A. Four common ORIGIN misconceptions

Misconception 1 — i means OSPF or IS-IS

It does not. The ORIGIN code is BGP metadata. A route can display i even when no OSPF adjacency exists anywhere near the originating router.

Misconception 2 — ? means the prefix owner is unknown

It does not. ? represents ORIGIN incomplete. Ownership and authorization are separate questions handled through routing policy, registry data and mechanisms such as RPKI—not the ORIGIN code.

Misconception 3 — ORIGIN tells you where the route entered your AS

It does not. For that operational question, inspect the BGP peer, next hop, AS_PATH, communities and route-policy evidence.

Misconception 4 — ORIGIN should be rewritten to make a path primary

It can influence best-path selection, but it is usually a poor first choice for expressing enterprise policy because LOCAL_PREF is clearer and evaluated earlier. Rewriting ORIGIN can also make troubleshooting harder by obscuring how the route was originally introduced.


15B. When set origin is useful in a lab

set origin is valuable for education because it isolates the ORIGIN comparison without requiring multiple route-injection mechanisms. In production, use it only when a documented interconnection or migration design calls for it.

Before deploying a route map that changes ORIGIN, record:

  • which prefixes match;
  • whether the route map is inbound or outbound;
  • whether other route-map clauses continue processing;
  • the intended best-path effect;
  • which routers can observe the rewritten value;
  • how the change will be rolled back.

That discipline prevents a one-line policy change from becoming an invisible tie-breaker across a large prefix set.


15C. Observability: how to prove ORIGIN actually caused the decision

When investigating a production path-selection question, do not stop at seeing i on the best path and ? on an alternate path. Prove that all earlier criteria tied.

A defensible evidence chain is:

1. Both paths valid and next hops reachable
2. Weight equal
3. LOCAL_PREF equal
4. Neither path wins on local origination
5. AS_PATH length equal
6. ORIGIN differs
7. Selected path matches the documented ORIGIN preference

If any earlier item differs, ORIGIN may be visible but not causative.

This distinction matters for incident reports. Saying “the route won because of ORIGIN” is a technical claim that should be supported by the full comparison, not merely by noticing different origin codes after the fact.

For automated validation, collect structured path attributes before and after the change and compare only the fields relevant to the intended experiment. That approach will be developed further in the automation module.


15D. Why the EGP origin code still exists

The e value is easy to misread because modern engineers rarely operate the historical Exterior Gateway Protocol that gave the value its name. BGP retained the three-value ORIGIN field for protocol compatibility and path semantics even as EGP itself disappeared from ordinary production use.

That means a modern route displaying e should trigger a policy investigation rather than an assumption that an ancient EGP session is running somewhere. It may have been created by explicit policy, migrated configuration, imported routing information or a lab exercise.

The operational response is the same evidence-based method used throughout this series: identify the neighbor that supplied the path, inspect the full set of attributes, trace any route map that can rewrite ORIGIN, and compare the observed value against the intended policy. Historical names are not substitutes for current-state evidence.


16. Knowledge check

Q1

Does ORIGIN i mean the route is currently learned through OSPF?

Answer: No. It is the BGP ORIGIN attribute value. On Cisco, routes introduced with a BGP network statement commonly display i, regardless of which IGP may exist elsewhere.

Q2

Which ORIGIN value is preferred if all earlier criteria tie?

Answer: IGP is preferred over EGP, and EGP is preferred over incomplete.

Q3

Is ORIGIN the same as ORIGINATOR_ID?

Answer: No. ORIGIN is a base path attribute; ORIGINATOR_ID is a route-reflection loop-prevention attribute.


17. Sources

  1. RFC 4271 — *A Border Gateway Protocol 4 (BGP-4)* — ORIGIN attribute values and Decision Process.
  2. Cisco — *IP Routing Configuration Guide, Cisco IOS XE 17.18.x: Configuring BGP* — origin type in the Cisco best-path sequence.
  3. Cisco — *Connecting to a Service Provider Using External BGP, IOS XE 17.x* — route-map policy and BGP output origin codes.
  4. Cisco — route-map command references for set origin.

Accessed: 2026-08-11.


Day 13 — Cisco BGP Weight: Local-Only Best-Path Policy Without AS-Wide Propagation

Learning objective

Use Cisco Weight to prefer one path on a single IOS XE router, prove that the decision remains local to that router, and distinguish Weight from standards-based attributes such as LOCAL_PREF.


Cisco BGP Weight: Local-Only Best-Path Policy Without AS-Wide Propagation
Cisco BGP Weight: Local-Only Best-Path Policy Without AS-Wide Propagation



1. Opening — a powerful knob with a deliberately small scope

An enterprise edge router receives the same destination from two external peers. Both routes are valid. Both have reachable next hops. Their standard BGP attributes are otherwise equal enough that the router reaches late tie-breakers.

The operator wants only this router to prefer ISP-A. Other routers in the autonomous system must not inherit that preference.

Cisco Weight is designed for exactly that kind of local decision.

Weight is evaluated before LOCAL_PREF in Cisco's best-path sequence. The higher value wins. The important operational boundary is that Weight is local to the Cisco router: it is not a BGP path attribute defined by RFC 4271 and it is not advertised to BGP peers.

That makes Weight useful—and dangerous. It can create a preference that is invisible to the rest of the AS unless engineers deliberately inspect the router where it was applied.


2. Standards behavior versus Cisco behavior

RFC 4271 defines the standard BGP path attributes used by interoperable BGP speakers. Weight is not among them.

Cisco IOS and IOS XE add Weight as a local selection value. Cisco documentation places it at the beginning of the Cisco best-path algorithm: among otherwise eligible paths, the path with the highest Weight is preferred before the algorithm evaluates LOCAL_PREF and later attributes.

Two consequences follow:

  1. A Weight change can override a LOCAL_PREF difference on the same router.
  2. The Weight change does not travel in an UPDATE to another router.

Do not describe Weight as a transitive, non-transitive, well-known or optional BGP attribute. Those classifications apply to BGP path attributes; Weight is a Cisco-local selection property.


3. Common sources of Weight

Cisco IOS XE can assign Weight in more than one way.

Neighbor-level Weight

A neighbor-level setting applies a local Weight to routes learned from that peer:

neighbor 192.0.2.1 weight 250

This is simple but broad: every applicable route from the neighbor receives the configured value.

Route-map set weight

A route map can set Weight selectively based on prefix, AS path, community or another supported match:

route-map PREFER-SERVICE permit 10
 match ip address prefix-list SERVICE
 set weight 300

Cisco documentation notes that a Weight assigned by set weight can override Weight assigned with the neighbor-level command for matching routes.

Locally originated routes

Cisco commonly displays Weight 32768 for locally originated BGP routes. Received routes normally show Weight 0 unless policy changes it. Day 15 will isolate the separate “locally originated” best-path step rather than relying only on the default Weight difference.


4. Scope comparison: Weight versus LOCAL_PREF

| Property | Weight | LOCAL_PREF |
|---|---|---|
| Standard BGP attribute | No | Yes |
| Cisco-local selection value | Yes | No |
| Advertised to iBGP peers | No | Yes, within the AS |
| Advertised to ordinary eBGP peers | No | No |
| Higher value preferred | Yes | Yes |
| Best use | One-router preference | AS-wide exit policy |

If several routers must agree on the preferred exit, LOCAL_PREF is normally the more appropriate mechanism. If only one Cisco router should prefer a path, Weight can be intentionally narrow.


5. Scenario — Harbor Edge Local Preference

Illustrative lab — not a real incident.

Topology


Cisco BGP Weight: Local-Only Best-Path Policy Without AS-Wide Propagation
Cisco BGP Weight: Local-Only Best-Path Policy Without AS-Wide Propagation


        ISP-A R1                    ISP-B R2
         AS65010                     AS65020
     192.0.2.1/30              198.51.100.1/30
            \                      /
             \ eBGP          eBGP /
              \                  /
               EDGE-R3 AS65030
       192.0.2.2/30   198.51.100.2/30
                       |
                203.0.113.0/24
              learned from both peers

Both providers advertise 203.0.113.0/24. In the closed lab, both paths are deliberately constructed with equal higher-priority standard attributes so the Weight change is easy to observe.

Objective

  • Establish both eBGP sessions.
  • Verify both paths are eligible.
  • Apply Weight 250 to routes from ISP-A.
  • Confirm R3 prefers ISP-A locally.
  • Confirm no BGP UPDATE contains “Weight”.
  • Remove the setting and verify the original decision process returns.

6. Prerequisites

  • Three IOS XE-capable nodes.
  • IOS XE 17.18.x documentation baseline.
  • Direct eBGP from R3 to R1 and R2.
  • Both provider next hops reachable.
  • Console or out-of-band management.
  • A lab-only prefix, 203.0.113.0/24, originated by each provider for controlled comparison.

The dual origination is intentionally artificial for attribute training. Do not model real prefix ownership from it.


7. Baseline configuration

R1 — AS65010

hostname R1
interface GigabitEthernet0/0
 ip address 192.0.2.1 255.255.255.252
 no shutdown
!
ip route 203.0.113.0 255.255.255.0 Null0
!
router bgp 65010
 bgp router-id 10.10.10.1
 neighbor 192.0.2.2 remote-as 65030
 address-family ipv4 unicast
  network 203.0.113.0 mask 255.255.255.0
  neighbor 192.0.2.2 activate
 exit-address-family

R2 — AS65020

hostname R2
interface GigabitEthernet0/0
 ip address 198.51.100.1 255.255.255.252
 no shutdown
!
ip route 203.0.113.0 255.255.255.0 Null0
!
router bgp 65020
 bgp router-id 10.20.20.2
 neighbor 198.51.100.2 remote-as 65030
 address-family ipv4 unicast
  network 203.0.113.0 mask 255.255.255.0
  neighbor 198.51.100.2 activate
 exit-address-family

R3 — AS65030

hostname R3
interface GigabitEthernet0/0
 ip address 192.0.2.2 255.255.255.252
 no shutdown
interface GigabitEthernet0/1
 ip address 198.51.100.2 255.255.255.252
 no shutdown
!
router bgp 65030
 bgp router-id 10.30.30.3
 neighbor 192.0.2.1 remote-as 65010
 neighbor 198.51.100.1 remote-as 65020
 address-family ipv4 unicast
  neighbor 192.0.2.1 activate
  neighbor 198.51.100.1 activate
 exit-address-family

8. Verification before modification

On R3:

show ip bgp summary
show ip bgp 203.0.113.0
show ip route 203.0.113.0
show ip route 192.0.2.1
show ip route 198.51.100.1

Confirm:

  • both sessions are Established;
  • both paths to the prefix exist;
  • both next hops resolve;
  • received Weight is the unmodified default for both paths;
  • whichever path is best before the experiment is recorded as baseline evidence.

Do not invent which peer wins the late tie-breakers. The selected image and exact router IDs can influence the baseline result.


9. Controlled modification — prefer ISP-A locally

On R3:

configure terminal
router bgp 65030
 address-family ipv4 unicast
  neighbor 192.0.2.1 weight 250
 end

Use an inbound soft refresh or appropriate safe reprocessing method if the selected image requires routes to be re-evaluated for the new neighbor Weight.

Then verify:

show ip bgp 203.0.113.0
show ip route 203.0.113.0

Expected result

The path learned from 192.0.2.1 should display Weight 250 and become preferred if it remains otherwise eligible.

The peer does not receive an UPDATE containing Weight because Weight is not transmitted as a BGP path attribute.


10. Fault injection

Illustrative lab — not a real incident.

Lab name

Harbor Hidden Weight Override

Fault

An engineer applies Weight 400 to ISP-B while an existing AS-wide LOCAL_PREF policy intends ISP-A to be preferred.

router bgp 65030
 address-family ipv4 unicast
  neighbor 198.51.100.1 weight 400

Expected symptoms

  • R3 chooses ISP-B locally.
  • Other routers in AS65030 do not learn “Weight 400”.
  • A troubleshooting engineer who checks only route policy on another router may see no reason for R3's decision.

Lesson

A local implementation knob can override a standards-based policy on one router without creating an obvious AS-wide policy artifact.


11. Step-by-step troubleshooting

  1. Confirm the affected prefix.
  2. Verify both BGP paths are valid and next hops reachable.
  3. Inspect the full BGP path detail, including Weight and LOCAL_PREF.
  4. Determine whether Weight came from a neighbor command or route map.
  5. Check route-map precedence and matches.
  6. Compare R3's decision with another router in the AS.
  7. Remember that packet captures of BGP UPDATE messages will not show Weight.
  8. Identify whether the intended policy is router-local or AS-wide.
  9. Remove the smallest incorrect Weight policy.
  10. Re-evaluate the route and verify forwarding.

12. Root cause and correction

Root cause

The router-local Weight value took precedence over later selection attributes.

Correction

Remove or narrow the Weight policy if the intent is not local to that router. If the intended decision must be shared throughout the AS, implement an appropriately designed LOCAL_PREF policy instead.

Why the correction works

It removes the earlier Cisco-local selection override, allowing the path to be compared using the intended standard attributes and later tie-breakers.


13. Post-fix verification

show ip bgp 203.0.113.0
show running-config | section router bgp
show route-map
show ip route 203.0.113.0

Record:

  • Weight on each path;
  • selected best path;
  • next-hop resolution;
  • route installed in the RIB;
  • any change in traffic path if a forwarding endpoint is added.

14. Rollback

router bgp 65030
 address-family ipv4 unicast
  no neighbor 192.0.2.1 weight 250
  no neighbor 198.51.100.1 weight 400

If Weight was set by route map, rollback the route-map clause rather than only removing a neighbor command.

Preserve before/after show ip bgp evidence because the local-only nature of Weight can otherwise make the original cause hard to reconstruct.


15. Production lessons

Design lesson

Use Weight only when router-local policy is actually desired.

Operations lesson

A distributed troubleshooting workflow must include local path-selection state; an AS-wide policy review alone can miss Weight.

Change-management lesson

Because Weight is evaluated early, a small configuration change can redirect large traffic volumes immediately on that router.

Portability lesson

Weight is Cisco-specific. Designs intended to remain vendor-neutral should prefer standards-based policy mechanisms when possible.

Security lesson

Unexpected local preference overrides can defeat carefully designed containment or preferred-exit policy. Treat path-selection changes as high-impact routing changes.


16. Knowledge check

Q1

Why does another iBGP router not see the Weight value configured on R3?

Answer: Weight is a Cisco-local selection property, not a BGP path attribute carried in UPDATE messages.

Q2

Which path wins if one path has Weight 300 and another has higher LOCAL_PREF but Weight 0, assuming both are eligible?

Answer: On Cisco's documented selection order, the path with Weight 300 is evaluated as better before LOCAL_PREF is considered.

Q3

When is LOCAL_PREF usually a better tool than Weight?

Answer: When the preferred exit should be communicated consistently to other BGP speakers inside the same AS.


17. Sources

  1. RFC 4271 — *A Border Gateway Protocol 4 (BGP-4)* — standard BGP path attributes and Decision Process.
  2. Cisco — *IP Routing Configuration Guide, Cisco IOS XE 17.18.x: Configuring BGP* — Cisco path-selection ordering.
  3. Cisco — *Connecting to a Service Provider Using External BGP, IOS XE 17.x* — configuring neighbor Weight and policy examples.
  4. Cisco — *Cisco IOS IP Routing: BGP Command Reference* — neighbor weight syntax and behavior.

Accessed: 2026-08-11.


Day 10 — BGP Route Origination: Network Statements, Redistribution and Safe Export Boundaries

Learning objective

Originate prefixes deliberately using BGP network statements and policy-controlled redistribution, prove the underlying source route before advertisement, and recognize the blast radius created by unfiltered redistribution.



BGP Route Origination



1. Opening — BGP cannot advertise what you never intended to originate

Many routing incidents begin before best-path selection, communities or traffic engineering ever matter.

The route was simply originated incorrectly.

An engineer may configure a network statement for a prefix that is not actually present in the local routing table. Another may redistribute connected routes and unintentionally expose infrastructure links. A static route used only as an origination anchor may disappear when tracking changes. A route map may be removed while the redistribute command remains, turning a narrow export policy into a broad one.

Route origination is therefore a control boundary.

The core question is:

> Which local reachability is this BGP speaker authorized to transform into BGP NLRI?

Day 10 treats that question as an engineering and security problem, not merely a CLI exercise.


2. Standards behavior versus Cisco origination methods

RFC 4271 defines BGP's protocol behavior but does not standardize Cisco CLI such as network or redistribute connected.

Cisco IOS XE provides multiple ways to place local reachability into the BGP table. Common mechanisms include:

  • the BGP network command;
  • redistribution from another routing source;
  • aggregate-address for aggregation;
  • conditional route injection and specialized features in more advanced designs.

This post focuses on the first two. Aggregation is Day 11.

Always distinguish:

  • protocol route advertisement from RFC behavior;
  • local route origination mechanism from Cisco implementation behavior.

3. The BGP network statement is not an interface-enabling command

In several interior routing protocols, a network command historically influences interfaces or protocol participation. In BGP, the purpose is different.

Cisco documents the BGP network command as specifying a network local to the AS and adding it to the BGP routing table when the relevant local route conditions are met.

Operationally, engineers should verify the intended prefix exists in the local routing information used by the implementation before expecting it to be originated.

For a documentation lab, a common pattern is:

ip route 203.0.113.0 255.255.255.0 Null0

followed by:

router bgp 65010
 address-family ipv4 unicast
  network 203.0.113.0 mask 255.255.255.0

The Null0 route is an origination anchor, not proof of useful end-to-end service.


4. Why exact-prefix verification matters

Suppose an engineer intends to originate 203.0.113.0/24 but only has more-specific routes such as two /25s.

Do not assume BGP will synthesize the /24 simply because all addresses are covered by more-specific reachability. Route origination and aggregation are separate mechanisms.

Before changing BGP, verify:

show ip route 203.0.113.0 255.255.255.0

Then verify BGP:

show ip bgp 203.0.113.0 255.255.255.0

This distinction is fundamental to troubleshooting “network statement configured, route not advertised” incidents.


5. Redistribution is powerful because it changes the trust boundary

Redistribution tells BGP to import routes learned from another source or routing domain.

Examples include:

  • connected;
  • static;
  • OSPF;
  • IS-IS;
  • EIGRP;
  • other supported sources.

The risk is obvious: a broad route source can contain far more prefixes than you intend to export externally.

This is why production redistribution should normally be paired with explicit policy that defines the permitted prefixes and, where needed, the attributes applied to them.

A rule worth adopting:

> Never treat redistribute connected as a harmless shortcut on an Internet-facing BGP edge.

Connected routes may include transit links, management networks, infrastructure addressing and temporary interfaces.


6. Use policy to constrain redistribution

A safe teaching pattern is:

  1. define a prefix list containing only authorized routes;
  2. match it in a route map;
  3. attach the route map to redistribution;
  4. verify the resulting BGP table before advertising externally.

Example:

ip prefix-list BGP-ORIGIN-CONNECTED seq 10 permit 203.0.113.128/25
!
route-map REDIST-CONNECTED permit 10
 match ip address prefix-list BGP-ORIGIN-CONNECTED
!
router bgp 65010
 address-family ipv4 unicast
  redistribute connected route-map REDIST-CONNECTED

Full route-map engineering comes later. Here the policy exists to enforce an origination boundary.


7. Origin code is not the whole route-source story

Cisco BGP displays familiar origin codes such as:

  • i
  • e
  • ?

A route introduced with a BGP network statement is commonly associated with IGP origin in BGP terminology, while redistributed routes commonly appear with INCOMPLETE origin unless policy changes behavior.

Do not over-interpret these labels. i does not mean the prefix is literally learned through your current IGP. ? does not mean the route is necessarily broken. These are standardized BGP ORIGIN semantics and historical naming.

The actual source of local reachability must be verified from routing configuration and the RIB.


8. Scenario — Beacon Origination Boundary

Illustrative lab — not a real incident.

Objective

R1 originates one service prefix with a BGP network statement and one connected service prefix through policy-controlled redistribution. A second connected infrastructure prefix must never enter BGP.

Topology

BGP Route Origination


BGP Path-Attribute Taxonomy: Mandatory, Discretionary, Transitive and Non-Transitive
Service A: 203.0.113.0/25  -- static Null0 anchor
Service B: 203.0.113.128/25 -- loopback connected
Transit:   192.0.2.0/30     -- infrastructure, must not be originated

       R1 AS65010 ---------------- R2 AS65020
        192.0.2.1/30            192.0.2.2/30

Success criteria

R2 receives:

  • 203.0.113.0/25;
  • 203.0.113.128/25.

R2 must not receive 192.0.2.0/30 as an originated customer/service prefix from R1.


9. Prerequisites

  • Two IOS XE-capable nodes.
  • IOS XE 17.18.x documentation baseline.
  • One loopback on R1 for the connected service prefix.
  • One static Null0 route for the network-statement prefix.
  • Console/management access.

10. Baseline configuration

R1

hostname R1
!
interface GigabitEthernet0/0
 ip address 192.0.2.1 255.255.255.252
 no shutdown
!
interface Loopback100
 ip address 203.0.113.129 255.255.255.128
!
ip route 203.0.113.0 255.255.255.128 Null0
!
ip prefix-list BGP-ORIGIN-CONNECTED seq 10 permit 203.0.113.128/25
!
route-map REDIST-CONNECTED permit 10
 match ip address prefix-list BGP-ORIGIN-CONNECTED
!
router bgp 65010
 bgp router-id 10.10.10.1
 neighbor 192.0.2.2 remote-as 65020
 address-family ipv4 unicast
  network 203.0.113.0 mask 255.255.255.128
  redistribute connected route-map REDIST-CONNECTED
  neighbor 192.0.2.2 activate
 exit-address-family

R2

hostname R2
interface GigabitEthernet0/0
 ip address 192.0.2.2 255.255.255.252
 no shutdown
!
router bgp 65020
 bgp router-id 10.20.20.2
 neighbor 192.0.2.1 remote-as 65010
 address-family ipv4 unicast
  neighbor 192.0.2.1 activate
 exit-address-family

11. Verification before modification

R1 source-route checks

show ip route 203.0.113.0 255.255.255.128
show ip route 203.0.113.128 255.255.255.128
show ip route 192.0.2.0 255.255.255.252

R1 BGP checks

show ip bgp 203.0.113.0
show ip bgp 203.0.113.128
show ip bgp 192.0.2.0
show ip bgp neighbors 192.0.2.2 advertised-routes

R2 checks

show ip bgp
show ip bgp 203.0.113.0
show ip bgp 203.0.113.128

The infrastructure transit prefix must not appear as an intentionally originated BGP service route.


12. Controlled modification — remove the redistribution filter

This is intentionally risky and must be limited to the isolated lab.

On R1, change:

router bgp 65010
 address-family ipv4 unicast
  no redistribute connected route-map REDIST-CONNECTED
  redistribute connected

Expected risk

Additional connected prefixes may become eligible for redistribution into BGP, including the inter-router transit network.

The exact resulting advertisement set depends on platform behavior, route state and outbound policy. The purpose is to demonstrate why unfiltered redistribution expands the origination boundary.


13. Fault injection

Illustrative lab — not a real incident.

Lab name

Beacon Origination Boundary — Connected Leak

Fault

Replace policy-controlled connected redistribution with unfiltered redistribute connected.

Symptom

R2 receives a prefix that was never intended to be part of the service advertisement set.

Security implication

Internal infrastructure addressing can escape into an external routing relationship.


14. Troubleshooting workflow

  1. Identify the unexpected prefix on R2.
  2. Determine from which neighbor it was learned.
  3. Inspect the route's origin and AS_PATH, but do not assume the origin code reveals the exact local source.
  4. On R1, identify the source of the prefix in the RIB.
  5. Inspect BGP origination configuration.
  6. Check whether redistribution is filtered.
  7. Check whether the route map still exists and is still attached.
  8. Restore the narrow redistribution policy.
  9. Verify the unintended prefix is withdrawn.
  10. Confirm the two intended service prefixes remain advertised.

15. Root cause and correction

Root cause

The redistribution command was broadened from a route-map-controlled import to unrestricted connected-route redistribution.

Corrected configuration

configure terminal
router bgp 65010
 address-family ipv4 unicast
  no redistribute connected
  redistribute connected route-map REDIST-CONNECTED
 end

Why it works

The route map constrains connected routes imported into BGP to the prefix list explicitly authorized for service origination.


16. Post-fix verification

On R1:

show ip bgp
show ip bgp neighbors 192.0.2.2 advertised-routes

On R2:

show ip bgp

Confirm:

  • both intended /25 service prefixes remain;
  • the transit /30 is absent;
  • the eBGP session remains stable.

17. Rollback

The rollback is the same as the correction: remove unrestricted redistribution and restore the route-map-controlled command.

Before any production redistribution change, save:

  • current BGP table;
  • current advertised routes;
  • relevant RIB entries;
  • prefix-list/route-map configuration;
  • a candidate rollback configuration.

18. Production design patterns

Prefer explicit origination

For a small set of aggregate/service prefixes, explicit network statements backed by deliberate RIB anchors are often easier to audit than broad redistribution.

If redistributing, use allow-list policy

An outbound safety policy can add another boundary, but do not rely on one layer alone when origination can be constrained at the source.

Monitor advertised prefixes

Prefix-count monitoring is useful, but exact-prefix validation is stronger. One accidental /30 may be operationally significant without noticeably changing a large prefix count.

Separate reachability from ownership

The fact that a route exists in the RIB does not mean BGP is authorized to advertise it to the Internet or a partner.


19. Production lessons

Configuration lesson

A BGP network statement and redistribution are different origination mechanisms with different failure modes.

Troubleshooting lesson

Prove the source route first, then prove BGP origination, then prove neighbor advertisement.

Security lesson

Redistribution is an export boundary. Unfiltered redistribution can expose infrastructure or private routes.

Change-management lesson

Removing a route map without removing or correcting the associated redistribute command can broaden behavior unexpectedly. Cisco documentation explicitly warns about redistribution continuing when filtering is removed.

Monitoring lesson

Track both intended advertisements and forbidden-prefix absence.


20. Knowledge check

Q1

A network 203.0.113.0 mask 255.255.255.0 statement exists, but the route is not in BGP. What should you verify first?

Answer: Verify that the intended source prefix exists in the local routing information required by the Cisco implementation, then inspect the BGP table and policy.

Q2

Why is redistribute connected risky on an edge router?

Answer: It can import every eligible connected route into BGP, including transit, infrastructure or management prefixes that were never meant for external advertisement.

Q3

Does BGP origin code ? automatically mean the route is invalid?

Answer: No. It commonly reflects INCOMPLETE origin semantics, often associated with redistribution. Validity and policy acceptance are separate questions.


21. Sources

  1. RFC 4271 — *A Border Gateway Protocol 4 (BGP-4)* — protocol route advertisement and ORIGIN behavior.
  2. Cisco — *IP Routing Configuration Guide, Cisco IOS XE 17.x: Configuring a Basic BGP Network* — network, route origination, aggregation and redistribution guidance.
  3. Cisco — *Cisco IOS IP Routing: BGP Command Reference* — BGP verification commands.
  4. Cisco IOS XE documentation warning on redistribution CLI removal — removing filtering while leaving redistribution can produce unexpected redistribution.

Accessed: 2026-08-11.


Featured Post

Day 41 — BGP Confederations: Sub-AS Design, External View and Migration

1. Opening Confederations are another way to scale BGP inside a large administrative domain. They divide the domain into member autonomous systems while presenting a single confederation identifier to external peers. They are powerful, but their operational model is more complex than simply 'using private ASNs inside.' The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture . 2. Concept and standards behavior RFC 5065 defines AS_CONFED_SEQUENCE and AS_CONFED_SET and how member-AS relationships are represented. Confederation external sessions have eBGP-like properties inside the confederation, while the confederation is presented externally as one AS. Modern guidance must also account for the fact that RFC 9774 prohibits new origination of AS_SET/AS_CONFED_SET in ordinary aggregation c...