4 ms·
This article challenges partisans in the BGP vs OSPF debate but it's deeper than just pointing out that the wire protocol is quite superficial. Obviously, at sc
by vii 6y ago
This article challenges partisans in the BGP vs OSPF debate but it's deeper than just pointing out that the wire protocol is quite superficial. Obviously, at scale you have to control OSPF to avoid update floods. But BGP also needs to be controlled so bad entries aren't trusted and propagated; for example, https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-knocked-large-parts-of-the-internet-offline-today/ https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...
The deeper message in the article is linked https://elegantnetwork.github.io/posts/Network-Validation-with-Vagrant/ https://elegantnetwork.github.io/posts/Network-Validation-wi... and it's about simulating infrastructure. This is so important for large scale systems as complex aggregate behaviour easily emerges from myopic error handling strategies that only take into account the local view from each actor. This is true for any large scale system not just networks!
The wire protocol is the surface; what information is propagated is essential and what lies underneath. Simulation illuminates these depths. Whiteboard discussions are great for debugging and communicating specific cases - for example, aggregate responses to retry backoff rules that are locally sensible can cascade into catastrophes if many actors follow the same process. However, it's hard as human beings to find these cases without the help of tools. Tools are so important!
- tptacek 6y agoThe BGP problem you're linking to owes to it being run at Internet scale by diverse teams, not as an interior routing protocol.
- vii 6y agoYes. I linked to it because it touched on BGP optimisation techniques, and illustrates the dangers of automatically passing on each update. Is there a better link to a public postmortem that illustrates these problems in a datacenter context? When we were operating the Facebook Moonshot project we found that even internally in controlled datacenters, at a scale less than Amazon's, there is plenty of diversity
- jlgaddis 6y ago> Obviously, at scale you have to control OSPF to avoid update floods. But BGP also needs to be controlled so bad entries aren't trusted and propagated; Well, sure, but why would this be an issue inside your datacenter? As tptacek mentions, that has nothing to do with the choice of BGP vs. OSPF as your IGP.