3 ms·
All great points- it sounds like you have a similar cultural awareness of the telco space. I'll reply to a few things that caught my brain's attention: > All t
by 1992spacemovie 2y ago
All great points- it sounds like you have a similar cultural awareness of the telco space. I'll reply to a few things that caught my brain's attention:
> All the OOB in the world will not help you if you cannot reach the management entity (eg IP-enabled PSU, terminal server, etc).
In _healthy_ OOB situations, all of the adjacent OOB infrastructure should be reachable, even if the entire core IP network is completely tanked. The only scenario where this would not apply in my eyes would be a power outage that whacks an entire site including the OOB gear. But in that scenario OOB doesn't help you.
> Next, study the failure modes of the elements. In the Rogers outage, a lack of route filters crashed a core router. That's a vague word, "crashed". Are we talking core dumps and SEGVs? Are we talking response times that skyrocketed, leading to peers timing out? Rogers really need to understand that. Typically in telco networks when nodes get "congested" like this there are escape valves built into the control plane protocol, eg a response that says "please back off and retry in rand(300)". They need to have a conversation with Cisco/Juniper etc and their router gurus about this.
Typically the "crash" is memory exhaustion due to incorrectly configured filtering between either routing protocols, or someone blasting a BGP peer with a large number of unexpected routes. As a former support engineer for BIGCO-ROUTER-COMPANY (either C.. or J..), I can't tell the number of times I've seen people melt down a large sized network due to either exceeding a defined prefix limit (limiting number of routes allowed), or accidentally nuking an ACL controlling route-redistribution, and either cratering all connectivity (no routes), or dump all routes unrestrictedly (no filter), with the latter resulting in memory exhaustion. Luckily, everyone these days working with big routers are culturally conditioned to do change-commit confirmation - if you make a change that blows the box up and isolates it, it will automatically revert the change after a defined period of time.
> Finally, the telco industry (or what's left of it) needs to do some introspection about the direction it is pulling vendors. For the last 15 years, telcos have been convinced that if only that can ingest some of that sweet, sweet cloud juice, their software costs will drop, they can slash operations costs, and watch the share price go brrr. Problem is, replacing legacy systems with ones cobbled together by vendors from a patchwork of kubernetes and prayers is guaranteed not to lead to the level of reliability that telcos and their regulators expect. If I'm a Rogers' operations manager and my network dies, I don't want to hear that some dude in India has to spend the next week picking through a service mesh and experimenting with multus to decide if turning if off and on again is gonna work.
I think your perception of the quality of a K8 telco stack is a bit off to be candid. They are not cobbling together random stacks from unvetted vendors/sources. Nearly every telco K8 stack these days is using an off the shelf K8 vendor, and off the shelf K8-compatible services on top, again from (reputable) vendors.
At the end of the day this was a failure of culture and management. The technology is a side conversation.