10 ms·
"Out of Band" network management is not trivial
- walterbell 2y ago> hardened in-band management What would this look like in practice? Management interfaces like BSPs don't have a great security track record.
- stingraycharles 2y agoI can only assume it’s based on VLAN for security (and probably dedicated ports assigned to VLANs so regular ports are never able to access the VLAN), but other than that, I have a hard time envisioning in-band management that doesn’t lock you out when the network goes down. It would protect you against things like DDoS attacks, and you can even assign dedicated (prioritized) access for these management ports. It’s an economical decision I suppose.
- bigcat12345678 2y agoWho said it is trivial?... Edit: The article take a title and describe some straightforward technical and business investments to make oob management network work.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- gavindean90 2y agoI’m reminded of when an old AT&T building went on sale as a house, and one of its selling points was that you could get power from two different power companies if you wanted. This highlighted to me the level of redundancy required to take such things seriously. It probably cost the company a lot to hook up the wires, and I doubt the second power company paid anything for the hookup. Big Bell did it there, and I’m sure they did it everywhere else too. Edit: I bet it had diesel generators when it was in service with AT&T to boot.
- thakoppno 2y agohttps://www.realtor.com/realestateandhomes-detail/13229-Southview-Ln_Dallas_TX_75240_M97272-56068 https://www.realtor.com/realestateandhomes-detail/13229-Sout... Listing removed a couple weeks ago.
- woleium 2y agoCrypto Collective eh?
- Scoundreller 2y ago> Edit: I bet it had diesel generators when it was in service with AT&T to boot. That's where AT&T screwed up in Nashville when their DC got bombed. They relied on natural gas generators for their electrical backup. No diesel tank farm. Big fire = fire department shuts down natural gas as wide as deemed necessary and everything slowly dies as the UPS batteries die. They also didn't have roll-up generator electrical feed points, so they had to figure out how to wire those up once they could get access again, delaying recovery. https://old.reddit.com/r/sysadmin/comments/kk3j0m/nashville_bombing_occurred_outside_of_a_major_att/?rdt=64520 https://old.reddit.com/r/sysadmin/comments/kk3j0m/nashville_...
- m463 2y agoInteresting. I've seen some power outages in california, and noticed that comcast/xfinity had these generator trailers rolled up next to telephone poles, probably powering the low voltage network infrastructure below the power lines.
- yaantc 2y ago> I bet it had diesel generators when it was in service with AT&T to boot. 20 to 25 years ago I visited a telecom switch center in Paris, the one under the Tuileries garden next to the Louvre. They had a huge and empty diesel generators room. They had all been replaced by a small turbine (not sure it's the right English term), just the same as what's used to power an helicopter. It was in a relatively small soundproof box, with a special vent for the exhaust, kind of lost on the side of a huge underground room. As the guy in charge explained to us, it was much more compact and convenient. The big risk was in getting it started, this was the tricky part. Once started it was extremely reliable.
- synack 2y agoWith launch costs dropping, I wonder if there’s a market for a low bandwidth “ssh via satellite” service. Could use AWS Ground Station to connect to your VPC.
- rlt 2y agoWhy not Starlink? ~$100/month/site is pretty low cost.
- synack 2y agoIf this is for use during outages, I want to know exactly what network path is used, ideally with as few hops as possible. Starlink can’t guarantee that.
- yusyusyus 2y agowhy? from my pov, once i’ve bought the service from the provider, their job is to deliver however they can; not my business, not my problem. my problem is making sure my redundancy (if required) isnt fate sharing.
- jen729w 2y agoRight, and it’s not as if you’d own the wired line anyway. That’d be leased just the same way your Starlink connection would be.
- erincandescent 2y agoBecause in networking, if you buy two uplinks and don't check the paths they're taking, fate demands that the fiber seeking back hoe just took out that one duct it turns out both of your "redundant" lines go down
- yusyusyus 2y agoeven with KMZs supplied, this still happens. complications in some cases. but an IP product (like starlink), i dont see the same equivalence. at what point does fate sharing analysis end in such a scenario?
- kkfx 2y agoApart from Rogers et alike, the main OOB/LOM issue is that's mostly only very old iron very few know, finding people who knows and finding non-hacky homegrown and not much tested solutions it's damn hard.
- goatsi 2y agoWhen I see out of band management at remote locations (usually for a dedicated doctors network run by the health authority that gets deployed at offices and clinics) it's generally analog phone line -> modem -> console port. Dialup is more than enough if all you need to do is reset a router config. Not 100% out of band for a telco though, unless they made sure to use a competitors lines.
- no_carrier 2y agoHere in Australia, POTS lines have been completely decommissioned, UK will be switched off by end of 2025 and I'm assuming there's similar timelines in lots of other countries.
- vladvasiliu 2y agoThey're on the way out in France, too. New buildings don't get copper anymore, only fiber. However, as I understand it, at least for commercial use, the phone company provides some kind of box that has battery-backing so it can provide phone service for a certain duration in case of emergency.
- tonyarkles 2y agoThe tricky part with that is that, at least in Canada, the RJ11 ports on the ONT are generally VoIP. They provide the appropriate voltages for a conventional POTS phone to work but digitize & compress the audio and send it along to the Telco as SIP or whatever. That works fine for voice but you're probably going to have a hard time using a conventional POTS modem over that connection. I've never tested it and am honestly pretty curious to see how well/poorly it would work.
- vladvasiliu 2y agoI’m pretty sure that’s the case for France, too. However, I’m only familiar with the emergency phone call use case, for which voice is enough. I’m not familiar with any legal obligation to provide data service, so I guess that if you need that, it’s up to you to negotiate SLAs or have multiple providers.
- transcriptase 2y agoIt’s trivial when you have the resources that come from being one of Canada’s 3 telecom oligopoly members. Unfortunately the CRTC is run by former execs/management of Bell, Telus, and Rogers, and our anti-competition bureau doesn’t seem to understand their purpose when they consistently allow these 3 to buy up and any all small competitors that gain even a regional market share. Meanwhile their service is mediocre and overpriced, which they’ll chalk up to geographical challenges of operating in Canada while all offering the exact same plans at the exact same prices, buying sports teams, and paying a reliable dividend.
- Scoundreller 2y agoIt's worse than that: 2 of the 3 telecom oligopoly members share (most) of their entire wireless network, with one providing most towers in the West, and the other in the East. I'm sure those 2 compete very hard with each other with that level of co-dependency.
- knocknock 2y agoMy previous org OOB used a data only SIM card from a different service provider. Curious why that wouldn't be a good solution?
- solatic 2y ago1. The risk, when you use a competitor's service, of your competitor cutting off service, especially at an inopportune time (like your service undergoing a major disruption, where cutting off your OOBM would be kicking you while you are down, but such is business). 2. The risk that you and your competitor unknowingly share a common dependency, like utility lines; if the common dependency fails then both you and your OOBM are offline. The whole point of paying for and maintaining an OOBM is to manage and compensate for the risks of disruption to your main infrastructure. Why would you knowingly add risks you can't control for on top of a framework meant to help you manage risk? It misses the point of why you have the OOBM in the first place.
- tonyarkles 2y agoMaybe 10-15 years ago there was a local Rogers outage that would have had the #2 failure you're describing. From what I recall, SaskTel had a big bundle of about 3,000 twisted pairs running under a park. Some of those went to a SaskTel tower, some to SaskTel residential wireline customers and some of those went to a Rogers facility. Along comes a backhoe and slices through the entire bundle.
- pharos92 2y agoI disagree, Out of Band Network Management (OOBM) is extremely trivial to implement. Most companies however don't see the value of OOBM until they have a major fault. The setup costs can be high, and the ongoing operational costs of OOBM infrastructure and links is also significant. I've built dozens of OOBM networks using fibre and 4G with the likes of Opengear. In instances, often deploying OOBM ahead of infrastructure rollouts so hardware can be delivered to site directly from factory, rather than go through a staging environment which adds time, cost and complexity.
- 1992spacemovie 2y agoOOB for carriers is significantly more complex; especially when you may be the only realistic access option in certain locations. However, given the rise of Starlink I think it becomes closer to "trivial" when the math becomes $100/mo/location + some minimal always-on OOB infrastructure on prem + cloud. Even in heavy-monopoly situations, you can usually guarantee the Starlink to Internet path due to the traffic bypassing the transport carriers on the ground (bent pipe to LEO sat) and landing at IXPs/near telco houses which egress direct to transit carriers.
- godelmachine 2y agoWe have a major incident wherein our firewall was totally down last month. The director at the end suggested that we need to have RS232 cable for out of band communication for such eventualities in the future. Makes one realize the reliability of RS232 in today’s day and age.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- 1992spacemovie 2y agoThere is OOB for carriers and OOB for non-carriers. OOB for carriers is significantly more complex and resource intensive than OOB for non-carriers. This topic (OOB or to forgo) has been beat to death over the last 20 years in the operator circles; the responsible consensus is trying to shave a % off operating expenses by cheaping out on your OOB is wrong. That said it does shock me that one of the tier-1 carriers in Canada was this... ignorant? Did they never expect it to rain or something? Wild.
- ianpenney 2y agoHam radio. Meshtastic. Knowing your neighborhood.
- benreesman 2y ago[flagged]
- clhodapp 2y agoWestern Electric! The Western Electric -> Western Digital substitution seems to be common for whatever reason
- benreesman 2y ago+1 for the correction, Western Digital is in fact not the brand that most of this happened under. Thank you Sir or Madame for keeping me honest in my strident claims!
- deleted 2y ago[deleted]
- Animats 2y agoIn the entire history of the Bell System, no electromechanical exchange was ever down for more than 30 minutes for any reason other than a natural disaster. With one exception, a major fire in New York City. Three weeks of downtime for 170,000 phones for that.[1] The Bell System pulled in resources and people from all over the system to replace and rewire several floors of equipment and cabling. That record has not been maintained in the digital era. The long distance system did not originally need the Bedminster, NJ network control center to operate. Bedminster sent routing updates periodically to the regional centers, but they could fall back to static routing if necessary. There was, by design, no single point of failure. Not even close. That was a basic design criterion in telecom prior to electronic switching. The system was designed to have less capacity but still keep running if parts of it went down. [1] https://www.youtube.com/watch?v=f_AWAmGi-g8 https://www.youtube.com/watch?v=f_AWAmGi-g8
- benjojo12 2y agoThat electro mechanical system also switched significantly less calls than the digital counterparts! Most modern day telcos that I have seen still have multiple power/line cards/uplinks in place and designed for redundancy. However the new systems can also just do so much more and are so more flexible that they can be configured out of existence just as easily! Some of this as well is just poor software, on some of the big carrier grade routers you can configure many things but the combination of things that you can figure may also just cause things to not work correctly, or even worse pull down the entire chassis, I don't have immediate experience on how good the early 2000s software was, but I would take a guess and say that configurability/flexibility has had a serious cost on reliability of the network
- atoav 2y agoAnd part of the reason why it is software is because people keep saying it is "just" software. Unreliability is unreliability even of it comes through software and we ahould treat broken software as broken, not as "just a software error".
- 2y ago
- Scoundreller 2y agoOne thing that was fascinating about the Rogers outage was on the wireless side: because "just" the core was down, the towers were still up. So mobile phones would try to make a connection to the tower just enough to connect but not be able to do anything, like call 9-1-1 without trying to fail-over to other mobile networks. Devices showed zero bars, but field test mode would show some handshake succeeding. (The CTO was roaming out-of-country, had zero bars and thought nothing of it... how they had no idea an enterprise-risking update was scheduled, we'll never know) Supposedly you could remove your SIM card (who carries that tool doohickey with them at all times?), or disable that eSIM, but you'd have to know that you can do that. Unsure if you'd still be at the mercy of Rogers being the most powerful signal and still failing to get your 9-1-1 call through. Rogers claimed to have no ability to power down the towers without a truck-roll (which is how another aspect where widespread OOB could have come in handy). Various stories of radio stations (which Rogers also owns a lot of) not being able to connect the studio to the transmitter, so some tech went with an mp3 player to play pre-recorded "evergreen" content. Others just went off-air. https://www.theregister.com/2022/07/25/canadian_isp_rogers_outage/ https://www.theregister.com/2022/07/25/canadian_isp_rogers_o...
- wannacboatmovie 2y ago> Supposedly you could remove your SIM card (who carries that tool doohickey with them at all times?) In sane handsets (ones where the battery is still removable), that tool was and still is a fingernail, which most have on their person. I believe the innovation of the need for a special SIM eject tool was bestowed upon us by the same fruit company that gave us floppy and optical drives without manual eject buttons over 30 years ago.
- rzzzt 2y agoYou could operate the ejection mechanism by hand both on optical and floppy disk drives with an uncurled paperclip (or a SIM card ejection tool were they to exist at that point in time). But I wouldn't ascribe the introduction of the motorized tray to the fruit company, it was the wordmark company: https://youtu.be/bujOWWTfzWQ https://youtu.be/bujOWWTfzWQ
- 2y ago
- jeffrallen 2y agoI love ChrisO so much, and it's funny but often he's talking about something I'm currently working on too. Thank you to Chris and to whoever posts his articles here.
- ralferoo 2y agoFrom TFA: > If your OOB network is your only way of managing things, you not only have to build a separate network, you have to make sure it is fully redundant, because otherwise you've created a single point of failure for (some) management. I'm not sure I necessarily agree with that. You can set up the network in such a way that you can route over the main network as a backup if your OOB network was down but the main network was up. Obviously, it's not quite as simple as sticking a patch cable between the two networks, but it can be close - you have a machine that's always on your OOB network, and it has an additional port that either configures itself over DHCP or has a hard-coded IP for the main net. But the important thing is that you never have that patched in, except for emergencies like your OOB network cable being severed but you still have access to the main network. If that does happen, you plug it in temporarily and use that machine as a proxy. There's no real reason for extra redundancy in the OOB, because if your main uplink is also severed, there's not really much you're going to be usefully configuring anyway!
- siebenmann 2y agoIn a lot of environments, you can at least choose to restrict what networks can be used to manage equipment; sometimes this is forced on you because the equipment only has a single port it will use for management or must be set to be managed over a single VLAN. Even when it's not forced, you may want to restrict management access as a security measure. If you can't reach a piece of equipment with restricted management access over your management-enabled network or networks, for instance because a fiber link in the middle has failed, you can't manage it (well, remotely, you can usually go there physically to reset or reconfigure it). You can cross-connect your out of band network to an in-band version of it (give it a VLAN tag, carry it across your regular infrastructure as a backup to its dedicated OOB links, have each location connect the VLAN to the dedicated OOB switches), but this gets increasingly complex as your OOB network itself gets complex (and you still need redundant OOB switches). As part of the complexity, this increases the chances an in-band failure affects your OOB network. For instance, if your OOB network is routed (because it's large), and you use your in-band routers as backup routing to the dedicated OOB routers, and you have an issue where the in-band routers start exporting a zillion routes to everyone they talk to (hi Rogers), you could crash your OOB network routers from the route flood. Oops. You can also do things like mis-configure switches and cross over VLANs, so that the VLAN'd version of your OOB network is suddenly being flooded with another VLAN's traffic. (I am the author of the original article.)
- kjellsbells 2y agoI worry that this misses the point a little. All the OOB in the world will not help you if you cannot reach the management entity (eg IP-enabled PSU, terminal server, etc). It is also insufficient to protect against second order thundering-herd-type problems (e.g.: you log in, stop a worker process, and upstream, traffic is directed away from the node to the others, and starts causing new problems). In telco operations, every MoP should have: an unambiguous linear sequence of steps, a procedure to verify that the desired result has been achieved, and a backout plan if things do go bad. This is drilled into you at every telco I ever worked at. Rogers' cardinal sin on the day of the outage was that they didn't have a backout plan at each step of the MoP. More structurally, networks have a dependency graph that you ignore at your peril. X depends on Y depend on Z, and so on. And yes, loops are quite possible! OOB management is an attempt to add new links to the graph that only get used in a crisis. These kind of pull-it-out-when-you-need-it solutions are fine, but have a tendency to fail just when you need them. For one, they don't get exercised enough, and two, they may have their own dependencies on the graph that are not realized until too late. So, what would this Internet rando prescribe? First order of business is to enumerate the dependency graph. I would wager that BGP, DNS, and the identity system are at or near the very top. Notice the deadly embrace of DNS and ID: if DNS is down, ID fails. Next, study the failure modes of the elements. In the Rogers outage, a lack of route filters crashed a core router. That's a vague word, "crashed". Are we talking core dumps and SEGVs? Are we talking response times that skyrocketed, leading to peers timing out? Rogers really need to understand that. Typically in telco networks when nodes get "congested" like this there are escape valves built into the control plane protocol, eg a response that says "please back off and retry in rand(300)". They need to have a conversation with Cisco/Juniper etc and their router gurus about this. Finally, the telco industry (or what's left of it) needs to do some introspection about the direction it is pulling vendors. For the last 15 years, telcos have been convinced that if only that can ingest some of that sweet, sweet cloud juice, their software costs will drop, they can slash operations costs, and watch the share price go brrr. Problem is, replacing legacy systems with ones cobbled together by vendors from a patchwork of kubernetes and prayers is guaranteed not to lead to the level of reliability that telcos and their regulators expect. If I'm a Rogers' operations manager and my network dies, I don't want to hear that some dude in India has to spend the next week picking through a service mesh and experimenting with multus to decide if turning if off and on again is gonna work.
- ChuckMcM 2y agoReminds me of a data center that said they had a backup connection and I pointed out that only one fiber was coming into the data center. They said, "Oh its on a different lambda[1]" :-) [1] Wave division multiplexing sends multiple signals over the same fiber by using different wavelengths for different channels. Each wavelength is sometimes referred to as a lambda.
- TwoNineFive 2y agoThe blog post is weird. "Rogers didn't even try, so OOB is hard." Also this sentence makes me question his IQ: "Some people have gone so far as to suggest that out of band network management is an obvious thing that everyone should have" Yes Chris, Rogers, the monopoly telco company of Canada, should have OOB network! They can afford it. Talking about the challenges of OOB is great, but the point the blog post is wrong and dumb. The report says "Rogers had a management network that relied on the Rogers IP core network". They had no OOB network. They didn't even try. This is a a symptom of Rogers status as a monopoly, negligence on the behalf of Rogers, and negligence on the behalf of the government who should have regulated OOB into existence. This is some serious clown car shit. One of the advantages that competitor networks provides is redundancy. Canada doesn't have that, so their networks will remain weak. This will probably happen again some day. Yes OOB is hard, but not even trying and then throwing up your hands and defending the negligent is stupid.