3 ms·
Someone following the news closer please fact-check me, but AFAIK, the chain of events: - The Optus network (the second largest telco network in Australia) wen
by fredwu 3y ago
Someone following the news closer please fact-check me, but AFAIK, the chain of events:
- The Optus network (the second largest telco network in Australia) went down
- Optus executives initially couldn't coordinate anything because they were... all on Optus
- The network outage lasted ~14 hours
- The network outage effected triple-zero, the country's emergency number, because Optus was rebooting towers and causing phones to still connect to them so they can't fallback to another carrier to dial 000
- After the network is finally back up, Optus blamed "a 3rd party" that caused the outage
- The "3rd party" was then turned out to be Singtel - Optus' parent company
- Singtel issued a statement to basically say Optus was wrong
- Optus then issued another statement saying the outage was caused by them using default configuration files on some of their Cisco routers
- The Australian senate summoned the Optus CEO for a hearing
- Here we are, the CEO resigned
EDIT:
To add more context, just over a year ago (in Sept 2022), Optus, under the same CEO's leadership, had a massive data breach: https://en.wikipedia.org/wiki/2022_Optus_data_breach https://en.wikipedia.org/wiki/2022_Optus_data_breach
- vermilingua 3y agoOther salient points: - The CEO told the senate hearing that she now carried Telstra and Vodaphone SIM cards (maybe practical, terrible optics) - Transport for NSW is heavily reliant on Optus so public transport was heavily impacted (not entirely Optus' fault, but public outrage means someone had to be the scapegoat)
- tsujamin 3y agoThe sad thing about this happening to Optus is that it _could_ have happened to any of the big telcos, this sort of technical cascading fault, but because it happened to Optus on the back of the data breach all the focus is on Optus, not on the wider lessons that ought to be learned. Would’ve been great to come out of an event like this looking at whether catastrophic fault conditions like this exist elsewhere in our national infra, but it feels like all that’ll happen is Optus gets the shit kicked out of them while other providers count their blessings
- paranoidrobot 3y agoWhat annoys me isn't so much that they had this outage. It's that they took so long to restore. Carrier systems are supposed to be multiply fault tolerant. Even if you have to put warm bodies physically in front of them, you should have enough of those people around to get the core networks up and going again within an hour or so by rolling back to a known-good configuration. Even if they've somehow managed to brick the systems, they should have enough hot/cold spares, and the ability to call $VENDOR to hand-deliver new ones if need be.
- simfree 3y agoThe model for network delivery has changed. Networks like Rogers in Canada, Optus in Australia, and dish in the USA outsource the core knowledge needed to recover their network to vendors like Nokia and Cisco. The employees locally lack the knowledge and the access to restore the network without guidance from external vendors outside the country. From an operating cost perspective, this is the right choice, but for reliability and sovereignty it's terrible. Sadly, with RCS, 5G Standalone and other new technologies that require operating servers with leading edge software, operators repeatedly choose to outsource the entire stack to an external vendor like Nokia rather than replicate that knowledge locally.
- deleted 3y ago[deleted]
- FireBeyond 3y agoI have a FirstNet SIM for my phone (the first responder network). I've never experienced it, and it is only during designated events (i.e. a switch has to be flipped at the carrier, so might not have worked here depending on the outage), but while my phone is nominally on AT&T during said events a few things are meant[1] to happen: Voice calls should be prioritized over other network traffic. Data should be the same. And my phone should roam from AT&T to VZW (or even TMO). [1] having said that, I've not looked too closely, and it sounds like (unsurprisingly) that might not exactly be the reality...
- resolutebat 3y agoNetworking people will be shocked, SHOCKED, to hear that the root cause appears to have been a broken BGP push. Taking 14 hours to recover from that is all on Optus though. https://en.wikipedia.org/wiki/2023_Optus_outage https://en.wikipedia.org/wiki/2023_Optus_outage And the damage went well beyond spotty emergency calls: for example, if you run a small business that relies on credit card payments, you were fucked if your terminals were on the Optus network. The situation was so bad that prepaid SIM cards for Telstra (the main competitor) were selling out in much of the country.
- robocat 3y agoFrom Wikipedia: Causes: A Border Gateway Protocol (BGP) routing problem played a role in the outage. Public data from CloudFlare showed a spike in BGP route announcements from the Optus network around the time the outage occurred — over 940,000 announcements in an hour from a node that normally makes less than 3,000 announcements per hour — indicative of a BGP routing problem. [snip] committee describes the outage as a gradual event triggered by loss of connectivity between neighbouring computer networks. The report suggests that approximately 90 edge provider routers disconnected as an automated protective measure against routing update overload. The failures occurred following a software upgrade at a North American Singtel exchange that caused one of the routers to disconnect. This, in turn, triggered Optus's routers to rapidly update its own routing tables which triggered the shutdown due to the pre-configured default threshold limits set by Cisco Systems being exceeded. The tabled report and Singtel stressed that the software upgrade was not the cause of the fault
- dhx 3y agoOptus' official position on the events that they tabled to the Australian parliament are is at [1]. In summary: > "This unexpected overload of IP routing information occurred after a software upgrade at one of the Singtel internet exchanges (known as STiX) in North America, one of Optus’ international networks. During the upgrade, the Optus network received changes in routing information from an alternate Singtel peering router. These routing changes were propagated through multiple layers of our IP Core network. As a result, at around 4:05am (AEDT), the pre-set safety limits on a significant number of Optus network routers were exceeded. Although the software upgrade resulted in the change in routing information, it was not the cause of the incident." > "It is now understood that the outage occurred due to approximately 90 PE routers automatically self-isolating in order to protect themselves from an overload of IP routing information. These self-protection limits are default settings provided by the relevant global equipment vendor (Cisco)." > "Several hypotheses and paths to restoration were explored over the period up to 10.30am." And then in later statements: > "Nokia is our managed services partner for our network, and they were involved from the very beginning in managing the incident and recovering the network; their staff are based in India in two locations"[2] One of the key problems appears to be heavy reliance on outsourced Nokia staff in India, who seemingly would have been disconnected from Optus' systems in Australia. Then within Australia for local Optus staff, perhaps staff had Optus-provided mobile phones and couldn't be reached if the mobile phone network was down. At the minimum, you'd like to think that on-call operational staff exist near all PE routers and have multiple communication means such as mobile phones with other carriers, satellite phone, fixed Internet connectivity not provided by Optus. The total outage duration was 6.5hrs to diagnose the problem and a further 3.5 hours to get 98% of connectivity re-established. Resolving the problem once diagnosed required physical presence at 14 sites across Australia to reset 90 PE routers (as part of "100+ devices"). [1] https://www.aph.gov.au/DocumentStore.ashx?id=2ed95079-023d-49d5-87fd-d9029740629b&subId=750333 https://www.aph.gov.au/DocumentStore.ashx?id=2ed95079-023d-4... [2] https://www.itnews.com.au/news/optus-had-not-contemplated-an-outage-as-big-as-november-8-602478 https://www.itnews.com.au/news/optus-had-not-contemplated-an...