40 ms·
Update about the October 4th outage
- reilly3000 5y agoSo their actual deployment process is quite rigorous and should have a tight blast radius. After lots of emulated and canary testing, their deployments are phased out over weeks. I don't see how a bad push could have done what happened yesterday. I found a paper that describes the process in detail. See page 10-11: https://web.archive.org/web/20211005034928/https://research.fb.com/wp-content/uploads/2021/03/Running-BGP-in-Data-Centers-at-Scale_final.pdf https://web.archive.org/web/20211005034928/https://research.... Phase Specification P1 Small number of RSWs in a random DC P2 Small number of RSWs (> P1) in another random DC P3 Small fraction of switches in all tiers in DC serving web traffic P4 10% of switches across DCs (to account for site differences) P5 20% of switches across DCs P6 Global push to all switches We classify upgrades in two classes: disruptive and non-disruptive, depending on if the upgrade affects existing forwarding state on the switch. Most upgrades in the data center are non-disruptive (performance optimizations, integration with other systems, etc.). To minimize routing instabilities during non-disruptive upgrades, we use BGP graceful restart (GR) [8]. When a switch is being upgraded, GR ensures that its peers do not delete existing routes for a period of time during which the switch’s BGP agent/config is upgraded. The switch then comes up, re-establishes the sessions with its peers and re-advertises routes. Since the upgrade is non-disruptive, the peers’ forwarding state are unchanged. Without GR, the peers would think the switch is down, and withdraw routes through that switch, only to re-advertise them when the switch comes back up after the upgrade. Disruptive upgrades (e.g., changes in policy affecting existing switch forwarding state) would trigger new advertisements/withdrawals to switches, and BGP re-convergence would occur subsequently. During this period, production traffic could be dropped or take longer paths causing increased latencies. Thus, if the binary or configuration change is disruptive, we drain (§3) and upgrade the device without impacting production traffic. Draining a device entails moving production traffic away from the device and reducing effective capacity in the network. Thus, we pool disruptive changes and upgrade the drained device at once instead of draining the device for each individual upgrade. Push Phases. Our push plan comprises six phases P1-P6 performed sequentially to apply the upgrades to agent/config in production gradually. We describe the specification of the 6 phases in Table 4. In each phase, the push engine randomly selects a certain number of switches based on the phase’s specification. After selection, the push engine upgrades these switches and restarts BGP on these switches. Our 6 push phases are to progressively increase scope of deployment with the last phase being the global push to all switches. P1-P5 can be construed as extensive testing phases: P1 and P2 modify a small number of rack switches to start the push. P3 is our first major deployment phase to all tiers in the topology. We choose a single data center which serves web traffic because our web applications have provisions such as load balancing to mitigate failures. Thus, failures in P3 have less impact to our services. To assess if our upgrade is safe in more diverse settings, P4 and P5 upgrade a significant fraction of our switches across different data center regions which serve different kinds of traffic workloads. Even if catastrophic outages occur during P4 or P5, we would still be able to achieve high performance connectivity due to the in-built redundancy in the network topology and our backup path policies—switches running the stable BGP agent/config would re-converge quickly to reduce impact of the outage. Finally, in P6, we upgrade the rest of the switches in all data centers. Figure 7 shows the timeline of push releases over a 12 month period. We achieved 9 successful pushes of our BGP agent to production. On average, each push takes 2-3 weeks
- marcosfelt 5y agoIf they have such a rigorous release process, what could have caused all of the dns records to get wiped?
- 1970-01-01 5y ago"Are you sure you want to remove ALL routes to AS32934? Type YES to confirm." Hey what is our internal BGP called again? AS32934? "Yeah" "OOK."
- boyrom3 5y agoMore details : https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/ https://engineering.fb.com/2021/10/05/networking-traffic/out...
- runawaybottle 5y agoIt was interesting to visit the subreddits of random countries (eg /r/Mongolia) and see the top posts all asking if fb/Insta/WhatsApp being down was local or global. I got the impression this morning that it was only affecting NA and Europe, but it looks like it was totally global. The numbers must be staggering of the number of people trying to login.
- geerlingguy 5y agoGotta love how painfully vague this is. Sounds like a PR piece for investors, not an engineering blog piece.
- johnduhart 5y agoI think you need to re-adjust your expectations, it's not reasonable to have a fully fleshed out RCA blog post available within hours of incident resolution. Most other cloud providers take a few days for theirs.
- padolsey 5y agoI mean, not an RCA per se, but info more akin to cloudflare's blog post would be v welcome IMHO: https://blog.cloudflare.com/october-2021-facebook-outage/ https://blog.cloudflare.com/october-2021-facebook-outage/
- bawolff 5y agoBoth posts have essentially the same info - the fb one just didn't include an explainer on how the internet works.
- gizmo385 5y agoThe Cloudflare post includes graphs depicting how the issue looked downstream and a rough initial timeline for the incident. The Facebook post says basically nothing more than that there was a networking mistake. I hope we get something more substantial and informative than that over the next couple of weeks, but it doesn't seem (at least from my searching) that Facebook is in the business of publicly posting in-depth post-mortems for their outages, which I personally find unfortunate.
- ignoramous 5y agoCloudflare can scratch the surface of the issue it wouldn't matter, it is a content marketing piece after all. Facebook, otoh, needs to be thorough.
- metissec98 5y agoWell that doesn't say a whole lot... I know it is early but they could use a little more detail. Even if it is just a timeline.
- dugo 5y agoAround the turn of the century, in a network the size of Europe, we had OOB comms to the core routers via ISDN/POTS. We experimented with mobile phones in the racks as well, much to the chagrin of the old telco guys running the PoPs.
- coliveira 5y agoThe best course of action is to split FB into separate companies. It is already neatly divided between instagram, WU and legacy facebook. That would be the best for the government to avoid disruptions.
- colechristensen 5y agoOf all the reasons to break up big companies, protecting consumers from Instagram downtime is not one of them.
- sydthrowaway 5y agoAny FB throwaway know if someone got fired for this?
- yalok 5y agoThe Bootcamp training at FB explicitly mentions that such things are not a fire-able offense - the attitude is around learning - if you managed to bring everything down, let’s learn together how you managed to do this… :)
- sydthrowaway 5y agoI don't think this is the case. Wasn't TechLead fired for SEV events?
- oversighzed 5y agoWhere did you hear that from? He doesn’t even say that in his video, he says he was fired for having side income on YouTube.
- sydthrowaway 5y agoSorry, I was thinking about the engineer who got PIPd and committed suicide.
- phreeza 5y agoI think it would be extremely unusual and counterproductive for someone in the trenches to get fired about this, as it is clearly a failure of procedure that this was even possible and so hard to recover from. Large companies I am aware of have a no-blame postmortem culture around this stuff. There may be people suffering consequences at a higher level in the SRE division, though I doubt this will happen in a timeframe of hours after the outage.
- deleted 5y ago[deleted]
- wyldfire 5y agoMove fast and NO CARRIER
- dev_tty01 5y ago>We also have no evidence that user data was compromised as a result of this downtime. No, that just happens during uptime.
- Jugurtha 5y agoThe first thing people here thought of was that it was the gouvernement denying access to these websites as it usually does for a number of reasons.
- judge2020 5y agoIt was pretty quickly deemed a global phenomenon, so no comments on posts about it said that. Also, enough people here on HN know how to investigate dns and bgp to have found the problem within the first 30 minutes, first with DNS then the revelation that every BGP route associated with them was withdrawn.
- rvz 5y agoIt has been painfully admitted by the Facebook mafia that they know that they are the internet and farming the data of an entire civilisation; further evidence that this deep integration of their services needs to be broken up. After all the scandals, leaks, whistleblowers etc it would take more than a DNS record wipe to take down the Facebook mafia.
- 0xy 5y agoKnowing almost nothing about networking, isn't the way Facebook handles networking somewhat of a monolithic anti-pattern? Why is a single update responsible for taking out multiple services and why wouldn't each product or even each region within each product have their own routes, for resiliency which can then be used to rollout changes slower? By having a large centralized and monolithic system, aren't they guaranteeing that mistakes cause huge splash damage and don't separate concerns?
- tw04 5y agoQuite the opposite. Back in the day you would've had to login dozens if not hundreds of routers individually to push the change, and it likely would've been caught after screwing up the first one. This is the result of SDN (software defined networking) and being able to push a change globally from one command. I recall major ISP's screwing up their routing tables in the past but never globally on this level.
- kanbara 5y agohttps://www.techjuice.pk/country-blocked-youtube-globally/ https://www.techjuice.pk/country-blocked-youtube-globally/ https://www.bleepingcomputer.com/news/security/major-bgp-leak-disrupts-thousands-of-networks-globally/ https://www.bleepingcomputer.com/news/security/major-bgp-lea...
- toast0 5y agohttps://www.itproportal.com/news/misconfigured-centurylink-database-caused-global-internet-outage/ https://www.itproportal.com/news/misconfigured-centurylink-d... https://www.bleepingcomputer.com/news/technology/ibm-cloud-global-outage-caused-by-incorrect-bgp-routing/ https://www.bleepingcomputer.com/news/technology/ibm-cloud-g... (this one isn't clear, maybe BGP hijacking, and if so, not sure who the responsible party was) https://www.catchpoint.com/blog/vodafone-idea-bgp-leak https://www.catchpoint.com/blog/vodafone-idea-bgp-leak (not sure how major this one was) You can practically search ISP bgp outage and get news about the last couple times they screwed up BGP and caused a big problem. Or service BGP and get a 50/50 chance of the service screwing up BGP or an ISP/country hijacking their routes and causing a big problem. BGP is one of the best ways to break things at scale.
- gannon- 5y agoThis is a funny post to have suggested at the bottom of the article: https://engineering.fb.com/2021/08/09/connectivity/backbone-management/ https://engineering.fb.com/2021/08/09/connectivity/backbone-...
- itronitron 5y agoLooks like the 'Failure Generator' was brought online.
- andrewxdiamond 5y agoThis more or less confirms what we’ve heard, and I appreciate the speed, but it’s incredibly lame from a details point of view. Will a real postmortem follow? Or is this the best we are gonna get?
- badtux 5y agoHaving been on the team that issued postmortems before, I can tell you that we said as little as possible in as vague a way as possible while meeting our minimum legal requirements. Actual Facebook customers (i.e. those who pay money to Facebook) will get a slightly more detailed release. But the whole goal is to give as little information as possible while appearing to be open. As an engineer that makes me growl, but that's how it is in this litigous world -- don't want to give someone a reason to sue.
- laegooose 5y agoHow would you explain that AWS, GCE, Cloudflare, GitLab publish very detailed post-mortems?
- _nalply 5y agoI don't know but perhaps they excluded damages in their ToS?
- cik 5y agoIt's a marketing strategy. Their target customer segment is technical. FaceBooks and Twitter's for the most part, aren't.
- ignoramous 5y agoYeah, also a chance some eng from AWS/GCP/Azure leaks actual details if they lie or if public statements are inadequate.
- fragmede 5y ago
- paxys 5y agoIt was quite ironic that while every Facebook property was offline there was an immense amount of misinformation about the incident perpetuated across the internet (including right here on HN) which everyone just believed as fact.
- tomrod 5y agoLike what?
- runawaybottle 5y agoWe almost went down the ‘this is a subterfuge to delete whistleblower evidence’ rabbit hole.
- colechristensen 5y agoI saw a couple of people clearly guessing something along these lines but none of them seemed to be claiming that it was actually happening, more like “isn’t it convenient that…”
- pbhjpbhj 5y agoThe timing was uncanny. I still don't see a reason why it couldn't have been an intrusion/rogue employee? Like someone had access to a system to push router firmware updates or something?
- colechristensen 5y agoA rogue employee would have been very easy to detect and that employee would have known this. The core network infrastructure involved is extremely sensitive and is quite unlikely to be accessible in a break in. Also IIRC an employee was posting on Reddit saying the incident started shortly after a network update was posted this morning. When you know more about the tech and systems involved a mistake seems infinitely more likely than sabotage.
- stephenhuey 5y agoEven though the angle grinder story wasn’t accurate, it’d still be interesting to know what percentage of the time to fix the outage was spent on regaining physical access: https://mobile.twitter.com/mikeisaac/status/1445196576956162050?s=21 https://mobile.twitter.com/mikeisaac/status/1445196576956162...
- trthomps 5y agoReading this statement all I can think of is this scene https://www.youtube.com/watch?v=15HTd4Um1m4 https://www.youtube.com/watch?v=15HTd4Um1m4
- shahsyed 5y ago> configuration changes on the backbone routers that coordinate network traffic between our data centers caused issues This could be anything, potentially. I'm not very knowledgeable in computer networking, but this could be as trivial as an incorrect update to a DNS record, right?
- toast0 5y agoBackbone routers don't usually deal with hostnames or DNS. This is pretty much saying they done broke BGP. And it sounds like they're saying that they broke it in a way that prevented accessing their data centers from the PoPs, and we know from the long downtime that it prevented accessing the BGP configuration system from darn near anywhere. It happened to also kill the announcements for anycast DNS.
- dodobirdlord 5y agoThere seems to be nothing uncertain about the immediate cause of the issue - Facebook revoked all of their BGP routes, and all of their IP addresses couldn't receive packets until they were restored.
- toast0 5y agoThe didn't revoke all their routes, FWIW, just a lot of them (including the anycast DNS routes)
- pbhjpbhj 5y agoWhat I don't understand is why, when a route is revoked, if there is no other route announced the routing table gets updated? It seems like either it's a black hole or it still works and there was a BGP error (or the route works but the resources aren't present, so traffic would be dropped). What's the reason for designing the system to revoke routes when no new route is announced? It strikes me it's like DNS when you get a SERVFAIL, why not try the prior IP address. The similarity in the design here suggests there may be common reasoning??
- eyelidlessness 5y agoOne of the things they restored was annoying sounds in the app every time I tap anything. Who knew that was DNS related!
- eyelidlessness 5y agoI’ll take my downvotes but I’d be happy for anyone to explain why.
- 1970-01-01 5y agoTL;DR We YOLO'd our BGP experiment to prod. It failed. https://web.archive.org/web/20210626191032/https://engineering.fb.com/2021/05/13/data-center-engineering/bgp/ https://web.archive.org/web/20210626191032/https://engineeri...
- r00tanon 5y ago"Post hoc ergo propter hoc"
- r00tanon 5y agoYes. It is true. If you enter Facebook into Facebook. It will break the internet.
- advpetc 5y agoJust out of curiosity, does Facebook have a status page? Like http://status.twitter.com http://status.twitter.com?
- paxys 5y agohttps://status.fb.com/ https://status.fb.com/ Although it only covers their API and business apps, not the site itself.
- jcims 5y agoIt was also down during the outage.
- advpetc 5y agoThat’s a bit sad
- toast0 5y agoFor a status page to be actually independent, it needs to have all it's requirements hosted on other infrastructure. fb.com authoritative DNS is the same as facebook.com, so it's going down when (FB) DNS goes down (and DNS is going down when BGP is broken, apparently). It looks like the status page is hosted on CloudFront though, so it got part of the way. (Of course, the other question is if it was updatable / updated during the outage)
- can16358p 5y agoPardon me if it's a stupid question, but out of curiosity: Is there any way to keep DNS up in case BGP goes down for any reason? Like a fallback nameserver hosted elsewhere/not affected by Facebook's ASs? Is it technically impossible or did Facebook just assume something like yesterday would never happen and kept things simple instead of complicating things?
- 5y ago
- r00tanon 5y agoRemember, remember, the 4th of October.
- supermatt 5y agoSounds like they could do with some updates to their risk-driven backbone management strategy! https://engineering.fb.com/2021/08/09/connectivity/backbone-management/ https://engineering.fb.com/2021/08/09/connectivity/backbone-...
- dave333 5y agoI thought DARPA designed the internet to survive nuclear war - no single point of failure - clearly Facebook's network breaks that rule. They need a DNS of last resort that doesn't update fast.
- paxys 5y agoFar easier to make a system resilient to bombing than a bad configuration update
- Telluur 5y agoOn your single point of failure. That might be true, but certainly isn't these days. Networks have grown so large and complex that the only reasonable way of managing them is through SDN, and a small mistake in configuration might results in a cascading effect on the whole infrastructure. That's also true for the entire (western) internet. We've ended up with a centralized market where a few key players, e.g. cloud providers/CDNs/DNS (Amazon/Google/Microsoft/Akamai/Fastly/Cloudflare) can easily break large parts of the internet. See Akamai outage in July.
- vishesh92 5y ago> We also have no evidence that user data was compromised as a result of this downtime. I am not sure why they had to mention this specifically. This makes it sound like an external attack.
- KZerda 5y agoThere were rumors early in the downtime that it was the responsibility of various outside groups. Saying, "no, your data was not impacted" is pretty standard in light of those rumors, even if they weren't the main ones spreading around after the initial reports.
- tsimionescu 5y agoIt doesn't make it sound like an attack, it's standard boilerplate to dispel any worries. It's natural for anyone to wonder downtime -> data loss? , so it's natural to reassure people that it wasn't the case.
- rbrbr 5y ago“ We also have no evidence that user data was compromised as a result of this downtime.” Well that happened already. No worries.
- raverbashing 5y agoThe badge story only shows how people are looking for "efficiency" where it doesn't matter, with predictable results. The badge system should be local to the building. There are few actual reasons (sure, besides "efficiency") of why badge control should be centralized. Even less reasons for it to be a subdomain of fb. Another option would be to keep the system but make it failsafe (but it seems the newer generation doesn't know what that means). If the network goes down keep it at the last config. Badge validation should be offline first and added/removed ones should be broadcast periodically. This is the same issue with smartlocks times the number of employees. Do you really want to add another point of failure between yourself and your home?
- Sebb767 5y agoHaving the badge system work from a single point has a lot of advantages for a company like FB: HR can update info from everywhere (they might not be in the same office), you can immediately deny or block a card everywhere, you have an audit log etc.. They're not having this for fun. Akso, it's likely not on an fb subdomain, but something like office.security.fb-infra.com (example). It just happens to be that fb-infra.com is using the Facebook DNS server.
- raverbashing 5y agoSure, this fits under the "few actual reasons" but think for a moment: does it make sense that access to a building is controlled only (keyword here) through a centralized location somewhere? Some DB who knows where? With no fallback? You might need a break-glass account/badge somewhere. Sure, the angle-grinder works, but probably cost you 2h maybe? > it's likely not on an fb subdomain, but something like office.security.fb-infra.com Thanks, yeah, makes sense
- robinson-wall 5y agoIt's still possible to design a badge system on an independent network (think just switched within a building) which syncs a local copy of the authoritative ldap from the corp domain, so your badge readers stay working if the link to the corp domain goes away. It's just more expensive and another thing to maintain, and still doesn't account for _all_ failure modes (what if you sync really frequently and a bad change was made deleting all accounts?)
- imgabe 5y agoIt just occurred to me to wonder if Facebook has a Twitter account and if they used it to update people about the outage. It turns out they do, and they did, which makes sense. Boy, it must have been galling to have to use a competing communication network to tell people that your network is down. It looks like Zuckerberg doesn't have a personal Twitter though, nor does Jack Dorsey have a public Facebook page (or they're set not to show up in search).
- melvinmt 5y ago> It looks like Zuckerberg doesn't have a personal Twitter though He does: https://twitter.com/finkd https://twitter.com/finkd
- imgabe 5y agoAh, I saw that one, but it wasn't verified so I figured it was an imposter. It has only a handful of tweets from 2009 and 1 from 2012, but it could really be him, I suppose.
- dane-pgp 5y agoYeah, that's kinda sus.
- nostromo 5y agoHis LinkedIn photo used to be this really awkward laptop camera photo of roughly this face: (-_-) It was amazing. I’m sad he remove it.
- e9 5y agoI’m not sure they are competing though. They serve different purposes and co-exist pretty well together.
- jell 5y agoThey have an official account. https://twitter.com/Facebook/status/1445061804636479493 https://twitter.com/Facebook/status/1445061804636479493 hint: "some people".
- andy-x 5y agoSuch a BS. FB imagining that they are their own Internet but failing in a most miserable way because they need actual Internet to communicate.
- vilius 5y agoFacebook just can't admit it went like this https://www.youtube.com/watch?v=uRGljemfwUE https://www.youtube.com/watch?v=uRGljemfwUE
- cheesecake_luvr 5y agoOn a side note: when I browse to that page in Firefox (92.0.1) from HN I can't go back to HN - the back arrow is disabled. What gives?
- DevoidSimo 5y agoDo you have the facebook container extension? That closes the current tab, opens a new tab with a container, then goes to the facebook link. Reopening the last closed tab works for me, although I haven't noticed this before since I always open links in a new tab.
- cheesecake_luvr 5y agoTried Edge, it works as expected. Tried turning off Facebook Container, it also works as expected. So you are right good Sir! Still, a bit unexpected behaviour though.
- PostThisTooFast 5y agoIf your business "relies on Facebook," it's already fucked. You should see this as a wake-up call to GTFO.
- Elyes-ghorbel 5y agoCould you please be more clear about ''no evidence that user data was compromised''
- niko001 5y agoIt would be interesting to estimate what dollar value can be ascribed to a x-hour FB outage, both in terms of lost ad revenue for FB itself as well missed conversions/revenue for businesses running ads on FB/IG.
- phtrivier 5y agoDoes anyone know if FB's advertisement contracts even have an SLA ? I can completely picture a world in which many people bought some ads yesterday morning (say, to promote an event that occured yesterday evening), the ads were never displayed to anyone, and FB will keep the money, thank you.
- wepple 5y agoDon’t forget WhatsApp users switching to Signal and possibly never returning
- lionkor 5y agoSo this is pure conspiracy theory, but to me this could be a security issue. What if something deep in the core of your infrastructure is compromised? Everything at risk? Id ask my best engineer, hed suggest to shut it down, and the best way to do that is to literally pull the plug on what makes you public. Tell everyone we accidentally messed up a BGP and thats it. But yeah, likely not.
- colordrops 5y agoSpeaking of conspiracies, one that is floating around is that this was done to cover up spread of information around the Pandora Leak.
- can16358p 5y agoEven though it's probably not that, I must admit the fact that I absolutely love reading theories like this.
- fragmede 5y agoBGP is public routing information and multiple external sources are able to confirm that aspect of the story. It makes for a good conspiracy theory but the BGP withdrawal is as real as the Moon landing.
- throw0101a 5y ago> It makes for a good conspiracy theory but the BGP withdrawal is as real as the Moon landing. I wasn't aware that Stanley Kubrick was now in NetOps. /s
- laurent92 5y agoIf Facebook had been under actual attack, and defended by taking itself off the internet… that would be the most hands-on approach to security.
- EricE 5y agoMany have pointed out that a couple of weeks ago Facebook had a paper out on how they had implemented a fancy new automated system to manage their BGP routes. Whoops! Never attribute to malice that which can more easily be explained by stupidity and all that.
- stormdennis 5y agoThe mobile whatsapp app should notify that the whatsapp servers are down and not allow you to just send messages that won't arrive for six hours
- herald67 5y agoyup, they should. I couldn't send the messages as well and thought my mobile had some issues and tried rebooting it
- dTal 5y agoThe app is designed under the assumption that Facebook servers are never down. If you can't reach the servers, the problem is assumed to be client-side, in which case they have decided the best UI is to keep retrying (not unreasonably in a mobile context). The only way to disambiguate "no internet service" (extremely common) with "Facebook dropped off the internet" (black-swan rare) is to ping some other, third party service. Unless that third party service has infrastructure as good as Facebook's, it will drown in pings the moment Facebook genuinely goes offline. I can see why Facebook wouldn't want to open that can of worms, if they even envisaged this failure mode (unlikely considering the chaos it caused).
- stormdennis 5y agoGood points. Then it should just tell the user that they appear to have no internet service.
- EricE 5y ago>The app is designed under the assumption that Facebook servers are never down. Which was and is a lame assumption. Stuff happens. SMTP wouldn't even be phased by this; it would just pick up where it left off. I've seen far too many applications fail in bizzare ways because people make unrealistic assumptions like "X will ALWAYS be there". Sure it's highly unlikely, but when you have multiple things making the same dumb assumptions, on the inevitable day when multiple things that need X and X is suddenly no longer there then you start to get cascading effects of Y that relied on something that relied on X not being there when it is assumed that it would always be there so now Y fails, and then something dependent in the same way on Y unexpectedly fails and so on. One should never assume that anything will "always" be available. That's an incredibly unrealistic assumption; and the more interconnected things become, the chances of these really nasty dependency chains/cascade failures skyrocket - leading to far worse outages and longer recovery times.
- herald67 5y agoDo you think DLT/ blockchain can minimize this from happening again in the future?
- go_prodev 5y agoI worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway. Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications. Sidenote: I asked the bank infrastructure team why the DR site was in the same earthquake zone, and they thought I was crazy. They said if there's an earthquake we'll have bigger problems to deal with.
- xchaotic 5y ago"I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway." so prod wasn't down and he fixed it in a hour and they fired the guy who knew how to fix such things so quickly. Idiot manager at the bank.
- go_prodev 5y agoI agree it was very heavy handed, but I suspect there was more at play (not the first mistake, and some regulatory reporting that may have looked bad for higher ups)
- datavirtue 5y agoHad a DBA once who was playing around with database projects in visual studio and he managed to hose the production database in the course of it. This caused our entire system to go down. Prostrate, he came before the COO expecting to be canned with much malice. The COO just asked if he learned his lesson and said all is forgiven.
- drcross 5y ago>DR downtime was about an hour, but the bank fired him anyway The US, not even once. The guy should have had "reload in 10", an outage window and config review. There must be more to this story than it being a firable offence for causing a P2 outage for an hour.
- crtasm 5y ago"To all the people and businesses around the world who depend on us, " ... yesterday was another example of why you shouldn't depend on us to such an extent.
- throwawaymanbot 5y agowho. cares. May the outage be longer. And May Mark be removed as its snakehead.
- dr_hooo 5y agoWhy is this non-post on the frontage? It's PR only
- s225197 5y agoInformaticien
- ggggg5000842 5y agogggg5000842