10 ms·
Facebook blames a server configuration change for yesterday’s outage
- halfnibble 8y agoWAT. A server configuration change? What kind of server configuration can affect presumably thousands of machines replicated across the globe? I'm trying to understand this.
- sarcasmatwork 8y agoBad configuration in a tunnel, IP, BGP etc. https://www.bleepingcomputer.com/news/technology/facebook-and-instagram-down-in-global-outage/ https://www.bleepingcomputer.com/news/technology/facebook-an...
- gizmo385 8y agoYes, but those are all reversible relatively quickly. They aren't something that I would think ought to take almost an entire day to resolve.
- ceejayoz 8y agoThe side-effects of such a thing might not be as easily reversible. I've had to sit around waiting a couple hours for a Percona database cluster to re-sync after a major networking whoops, and it only had a few hundred gigabytes of data.
- ceejayoz 8y agoIt's hardly unheard of. https://en.wikipedia.org/wiki/Cascading_failure https://en.wikipedia.org/wiki/Cascading_failure An organization Facebook's size isn't gonna be applying configuration changes to one server at a time over SSH, either. A server configuration can easily affect thousands of machines across the globe if it's deployed to them all.
- wyre 8y agoShy did I take so long for Facebook to release the cause of the outage? If they are applying configuration changes at a large level shouldnt it be fairly easy for them to figure out what was the cause?
- ceejayoz 8y agoThat's silly. Error rates show as elevated on https://developers.facebook.com/status/dashboard/ https://developers.facebook.com/status/dashboard/ until 11pm Pacific yesterday. The @facebook Twitter account sent out a statement basically within an hour of the start of the next business day.
- viraptor 8y agoPossibly because it doesn't matter to us really. The postmortem will be interesting to read if they publish it, but otherwise - it stopped working. Time to explain it to the peanut gallery is better spent dealing with the actual issue.
- mikewhy 8y ago"To err is human, but to really foul things up you need a computer."
- jhayward 8y agoMost failures of this type end up being a cascading resource exhaustion problem propagated by an un- or mis-analyzed feedback or dependency path. It is frankly amazing it doesn't happen more often. I'm excluding the other common type of long outage, the head-desking "failover didn't work, backups are horked, it'll take 10's of hours to restore/cold start" kind.
- groestl 8y agoIt takes an admin to bring down a host, but it takes a configuration management system to bring down a site.
- CodeSheikh 8y agoSeems like a lot of people were forced to be productive yesterday (:
- canada_dry 8y ago> Seems like a lot of people were forced to be productive yesterday Well... except for those companies dumb enough to farm out their intra-company communications to facebook. A friend of mine's law firm had to resort to - OMG - the phone - yesterday.
- thisacctforreal 8y agoWait, the law firm talks about their dealings over Facebook?
- evv 8y agoFacebook Workplace, presumably (the slack competitor)
- indigodaddy 8y agoNope, Slack was working fine...
- ajsharp 8y agoThis is fucking bananas. For nearly a decade, Facebook has been at the forefront of innovating how code is deployed at global scale. They presumably have gradual rollouts, automated rollbacks, anomaly detection, not to mention (I assume) loads of organizational safeguards in place to ensure this sort of thing never happens. Something else happened. This was not a configuration issue. Edit: If it was, I'd expect a post-mortem post-haste.
- shereadsthenews 8y agoHow did you determine that Facebook leads this space? I recently read an article about how Facebook distributes RPMs internally and it struck me as the kind of thing an insane person might have invented fifteen years ago. I mean, NFS in front of glusterfs? Also, RPMs???? Talk about bananas.
- muxator 8y agoCould you link the article?
- shereadsthenews 8y agohttps://www.slideshare.net/mobile/PhilDibowitz/centos-at-facebook https://www.slideshare.net/mobile/PhilDibowitz/centos-at-fac...
- muxator 8y agoThanks. What I read there seems sensibile to me.
- ajsharp 8y agoI'm mostly guessing based on what I've read over the years. They've published a hefty corpus of work regarding their deployment infrastructure and greater code review/quality approach. Here's a few examples: - https://code.fb.com/web/rapid-release-at-massive-scale/ https://code.fb.com/web/rapid-release-at-massive-scale/. - https://www.quora.com/How-does-Facebook-release-deploy-process-work-What-tools-and-processes-does-Facebook-use-to-push-changes-to-the-site# https://www.quora.com/How-does-Facebook-release-deploy-proce...
- Chris_Chambers 8y agoImagine actually believing Facebook will ever tell you the truth about anything.
- mattbeckman 8y agoLargely impacted by this outage was how it affected those who use Facebook Login as a convenient OAuth option. Good thing for developers to remember if someone asks them to avoid a native login option.
- fenwick67 8y agoTo everyone jumping to conclusions, remember that the words "server" and "configuration" can mean a whole host of things. It doesn't necessarily mean they mistyped their nginx config.
- stingraycharles 8y agoExactly. Doing an upgrade to an internal email service is a configuration change. Scaling down a cluster is a configuration change. Mitigating a DDoS attack by implementing a firewall rule is a configuration change. “Configuration” in this context is the high level system configuration, and can mean pretty much anything that falls under that.
- ajsharp 8y agoAlso, 'many people had trouble accessing our apps and services' is some ninja-level gaslighting: https://twitter.com/facebook/status/1106229690069442560 https://twitter.com/facebook/status/1106229690069442560
- wowGaslight 8y agoThat's not how the word gaslight works.
- traek 8y agoIn what way is that gaslighting?
- aboutruby 8y agoNot a big surprise as it's one of the harder things to test.
- chowes 8y agoMust be related to them merging the chat backends for WhatsApp, FB, and Instagram
- segmondy 8y agoI suspect this too, fastly integrating the systems they can't be broken up.
- mancerayder 8y ago.. but, but, didn't they wring the DevOps folks through coding challenges, sorting algos and whiteboard coding before hiring them? I heard that's the number 1 way to ensure uptime at FAANG. (Configuration changes, that's the source of my sarcasm)
- clanrebornx 8y agoI think those coding interviews don't apply to frontend, sysadmin and DevOps.
- cannedslime 8y agoKeep calm and blame dev ops!
- lousken 8y agoIt took very long to fix so I think this was related to their databases, maybe some data corruption.
- rachelbythebay 8y agoAh, the tao of reliability. Async too.
- cryptokernel 8y agoWhat's new with facebook? Just leaving this here: https://coincircle.com/l/50VVxbObg3 https://coincircle.com/l/50VVxbObg3