4 ms·
Major credit to Cloudflare for publishing a clear, honest, and detailed description of what happened. I wish more companies would do this. One thing I’d be int
by obeattie 8y ago
Major credit to Cloudflare for publishing a clear, honest, and detailed description of what happened. I wish more companies would do this.
One thing I’d be interested to know more about is why it took 17 minutes to fix. While you can and always should strive to make them less likely, outages are inevitable, so how you respond is crucial. Here the outage was very obviously caused by a deployment that I’d assume was supervised by humans – why did it take 17 minutes to roll back?
- blackbrokkoli 8y agoI'm not an expert, but is 17 minutes for: - shit is not working - is this an attack? - no it's us - how? - that's how - let's go back - have to get supervisor - roll back huge thing really that long?
- dx034 8y agoWith ~150 data centres, roll back alone probably took 5-10 minutes. Don't think 17 minutes is that long.
- Cthulhu_ 8y agoFor simple PagerDuty alerts I already need 15 minutes to open the app / logs and figure out what's going on.
- wool_gather 8y agoGood point, although the way it's described it sounds like the problem cropped up right after deploy. So they would have been watching it actively. But as said above, 17 minutes to notice, figure out what's going on, decide what to do, and propagate the resolution seems reasonable.
- dfcowell 8y agoNot to mention selectively purging all of the rules created (I assume) at every edge server in their network that were blackholing all of the traffic to the resolver. There's probably a command for it, but all in all 17 minutes seems quite a respectable turnaround time.
- Steltek 8y agoDid they really deploy to all 150 DCs at once? Why was this release not done in phases? Not even a canary?
- dx034 8y agoMaybe they did, the strategy is probably to deploy to small DCs first. But that also means the DDOS detection wouldn't trigger. So probably something that didn't get caught during such tests.
- jacksmith21006 8y agoIn 2000 the answer would be no. In 2018 I think it is. Things change in a time when you would freak getting up in the morning and say google.com did not work.
- bovermyer 8y agoHumans react, analyze, and communicate at the same speed we did in 2000. Our tools may have gotten better, but that only cuts down on part of the process. Automated processes can only mitigate so many edge cases. Even then, humans need to be involved, and that slows things down.
- throwaway413 8y agoDisagree. Have you ever seen a toddler use an iPad? More input/stimuli in 2018 = higher capacity to analyze said inputs. Or an overload of capacity which results in the epidemic of mental illnesses and psychotic breakdowns we witness in this post-social media society. Another example - I can record a video and broadcast internationally, translated on the fly into dozens of languages, effectively communicating with significantly more people than if I could not harness that technical capability. In 2000 that communication process would have been orders of magnitude longer. (Did they have on-the-wire translation then? Idk, just making an assumption to illustrate my point.)
- phyzome 8y agoYes, this exactly. And to expand on what happens around that: - monitoring system picks up irregularity (smoothed over some window of time, which delays alerting) - alert propagates to humans - humans may take time to notice alert (even a page takes a few seconds to read) - humans make decisions, may need to talk to other humans (all of what you said above) - humans evaluate correct procedure, double-check it (you don't want them making the wrong "fix" and making something else worse, do you?) - humans execute commands - commands take time to run on large collections of computers (running them completely in parallel can cause thundering herd issues, in some cases)
- TomAnthony 8y agoThe problem is so clear in their write up that I can understand your thinking. However, in reality as this was going down it was probably not that clear cut. Especially when you consider that they are getting DoS attacks every 2-3 minutes - so all deploys are going out into a hectic world and the dots maybe aren't that easy to connect under those circumstances.