26 ms·
Post Mortem of Google Outage on 14 December 2020
- koreanguy 6y agoall google products are interlinked
- polote 6y agooff-topic :Do we know what happened a day after that when Gmail returned "this email doesn't exist" ?
- rubyron 6y agoI’d say that’s solidly on-topic, and was more damaging to my business than the previous outtage.
- azornathogron 6y agoHere is the incident report for the Gmail problem: https://static.googleusercontent.com/media/www.google.com/en//appsstatus/ir/4et50yp2ckm8otv.pdf https://static.googleusercontent.com/media/www.google.com/en... It is linked from the Google Workspace status page here: https://www.google.com/appsstatus#hl=en-GB&v=issue&sid=1&iid=a8b67908fadee664c68c240ff9f529ab https://www.google.com/appsstatus#hl=en-GB&v=issue&sid=1&iid...
- polote 6y agothanks, so a change of an env variable put the most used email system in the world down during 6 hours, not sure I can believe that
- xyzelement 6y ago>> so a change of an env variable put the most used email system in the world down during 6 hours, not sure I can believe that I can totally believe it. In my experience, the bigger the outage the stupider-seeming the cause.
- jeffbee 6y agoGoogle has had a bunch of notorious outages caused by similar things, including pushing a completely blank front-end load balancer config to global production. The post mortem action items for these are always really deep thoughts about the safety of config changes but in my experience there, nobody ever really fixes them because the problem is really hard. For this kind of change I would probably have wanted some kind of shadow system that loaded the new config, received production inputs, produced responses that were monitored but discarded, and had no other observable side effects. That's such a pain in the ass that most teams aren't going to bother setting that up, even when the risks are obvious.
- jeffbee 6y agoActually now that I remember correctly, back when I was in that barber shop quartet in Skokie^W^W^W^W^W err, back when I was an SRE on Gmail's delivery subsystem, we actually did recognize the incredible risk posed by the delivery config system and our team developed what was known as "the finch", a tiny production shard that loaded the config before all others. It was called the finch to distinguish it from "the canary" which was generally used for deploying new builds. I wonder if these newfangled "best practices" threw the finch under the bus.
- jeffbee 6y agoDamn, that's awful. Should lead to a deep questioning of people who claim to be migrating your junk to "best practices" if the old config system worked fine for 15 years and the new one caused a massive dataloss outage on Day 1.
- rachelbythebay 6y agoNew one gets you promotions. Old one is boring and does not. Guess what people want to work on. Seen it happen too many times.
- derwiki 6y agoI read this and identified with it, and then read your username
- silentsea90 6y agoWouldn't paint all migrations with the same brush
- gojomo 6y agoThat Google is choosing to use a PDF (!) as the official incident-reporting media is as confidence-destroying as was the outage.
- dane-pgp 6y agoIt's almost as disappointing as the fact that their status page doesn't redirect from HTTP to HTTPS. (Presumably they wanted to make the status page depend on as few services as possible, to prevent a scenario where an outage also affects the status page itself, but whatever script they are using to publish updates to the page could also perform a check that the HTTPS version of the site is accessible, and if not, remove the redirect). Could we get the URL of the submission updated please? (Also, it would be nice if the submission form added an "Are you sure?" step when people submit HTTP links).
- gojomo 6y agoWow, surprising PDF partisan downvotes here!
- pdkl95 6y ago> A configuration change during this migration shifted the formatting behavior of a service option so that it incorrectly provided an invalid domain name, instead of the intended "gmail.com" domain name, to the Google SMTP inbound service. Wow... how was this even possible? Did they do any testing whatsoever before migrating the live production system? They misformatting the domain name should have broken even basic functionality tests. I wonder if they didn't actually test the literal "gmail.com" configuration, due to dev/testing environments using a different domain name? I had that problem when on my first Ruby on Rails project due to subtle differences between the development/test/production settings in config/environments/. Running "rake test" is not a substitute for an actual test of the real production system.
- NoodleIncident 6y agoThe nature of configuration is that it's different for prod and your testing environments. It doesn't make it impossible to test your prod config changes, but it's not that simple either.
- tpmx 6y agoThought: There should be an official Google status dashboard for "free" (paid for via personal ad targeting metadata) services like Search, Gmail, Drive, etc. With postmortems, too. Probably won't happen unless it's mandated by law, though.
- deleted 6y ago[deleted]
- Splendor 6y agoThere are several. Here's one example: https://downdetector.com/ https://downdetector.com/
- whermans 6y agoExisting dashboards showed the services as being up, since the frontends worked fine - as long as you did not try to authenticate.
- Privacy846 6y agoKeep your scare quotes.
- tpmx 6y agoI kinda thought the sentence within the parentheses explained it all.
- Privacy846 6y agoAren’t you clever, explaining to us that companies don’t do things for free.
- paxys 6y agoFree and paid versions of Google services share pretty much the entire stack, so the Google Workspace (previously G Suite) status page works just fine – https://www.google.com/appsstatus#hl=en&v=status https://www.google.com/appsstatus#hl=en&v=status.
- jpxw 6y agoTLDR: the team forgot to update the resource quota requirements for a critical component of Google’s authentication system while transitioning between quota systems.
- jrockway 6y agoThat's not the TL;DR is it? It seems that the quota system detected current usage as "0", and thus adjusted the quota downwards to 0 until the Paxos leader couldn't write, which caused all of the data to become stale, which caused downstream systems to fail because they reject outdated data.
- basicneo 6y agoIs that an accurate tl;dr? "As part of an ongoing migration of the User ID Service to a new quota system, a change was made in October to register the User ID Service with the new quota system, but parts of the previous quota system were left in place which incorrectly reported the usage for the User ID Service as 0."
- jamisteven 6y agothats one hell of a post mortem.
- mehrdadn 6y agoI'm skimming this but so far I still don't understand how any kind of authentication failure can or should lead to an SMTP server returning "this address doesn't exist". Does anyone see an explanation? Edit 1: Oh wow I didn't even realize there were multiple incidents. Thanks! Edit 2 (after reading the Gmail post-mortem): AH! It was a messed up domain name! When this event happened, I said senders need to avoid taking "invalid address" at face value when they've recently succeeded delivering to the same addresses. But despite the RFC saying senders "should not" repeat requests (rather than "must not"), many people had a lot of resistance to this idea, and instead just blamed Google for messing up implementing the RFC. People seemed to think every other server was right to treat this as a permanent error. But this post-mortem makes it crystal clear how that completely missed the point, and in fact wasn't the case at all. The part of the service concluding "address not found" was working correctly! They didn't mess up implementing the spec—they messed up their domain name in the configuration. Which is an operational error you just can't assume will never happen—it's exactly like the earlier analogy to someone opening the door to the mailman and not recognizing the name of a resident. A robust system (in this case the sender) needs to be able to recognize sudden anomalies—in this case, failure to deliver to an address that was accepting mail very recently—and retry things a bit later. You just can't assume nobody will make mistake operational mistakes, even if you somehow assume the software is bug-free.
- klysm 6y agoIf I had to guess I would say sometimes when you request something you aren't authorized to see you get a 404 because they don't want you to be able to tell what exists or not without any creds.
- mehrdadn 6y agoSending an email doesn't require "seeing" anything other than whether the server is willing to receive email for that address though, right? Also, that problem occurs due to lack of authorization, not due to authentication failure, right? Authentication failure is a different kind of error to the client—if the server fails to authenticate you, clearly you already know that it's not going to show you anything?
- physicsgraph 6y agoMy favorite quote: "prevent fast implementation of global changes." This is how large organizations become slower compared to small "nimble" companies. Everyone fails, but the noticeability of failure incentivizes large organizations to be more careful.
- Xorlev 6y agoSlowing down actuation of prod changes to be over hours vs. seconds is a far cry from the large org / small org problem. Ultimately, when the world depends on you, limiting the blast radius to X% of the world vs. 100% of it is a significant improvement.
- inopinatus 6y agoIt doesn’t have to be this way, but that’s partly a matter of culture. By aspiring to present/think/act as a monoplatform, Google risks substantially increasing the blast radius of individual component failure. A global quota system mediating every other service sounds both totally on brand, and also the antithesis of everything I learned about public cloud scaling at AWS. There we made jokes, that weren’t jokes, about service teams essentially DoS’ing each other, and this being the natural order of things that every service must simply be resilient to and scale for. Having been impressed upon by that mindset, my design reflex is instead to aim for elimination of global dependencies entirely, rather than globally rate-limiting the impact of a global rate-limiter. I’m not saying either is a right answer, but that there are consequences to being true to your philosophy. There are upsides, too, with Google’s integrated approach, notable particularly when you build end-to-end systems from public cloud service portfolios and benefit from consistency in product design, something AWS eschews in favour of sometimes radical diversity. I see these emergent properties of each as an inevitability, a kind of generalised Conway’s Law.
- simon1573 6y agoDo you blog? I really enjoyed reading that.
- 6y ago
- blaisio 6y agoHmm one thing that jumped out at me was the organizational mistake of having a very long automated "grace period". This is actually bad system architecture. Whenever you have a timeout for something that involves a major config change like this, the timeout must be short (like less than a week). Otherwise, it is very likely people will forget about it, and it will take a while for people to recognize and fix the problem. The alternative is to just use a calendar and have someone manually flip the switch when they see the reminder pop up. Over reliance on automated timeouts like this is indicative of a badly designed software ownership structure.
- sleepydog 6y agoI agree, and even if the grace period were a good idea, enforcement should have slowly ratcheted up over the grace period, rather than having full enforcement immediately after it expired.
- tantalor 6y agoThis is also called a "time bomb". It's a bad thing.
- NoodleIncident 6y agoWhat's insane to me is that a grace period is built in, but that the mere fact that this grace protection is active isn't a giant neon sign on their dashboards and alerts. I do see how it could slip through the cracks, since it was the reported usage that was wrong, not the quota itself.
- paxys 6y agoWe once found a very annoying bug which was caused because someone set a feature flag to a tiny rollout % and then left the company without updating it. It sat that way for 2 years before someone finally noticed.
- deleted 6y ago[deleted]
- sabujp 6y agos/post mortem/incident analysis/
- b0afc375b5 6y agoWhat is the difference between the two? I tried searching for 'post mortem vs incident analysis' but couldn't find anything.
- gojomo 6y agoSome org might have decreed internal fine distinctions, but in common parlance, both terms are used for the same kind of after-event writeups.
- ph0rque 6y agoWell, post mortem means "after death" in Latin. So it would seem the difference is, one can recover from an incident...
- gregoda 6y agoI'd tend to think that 'post-mortem' translates (in common usage) more accurately as 'after termination' -- a process performed at the completion of some other process. It's really a good idea to do post-mortems on successes as well as failures.
- papito 6y agoPost-mortem is just a gruesome term.
- gojomo 6y agoIn this domain, the two terms are synonyms. And neither "post-mortem" nor "incident report" appear on the log page, making either an equally fair synthesized title.
- deleted 6y ago[deleted]
- AdrianB1 6y agoI am happy to see Google makes mistakes too, even if theirs are in areas a lot more complex than what I see at my job :) Now on a serious note, the increasing complexity of the systems and architectures makes it more challenging to manage and makes failures a lot harder to prevent.
- ak217 6y agoI've worked on a number of systems where code was pushed to a staging environment (a persistent functional replica of production where integration tests happen) and sat there for a week before being allowed in production. A staging setup might have prevented this scenario, since the quota enforcement grace period would expire in staging a week before prod and give the team a week to notice and push an emergency fix.
- deleted 6y ago[deleted]
- papito 6y agoStaging is never hammered with production traffic, often not exposing problems. It's a sanity check for developers and QA, essentially. On the scale of Google, you test in production, but with careful staged canary deployments. Even a 1% rollout is more than most of us have ever dealt with.
- shadowgovt 6y agoIn this specific case, the problem they were into is that the quota system had been different for three months, but the differences were not being enforced. It's very unclear how they would have gone about Canary that period really, the canary should have been done on the quota side. The quota system should have enabled enforcement to 1% of the quota clients. but it turns out it's actually hard to configure that sort of thing. The ways you can slice subsets of Google infrastructure are absolutely holographic, and it costs engineering time to change the slices.
- deleted 6y ago[deleted]
- tpolm 6y agoit is interesting that is both cases (recent gmail and this one) it was a "migration": "As part of an ongoing migration of the User ID Service to a new quota system" "An ongoing migration was in effect to update this underlying configuration system" it was not a new feature, not a massive hardware failure, it was migrating part of the working product due to some unclear reason of "best practices". both of those migrations failed with symptoms suggesting that whoever was performing them did not have deep understanding of systems architecture or safety practices and there was no one to stop them from failing. Signs of slow degradation of engineering culture at Google. There will be more to come. Sad.
- cpncrunch 6y ago>both of those migrations failed with symptoms suggesting that whoever was performing them did not have deep understanding of systems architecture or safety practices and there was no one to stop them from failing. Can any single person at Google have a full understanding of all the dependencies for even a single system? I have no idea, as I've never worked there, but I would imagine that there is a lot of complexity.
- tpolm 6y agosomehow they managed to build complex systems like gmail, continuously develop new features there and not have massive outages due to "migrations" - suggests that something that they were doing right, they are no longer able to do
- caturopath 6y agoI'm pretty sure Google has had occasional severe outages for their whole history.
- joshuamorton 6y agoA single event is not data.
- 6y ago
- gregw2 6y agoAnyone know if there's a similar public outage report for that late November AWS us-east-1 outage?
- ignoramous 6y agoThat was a quote outage imposed by Linux defaults: https://news.ycombinator.com/item?id=25236057 https://news.ycombinator.com/item?id=25236057
- x87678r 6y agoI'm curious about big outages like this in big internet corps. Does anyone know if SREs in Europe fixed the problem or it relied on people in Mountain View? When an outage this big hits do devs get involved? I've worked as an SRE and it sucks to be fixing developer's bugs in the middle of the night.
- mikelward 6y agoYes, SRE teams typically have a sister SRE team in another continent and time zone.
- deleted 6y ago[deleted]
- asdfasgasdgasdg 6y agoI assume this is in the SRE book, but a tier one product like the identity service will have global SRE coverage (i.e. at least three SRE teams so that there is always an SRE group for whom it is daytime holding the pager). Devs are often involved in diagnosis, but are less often required for mitigation, as the mitigation is almost always to revert whatever change caused the problem. This is a simplification of course, but it gives an idea of the general pattern.
- smueller1234 6y agoI won't comment on the incident, but I can tell you that we have two, not three, SRE sibling teams each. That still gives awake-time coverage, but not working-hours coverage. We simply pay folks for the time spent oncall outside of working hours. (Google SRE)
- ryanobjc 6y ago4am is noon Europe time, so the sres in Europe would have gotten the page and been on top of their game. They fixed it pretty quick. Of course in a global outage, nothing is fast enough.
- guenthert 6y ago
- jiggawatts 6y agoI've been bitten by quota grace periods before, they're the "buffer bloat" of cloud platform management systems. You build something, it seems fine, and then a month later it keels over. My tiny disaster was caused by Amazon EFS. They provide a performance quota grace period of a month, during which all tests passed with flying colours. Turns out that EFS rations out IOPS per GB of stored data, and this particular application stored only a few MB at the time, because it was brand new and hadn't accumulated anything yet. I had a very angry customer calling up asking me to please explain why they were seeing an average of 0.1 IOPS... From: https://docs.aws.amazon.com/efs/latest/ug/performance.html https://docs.aws.amazon.com/efs/latest/ug/performance.html "The baseline rate is 50 MiB/s per TiB of storage (equivalently, 50 KiB/s per GiB of storage). AWS now provides a performance floor of 1 MiB/s, but at the time there was no floor. If I remember correctly, this application had something like 2 MIB of data, which was constantly being updated by various processes, so there was no quota being accumulated. The system performance went from something like 1 Gbps to 100 bytes per second instantly. It took 10 seconds for a 1 KiB I/O to complete. Fun times, fun times...
- crmd 6y agoThe linear IOP density model seems clever and logical but is a huge source of headaches because it includes a patently false assumption that IOPs scale in proportion to growth in object size. Performance quotas should be assigned at the object level (block device, file system, bucket) regardless of size.
- vlovich123 6y agoIt’s not because the underlying storage is 1 TiB disk, so you’re getting a fixed allocation that’s shared among many other 50mb/s clients (ie assuming flash and overall transfer rates of ~2gigabytes/, that lets them coex roughly 40 customers on 1 machine without over subscribing)? Isn’t that kind of pricing mode about the only one that would be feasible to implement to make this cost effective? How are you thinking it should work? Fixed cost per block transferred?
- 6y ago
- nullifidian 6y ago"I hated that app (vscode) on my last laptop, it was very slow and bloated, and I've been considering switching to something else. Now I don't care enough to see whether it's actually optimized or not, it's faster than my brain and that's quite enough." That's the sad part. That's how software becomes slower and slower with every year.
- herf 6y agoIn Thunderbird, OAuth2 login is still broken. The login page prompts for email again and again, never makes it to the password.
- aleph1 6y agoThis could be caused by Google not recognizing and blocking Thunderbird's default user agent; try toggling general.useragent.compatMode.firefox to true (which basically has TB emulate Firefox's user agent)
- BugWatch 6y agoThank you! This worked in my case.
- BugWatch 6y agoI am having the same difficulties, and it's giving me a headache. The only workaround was to change the authentication method to "Normal password", and enable the "Allow less secure apps" setting on the Gmail account. And that's a major PITA since I have 40-ish Gmail accounts. (Why so many? I tend to separate my different concerns for privacy and security reasons.)
- lbblack 6y agoI think what most software developers will come to find out with Cloud outages, it isn't necessarily that the tech is working incorrectly - it's most likely that it was not configured in the way it was intended...
- xvector 6y agoCan you imagine the stress of being on-call for this? Shudder!
- ChuckMcM 6y agoWell I guess the "Code Purple" got its own "Code Red"[1] The take away for me here is that maximizing resource utilization continues to be a hard problem and as you get better at it, the margin for errors is smaller and smaller. [1] Sorry its an inside Google joke.
- TedShiller 6y agoTo be honest, I never understood the point of companies publishing post mortems after outages. What is it supposed to accomplish? We the users don’t understand their infrastructure anyway, how are we expected to understand the post mortem? Besides, as the users why should we care why things went down? The fact is, they did. And that’s the only thing that matters. I don’t feel better knowing WHY I lost my services. I just feel bad BECAUSE I lost them. Or are post mortems supposed to reassure users that the outage won’t happen again? It doesn’t reassure that either because by definition unpredictable outages always happen due to something new and unpredictable. This post mortem certainly won’t stop the next outage from happening. We KNOW there will be more, we just don’t know when. Or are post mortems supposed to show that the company takes full responsibility for what happened? But they always do and are fully expected to. So it’s meaningless. No company would ever say “we don’t take responsibility for this error we caused”. Even in the case of massive data leaks, which cannot be reversed, companies always take full responsibility. And it doesn’t help anyone. The only thing post mortems show is that the company didn’t do their job or was careless or disorganized or confused. But we already know that, because they had an outage. So what’s the point?
- ithkuil 6y agoRecruiting?
- tpxl 6y agoA good post mortem is basically "We've made a mistake, try not to make the same mistake as us". Useless to end users, useful to engineers in similar situations.
- Terretta 6y agoLearning. If problems were solved, nobody would write code. That new code or config is being deployed shows these problems (new feature rollouts, migrations, scaled resilience, etc.) are not yet formally solved. As such, things not known will become known — and usually be revealed in prod. Similarly, in technology systems, there’s no such thing as human error, only uncaught error conditions. Postmortems capture learning, so the conditions can be caught next time. Publishing them shows the engineering organization understands and applies this learning loop.
- wodenokoto 6y agoIs there an ELI5 for this? I really don’t understand what was under a quota and what it means that the quota had a grace period. Did YouTube not update their authentication to a new version of the api, but they had a quota for old api calls that ran out?
- netdur 6y agoExactly how I predicted the problem could be https://news.ycombinator.com/item?id=25416544 https://news.ycombinator.com/item?id=25416544
- carlsborg 6y agoConfiguration errors strike again.
- Angostura 6y agoI'm not an infrastructure engineer - could someone explain the benefits of using quotas for something like an authentication service. It feels like something that shouldn't really need a quota - unless the idea is to monitor services that have run amok.
- halflings 6y agoAny service running way over quota could break all other services. (at the end of the day, resources are physically limited to what you have in your datacenters) Quotas are one way to isolate this impact to that service in particular. Of course, when it's a critical service like authentication, it hardly isolates anything... but I can't think of a better alternative.
- beezle 6y agoSlightly off topic - it would be refreshing if financial and other institutions provided similar public post mortems on incidents that affect large numbers of their clients. Recent ones that come to mind are Interactive Brokers and Robinhood.