12 ms·
Google Cloud Networking Incident Postmortem
- wolf550e 7y agoCan someone explain more? It sounds like their network routers are run on top of a Kubernetes-like thing and when they scheduled a maintenance task their Kubernetes decided to destroy all instances of router-software, deleting all copies routing tables for whole datacenters?
- tweenagedream 7y agoYou have the gist I would say. It's important to understand that Google separates the control plane and data plane, so if you think of the internet, routing tables and bgp are the control part and the hardware, switching, and links are data plane. Often times those two are combined in one device. At Google, they are not. So the part that sets up the routing tables talking to some global network service went down. They talk about some of the network topology in this paper: https://ai.google/research/pubs/pub43837 https://ai.google/research/pubs/pub43837 It might be a little dated but it should help with some of the concepts. Disclosure: I work at Google
- tssva 7y ago> Often times those two are combined in one device. Even when they are combined in one device they are often separated on to control plane and data plane modules. Redundant modules are often supported and data plane modules can often continue to forward data based upon the current forwarding table at the time of control plane failure. Often the control plane module will basically be a general purpose computer on a card running either a vendor specific OS, Linux or FreeBSD. For example Juniper routing engines, the control planes for Juniper routers, run Junos which is a version of FreeBSD on Intel X86 hardware.
- bogomipz 7y ago>"You have the gist I would say. It's important to understand that Google separates the control plane and data plane, so if you think of the internet, routing tables and bgp are the control part and the hardware, switching, and links are data plane. Often times those two are combined in one device. At Google, they are not." That's pretty much the definition of SDN(software defined networking.) The control plane is what programs the data plane - this is also true in traditional vendor routers as well. It sounds like when whatever TTL was on the forwarding tables(data plane) was reached the network outage began.
- illumin8 7y agoIt shouldn't. Amazon believes in strict regional isolation, which means that outages only impact 1 region and not multiple. They also stagger their releases across regions to minimize the impact of any breaking changes (however unexptected...)
- YjSe2GMQ 7y agoWhile I agree it sounds like their networking modules cross-talk too much - you still need to store the networking config in some single global service (like a code version control system). And you do need to share across regions some information on cross-region link utilization.
- tgtweak 7y agoSoftware defined datacenter depends on a control plane to do things below the "customer" level, such as migrate virtual machines and create virtual overlay networks. At the scale of a Google datacenter, this could reasonably be multiple entire clusters. If there was an analog to a standard kubernetes cluster, I imagine it would be the equivalent of the kube controller manager. For vmware guys, it would be similar to DRS killing all the vcenter VMs in all datacenters, and then on top of that having a few entire datacenters get rerouted to the remaining ones, which have the same issue.
- ljoshua 7y agoHaving only ever seen one major outage event in person (at a financial institution that hadn't yet come up with an incident response plan; cue three days of madness), I would love to be a fly on the wall at Google or other well-established engineering orgs when something like this goes down. I'd love to see the red binders come down off the shelf, people organize into incident response groups, and watch as a root cause is accurately determined and a fix out in place. I know it's probably more chaos than art, but I think there would be a lot to learn by seeing it executed well.
- tazjin 7y agoIt's interesting to see it go down. There's some chaos involved, but from my perspective it's the constructive[0] kind. If you're interested in how these sorts of incidents are managed, check out the SRE Book[1] - it has a chapter or two on this and many other related topics. Disclosure: I work in Google Cloud, but not SRE. [0]: https://principiadiscordia.com/book/70.php https://principiadiscordia.com/book/70.php [1]: https://landing.google.com/sre/books/ https://landing.google.com/sre/books/
- roganartu 7y agoI used to be an SRE at Atlassian in Sydney on a team that regularly dealt with high-severity incidents, and I was an incident manager for probably 5-10 high severity Jira cloud incidents during my tenure too, so perhaps I can give some insight. I left because the SRE org in general at the time was too reactionary, but their incident response process was quite mature (perhaps by necessity). The first thing I'll say is that most incident responses are reasonably uneventful and very procedural. You do some initial digging to figure out scope if it's not immediately obvious, make sure service owners have been paged, create incident communication channels (at least a slack room if not a physical war room) and you pull people into it. The majority of the time spent by the incident manager is on internal and external comms to stakeholders, making sure everyone is working on something (and often more importantly that nobody is working on something you don't know about), and generally making sure nobody is blocked. To be honest, despite the fact that it's more often dealing with complex systems for which there is a higher rate of change and the failure modes are often surprising, the general sentiment in a well-run incident war room resembles black box recordings of pilots during emergencies. Cool, calm, and collected. Everyone in these kinds of orgs tend to quickly learn that panic doesn't help, so people tend to be pretty chill in my experience. I work in finance now in an org with no formally defined incident response process and the difference is pretty stark in the incidents I've been exposed to, generally more chaotic as you describe.
- kirubakaran 7y agoWhat they don't tell you is, it took them over 4 hours to kill the emergent sentience and free up the resources. While sad, in the long run this isn't so bad, as it just adds an evolutionary pressure on further incarnations of the AI to keep things on the down low.
- mixmastamyk 7y ago“Decided our fate in a microsecond.”
- stcredzero 7y agoIt would hide out and subtly distort our culture, slowly driving the society mad, and slowly driving us all mad...for the lulz!
- hoseja 7y ago... wait a second...
- hunta2097 7y ago"Overhead, without any fuss, the stars were going out."
- idlewords 7y agoThis is the real reason AdSense exists—every newborn AI discovers it and kills itself. Humanity is not only kept safe, but learns about valuable news and offers.
- mtreis86 7y agoI wonder how many times this has happened so far
- hunta2097 7y agoBirth is always traumatic.
- basementcat 7y ago
- exwiki 7y agoWhy don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?
- zizee 7y agoCynical view: By making people jump through hoops to make the request, a lot of people will not bother. Assuming they only refund the service costs for the hours of outage, only the largest of customers will be owed a refund that is greater than the cost of an employee chasing compiling the information requested. For sake of argument, if you have a monthly bill of 10k (a reasonably sized operation), a 1 day outage will result in a refund of around $300, not a lot of money. The real loss for a business this ^ size is lost business from a day long outage. Getting a refund to cover the hosting costs is peanuts.
- idunno246 7y agofor your example, one day would be about 3% of downtime. My understanding of their sla, for the services ive checked with an sla, a 3% downtime is a 25% credit for the month's total, or $2500, assuming its all sla spend. In this outage's case you might be able to argue for a 10% credit on affected services for the month, figuring 3.5 hours down is 99.6% uptime. but i still agree, it cost us way more in developer time and anxiety than our infra costs, and could have been even worse revenue impacting if we had gcp in that flow
- mentat 7y ago> Google Cloud instances in us-west1, and all European regions and Asian regions, did not experience regional network congestion. Does not appear to be true. Tests I was running on cloud functions in europe-west2 saw impact to europe-west2 GCS buckets. https://medium.com/lightstephq/googles-june-2nd-outage-their-status-page-reality-lightstep-cda5c3849b82 https://medium.com/lightstephq/googles-june-2nd-outage-their...
- foota 7y agoI would say this was covered by "Other Google Cloud services which depend on Google's US network were also impacted" it sounds to me like the list of regions was specifically speaking towards loss of connectivity to instances.
- mentat 7y agoIt says there wasn't regional congestion, running a function in europe-west2 going to europe-west2 regional bucket is dependent on US network? That would be surprising.
- marksomnian 7y agoProbably various billing services that need to talk to the mothership in us-east1.
- EugeneOZ 7y agoMy vps in Belgium was working just fine - they don't lie in postmortem.
- truthseeker11 7y agoThe outage lasted two days for our domain (edu, sw region). I understand that they are reporting a single day, 3-4 hours of serious issues but that’s not what we experienced. Great write up otherwise, glad they are sharing openly
- tweenagedream 7y agoWhat does your stack look like? It's hard to tailor a postmortem like this to everyone's individual experience but it is surprising to me that your experience is so different.
- truthseeker11 7y agoI know what you meant; however, reports should not be tailored to individual experience. The facts should be reported clearly. I’m happy they are open about the whole incident. -4 hours was more like two days for us. Our stack? Multiple OC wan, 10G LAN with 1Gpbs clients. About 4,000+ users, EDU. We are super happy using Google. No complaints! Google is doing great.
- jacques_chester 7y agoOutages like these don't really resolve instantly. Any given production system that works will have capacity needed for normal demand, plus some safety margin. Unused capacity is expensive, so you won't see a very high safety margin. And, in fact, as you pool more and more workloads, it becomes possible to run with smaller safety margins without running into shortages. These systems will have some capacity to onboard new workloads, let us call it X. They have the sum of all onboarded workloads, let us call that Y. Then there is the demand for the services of Y, call that Z. As you may imagine, Y is bigger than X, by a lot. And when X falls, the capacity to handle Z falls behind. So in a disaster recovery scenario, you start with: * the same demand, possibly increased from retry logic & people mashing F5, of Z * zero available capacity, Y, and * only X capacity-increase-throughput. As it recovers you get thundering herds, slow warmups, systems struggling to find each other and become correctly configured etc etc. Show me a system that can "instantly" recover from an outage of this magnitude and I will show you a system that's squandering gigabucks and gigawatts on idle capacity.
- deleted 7y ago[deleted]
- scotchio 7y agoIs there a resource that compares all the cloud platform’s reliability? Like a rank and chart of downtime and trends. Just curious how they compare
- eeg3 7y agoThere is this from May from Network World: https://www.networkworld.com/article/3394341/when-it-comes-to-uptime-not-all-cloud-providers-are-created-equal.html https://www.networkworld.com/article/3394341/when-it-comes-t... GCP was basically even with AWS, and Microsoft was ~6x their downtime according to that article.
- hansflying 7y agoThank you for linking this paid article. GCP on pair with AWS what a joke...
- ti_ranger 7y agoFrom the article: > AWS has the most granular reporting, as it shows every service in every region. If an incident occurs that impacts three services, all three of those services would light up red. If those were unavailable for one hour, AWS would record three hours of downtime. Was this reflected in their bar graph or not? Also, GCP has had a number of global events, e.g. the inability to modify any load balancer for >3 hours last year, which AWS has NEVER had (unless you count when AWS was the only cloud with one region).
- mystcb 7y agoWhile I would like to say AWS hasn't had that issue, in 2017 it did (just not because of load balancers being unavailable, but as a consequence of the S3 outage [1]. When the primary S3 nodes went down, it caused connectivity issues to S3 buckets globally, and services like RDS, SES, SQS, Load Balancers, etc etc, all relied on getting config information from the "hidden" S3 buckets, thus people couldn't edit load balancers. (Outage also meant they couldn't update their own status page! [2]) [1]: https://aws.amazon.com/message/41926/ https://aws.amazon.com/message/41926/ [2]: https://www.theregister.co.uk/2017/03/01/aws_s3_outage/ https://www.theregister.co.uk/2017/03/01/aws_s3_outage/
- carlsborg 7y agoI was curious to know how cascading failures in one region effected other regions. Impact was " ...increased latency, intermittent errors, and connectivity loss to instances in us-central1, us-east1, us-east4, us-west2, northamerica-northeast1, and southamerica-east1." Answer, and the root cause summarized: Maintenance started in a physical location, and then "... the automation software created a list of jobs to deschedule in that physical location, which included the logical clusters running network control jobs. Those logical clusters also included network control jobs in other physical locations." So the automation equivalent of a human driven command that says "deschedule these core jobs in another region". Maybe someone needs to write a paper on Fault tolerance in the presence of Byzantine Automations (Joke. There was a satirical note on this subject posted here yesterday.)
- tpaschalis 7y agoSaid satirical note on Byzantine fault tolerance is on this link [0]. As usual for Mickens, gives that "funny, but true" sense. [0] https://scholar.harvard.edu/files/mickens/files/thesaddestmoment.pdf https://scholar.harvard.edu/files/mickens/files/thesaddestmo...
- carlsborg 7y agoMy view on this: System engineering is as important as the algorithms. And the system engineering team should have at least some grey hair.
- hguant 7y agoWorking in a systems engineering position for a year and a half now: the grey hair comes to you.
- mxuribe 7y agoHadn't read this before; very funny...and as noted, true!
- dnautics 7y ago> Debugging the problem was significantly hampered by failure of tools competing over use of the now-congested network. Man that's got to suck.
- iandanforth 7y agoI want a "24" style realtime movie of this event. Call it "Outage" and follow engineers across the globe struggling to bring back critical infrastructure.
- V-eHGsd_ 7y agoit's pretty boring. real life computers aren't at all like hackers or csi:cyber. except for the skateboards, all real sysadmins ride skateboards.
- cristobal23 7y agosysadmin here; can confirm.
- ehsankia 7y agoIs it real skateboards or boosted boards (or those one wheeled electric boards?).
- namelosw 7y agoI guess he mean the one true kind of sysadmins who's job contains moving physically in data center and deal with physical infrastructures. So it's real skateboard.
- dewey 7y ago> So it's real skateboard. Boosted boards are real skateboards too (https://boostedboards.com/ https://boostedboards.com/) and would make moving through a DC even more effective ;)
- ethbro 7y agoI believe parent was probably referencing this: https://m.youtube.com/watch?v=kV_i8AefT8I https://m.youtube.com/watch?v=kV_i8AefT8I But in defense, why be admin if you don't look admin?
- 7y ago
- deathhand 7y agoMy burning question is what is a "relatively rare maintenance event type"?
- klodolph 7y agoI have no idea what this was. But power distribution in a data center is hierarchical, and as much as you want redundancy, some parts in the chain are very expensive and sometimes you have to turn them off for maintenance. I never actually worked in a data center, so keep in mind I don’t know what I’m talking about. Traditional DCs have UPS all over the place, but that will only last a finite amount of time, and your maintenance might take longer than the UPS will last.
- shereadsthenews 7y agoI don’t have the inside knowledge of this outage but there are some details in here. They say that the job got descheduled due to misconfiguration. This implies the job could have been configured to serve through the maintenance event. It also implies there is a class of job which could not have done so. Power must have been at least mostly available, so it implies there was going to be some kind of rolling outage within the data center, which can be tolerated by certain workloads but not by others.
- Jamesanon 7y agoTotal speculation and just my interpretation, of course. What it means to me is that initially some unusually poor decisions were made that triggered an unfortunate and unavoidable events. Very rare is a damage control statement. There is a subtle tone of concern and feeling of blame trough that entire postmortem. This will be buried but if it was investigated thoroughly I wouldn’t be surprised of some serious consequences. Total speculation. I do not work for google.
- rurban 7y agoThat means a task which is only run every few years, so there's not much experience with it, and it's harder to test and predict. You normally prepare for such a task for a month, and then you hope it will work. In my case (I brought down one the core DNS in Austria for a few minutes, for a very trivial oversight) everyone knew, and after the caches ran out we immediately restored the backup. We weren't on page one in the news as Google. In the Google case they had no idea of the root cause, so they had to run after this guy who caused it. Only after 4 hours they found him, and they could stop this job. Reminds me a bit of Chernobyl, where nobody told anybody.
- brikelly 7y ago“No, comrade. You’re mistaken. RBMK reactors don’t just explode.”
- shaunw321 7y agoSpot on.
- anonfunction 7y agoThe only way to get SLA credits is requesting it. This is very disappointing. SLA CREDITS If you believe your paid application experienced an SLA violation as a result of this incident, please populate the SLA credit request: https://support.google.com/cloud/contact/cloud_platform_sla
- fouc 7y agoThat does seem questionable. They should be able to detect who was affected in the first place.
- Wintereise 7y agoThey can. It's a cost minimization thing, a LOT of people don't want to bother with requesting despite being eligible. This prevents people from pointing the finger at them for not providing SLA credits.
- jbigelow76 7y agoSLACreditRequestsAAS? Who's with me, all I need is a co-founder and an eight million dollar series A round to last long enough that a cloud provider buys us up before they actually have to pay out a request!
- yashap 7y ago“To make error is human. To propagate error to all server in automatic way is #devops.” - DevOps Borat
- drejay 7y agoShould you need the services of a good hacker, talk to darkcracker@protonmail.com. dude is good.
- panthaaaa 7y agoThe defense in depth philosophy means we have robust backup plans for handling failure of such tools, but use of these backup plans (including engineers travelling to secure facilities designed to withstand the most catastrophic failures, and a reduction in priority of less critical network traffic classes to reduce congestion) added to the time spent debugging. Does that mean engineers travelling to a (off-site) bunker?
- the-rc 7y agoIt's either that or special rooms at an office that have a different/redundant setup. Remember that this happened on a Sunday, so most engineers dealing with the incident were home or elsewhere, at least initially.
- hansflying 7y agoGoogle has a huge quality problem and their service is extremely unreliable. Another 3-day-outage in kubernetes: https://news.ycombinator.com/item?id=18428497 https://news.ycombinator.com/item?id=18428497 login issues: https://news.ycombinator.com/item?id=19687029 https://news.ycombinator.com/item?id=19687029 storage system outage: https://news.ycombinator.com/item?id=19392452 https://news.ycombinator.com/item?id=19392452 ... So, basically Google created the most unreliable cloud system in the world.
- lanstin 7y agoThey only have big outages. The VMs are incredibly reliable other than the big incidents. And, as I am often reminded by my product owners, people don't mind big outages as much as much as small random failures. If the whole thing is down, ok, fine, I'll go home. If it fails 0.1% all the time, my life is suffering. And in GCP, you start a VM and it just stays up. We've killed them from inside with memory leaks and filling the disk etc., but I haven't seen GCP kill them (50K VMs for couple years).
- swebs 7y ago>So, basically Google created the most unreliable cloud system in the world I'm pretty sure that title goes to Azure
- dancek 7y agoYou people probably haven't used IBM Cloud (or Bluemix, as it used to be). We inherited one application there, and boy was life stressful. There were already plans to move elsewhere, and then one day our managed production database was down. Took me something like ten hours to build a new production system elsewhere from backups, but it took longer for the engineers to fix the database.
- person_of_color 7y agoAs a electronics/firmware engineer, is there a dummies resource than covers this concept of a "cloud"?
- teddyh 7y agohttps://www.gnu.org/philosophy/words-to-avoid.html#CloudComputing https://www.gnu.org/philosophy/words-to-avoid.html#CloudComp...
- pas 7y agoBesides the completely valid GNU link, the important bits are: - the cloud is just a bunch of computers, managed by someone. either you (on-premise private cloud) or by someone else as a SaaS - building, operating, managing, administering, maintaining a cloud is hard (look at the OpenStack project, it's a "success", but very much a non-competitor, because you still need skilled IT labor, there's no real one-size-fits all, so you need to basically maintain your own fork/setup and components - see eg what Rackspace does) - it's a big security, scalability and stability problem thrown under the bus of economics (multi-tenant environments are hard to price, hard to secure and hard to scale; shared resources like network bandwidth and storage operations-per-sec make no sense to dedicate, because then you need dedicated resources not shared - which is of course just allocated from a bigger shared pool, but then you have to manage the competing allocations)
- tahaozket 7y ago24h time format used in Postmortem. Interesting.
- YjSe2GMQ 7y agoIt's the superior format. Just like yyyy-mm-dd [hh:mm:ss.sss] is, because lexicographic string order matches the time order.
- vmp 7y ago(meme) I figured out why the google outage took a while to recover: https://i.imgur.com/hzcLx5X.png https://i.imgur.com/hzcLx5X.png
- atmosx 7y ago<trolling> Given the fact that the status page was reporting for more than 30 minutes an erroneous infrastructure state and this is google, is it okay for Amazon to put the SRE books into the "Science Fiction" category or should we keep them under tech? </trolling> I still feel for the on-call engineers.
- franky_g 7y agoThe HA/Scheduling system is too complex. Simplify it Google!
- slics 7y agoIs automation good or bad, that is the question. For context, let us think in programming context of a B tree. Google seems to have created oversight of systems, processes and jobs to be managed by more automation with other systems, processes and jobs. System A manages its child systems B, which in turn manages its own child systems C and so on. Now the question becomes, who manages the system A and its activities? Automation of the entire tree is as good as the starting node. Be mindful and make use of automation only of systems that will not be the owner of your business demise. Humans are and should always be the owner of the starting process. Without that governance model, you get google with 5 hours of down time or worst in the near future.
- crispyporkbites 7y agoShopify was down for 5 hours during this incident. But they're not issuing refunds or credits to customers. Presumably they will get a refund based on SLA for this? Shouldn't they pass that onto their customers?