49 ms·
Post-Mortem for Google Compute Engine’s Global Outage on April 11
- brianwawok 10y agoThis is a very good Post-Mortem. As I assumed it was kind of a corner case bug meet corner case bug met corner case bug. This is also why I am of afraid of a self driving cars and other such life critical software. There are going to be weird edge cases, what prevents you from reaching them? Making software is hard....
- reustle 10y agoI wouldn't worry so much. I'm sure self driving cars are going to save a lot more lives than they are going to end. Humans are terrible drivers, and the software will only get better.
- VonGuard 10y agoYeah, remember, auto-pilot in a plane needs to be 100% reliable, or everyone dies. A car needs to be, I dunno, 80%? Compared to a bad human driver, who still drives every damn day, a computer need only be about 60% reliable to be better. People suck at driving. Even a shitty self-driving car will save a ton of lives simply by obeying traffic laws.
- dreamcompiler 10y agoAn auto-pilot for an airplane is a considerably easier problem to solve. No lanes; no pedestrians; very little other traffic; three spatial degrees of freedom. That's why auto-pilots for airplanes have existed for almost a century but we're just now beginning to get self-driving cars. Humans are still better at dealing with the full panoply of crap that road driving throws at us.
- infinotize 10y agoAircraft autopilots also rely on experienced and licensed pilots to operate them and be responsible for the aircraft at all times. Self driving cars have assume the operator is not particularly capable nor paying attention to anything happening on the road.
- duaneb 10y agoDo they rely on the pilot? I was under the impression it was entirely hands off.
- durandal1 10y agoThere are many conditions unders which an aircraft autopilot will simply disconnect without previous warning and hand over controls to the pilot.
- Eugr 10y agoCurrent generation of autopilots doesn't handle traffic avoidance or make any routing decisions - they just follow pre-programmed routes at pre-programmed speed and altitude (or climb/descent profile). But even this relatively simple level of automation causes problems - pilots start to rely on automation too much, and when things go south they are not capable to deal with it. Airlines recognize it, and put more emphasis on hand-flying during training and routine operations, so pilots don't lose their basic piloting skills. It's not a new problem - there is an excellent training video from 1997 - "Children of the Magenta": https://www.youtube.com/watch?v=pN41LvuSz10 https://www.youtube.com/watch?v=pN41LvuSz10
- schwarrrtz 10y agoThis video also inspired an excellent podcast from 99 Percent Invisible about the challenges and dangers of automation. I would highly recommend listening. http://99percentinvisible.org/episode/children-of-the-magenta-automation-paradox-pt-1/ http://99percentinvisible.org/episode/children-of-the-magent...
- 10y ago
- boydc 10y agoSelf-driving car could be better than human in average. But as long as there are human drivers who drive better than self-driving software, it would be disaster for these drivers. We definitely do not want some technique than do good for majority but do horrible things for minority, right?
- scarecrowbob 10y agoI;m not following your argument here. A single driver's ability isn't the only risk factor... if I'm a great driver but every one else sucks (that's how it for everyone already, right :D ), then an overall increase in the population's driving ability helps me, right?
- boydc 10y agothe overall increase helps you indeed. But do you want use self-driving software if you are a great driver(or you think you are)? I do not because I want to be more safer by driving myself. If great drivers like to drive themselves. Others wants too because they do not trust these great drivers. In everyone driver's eyes, there are only two kinds of drivers 1 ) bad driver slower than me. 2 ) mad driver faster than me.
- jchrisa 10y agoWhy would it be a disaster?
- llamataboot 10y agoThis would be true only if your driving ability only affected your chance to die, but your driving ability has an effect on everyone else's safety on the road as well!
- boydc 10y agoConsider you are a damn good driver, better than self-driving software. If self-driving software can reduce your risk by giving you a safer environment(by replacing lots of bad drivers), but it will increase your risk when handling risks(because it's not as good as you). Would you like to choose self-driving software? The point here is, no matter how good the driving environment goes, I do not want to lost any chance to survive(If I'm a good driver).
- zdkl 10y agoNot sure low-probability/high-damage events are comparable to high-prob/"low"-dmg in the first place and that's not a trivial question to handle in real-time
- ceejayoz 10y ago> Yeah, remember, auto-pilot in a plane needs to be 100% reliable, or everyone dies. First, they're not anywhere near 100% reliable. They can fail on their own, and they'll also intentionally shut themselves off if the instruments they rely on fail. https://en.wikipedia.org/wiki/Air_France_Flight_447 https://en.wikipedia.org/wiki/Air_France_Flight_447 Second, an autopilot failure shouldn't lead to death if the pilots are competent and paying attention.
- duaneb 10y agoWhat does a failed autopilot look like? Would a pilot do any better with an "aerodynamic stall"? I know little about planes and it seems like that'd be a big problem with or without a pilot driving.
- ceejayoz 10y agoAll pilots are trained to recover from a stall. A failed autopilot could look like all sorts of things, from just automatically disconnecting itself (usually with a loud warning alert) to issuing incorrect instructions (which is why the pilots are supposed to be awake and alert while it's engaged, watching the instruments).
- neurotech1 10y agoActually a key finding in AF447 was that pilots were not trained on how to recognize and recover from a high altitude stall. It is not like flying a Cessna 150. The junior first officer didn't realize the aircraft had stalled. Pilots were trained on the procedure for recovering from a low altitude stall; 100% or TOGA thrust and power out of it while minimizing altitude loss. Training has now changed for both low altitude and high altitude stall recovery.
- ceejayoz 10y agoIIRC, part of the issue was that the two pilots issued contradictory joystick commands, which the plane averaged to zero. Which is a bit terrifying.
- duaneb 10y agoBut autopilot for planes is actually much easier than negotiating traffic with irrational humans with road rage. You can coordinate with air traffic for takeoffs and landings, and there is very little to run into at tens of thousands of feet.
- raverbashing 10y ago> auto-pilot in a plane needs to be 100% reliable, or everyone dies Actually it is much more simpler than a self-driving car. And if there is a problem it disengages.
- vacri 10y agoAutopilot in a plane usually doesn't involve autonavigation, whereas autodrive in a car generally requires navigation. Autodrive in a car without navigation is basically 'cruise control'.
- dreamcompiler 10y ago'Cruise control' does not (at least it didn't until very recently) even attempt to avoid collisions with neighboring cars or keep the car in its lane. Even without navigation, autodrive in a car is a considerably more difficult problem than either cruise control or autopilot in a plane.
- vacri 10y agoMy point is that 'autopilot' is more like 'cruise control', and except in advanced cases, is not analagous to 'autodrive'.
- brianwawok 10y agoI know this logically. But emotionally I know how many bugs I have written in my life. I know software devs are human.. aka I know how the sausage is made.
- kohanz 10y agoYes, but have you been part of the development & testing effort for mission-critical software (e.g. a class 1 or 2 medical device?). It's not true in all cases, but for the most part the level of QA that goes into the devices before release is significantly higher than that of your average product. This is why regulation is required.
- deleted 10y ago[deleted]
- praxulus 10y agoGCE downtime just means people lose money, it's not life-or-death. Skimping on QA in order to reduce costs and get to market faster is a perfectly reasonable decision when the consequences are so mundane.
- flurdy 10y agoI understand what you mean but that is generalising too much what people use GCE, public clouds, self hosted servers for, and especially going forward. It is not all convenience applications, game backends etc. What people these days use AWS/GCE for is so varied, even public sector use AWS Gov Region for example. Downtime consequences is not just money lost but can be life-and-death and for some application they need solid QA even if hosted in a public cloud. It may (emphasise 'may') be how they share medical data via GCE/AWS that gets delayed just before a surgery (ok, edge case) or how they update bugs in a critical GPS model that happen to be used by an ambulance, or even a taxi used by pregnant lady that is about to drop, etc. Or simple general medical self diagnosis information site that by chance could have saved someone in that time slot. Or any other random non medical usage which involves a server and data of some kind that happen to be in GCE. Yes critical real time systems often are on-premise or in self hosted data centres, but more and more are not especially if viewed as not critical but in some cases indirectly are critical.
- ambago 10y agoEspecially since "driver-error" is the cause of 94% of motor vehicle crashes in the U.S.[1], with 32,675 people killed and 2.3 million injured in 2014.[2] Worldwide, motor-vehicle crashes cause over 1.2 million deaths each year and are the leading cause of death for people between the ages of 15-29 years old.[3] It's estimated that self-driving cars could reduce vehicle crashes by approximately 90%! [4] [1] http://www-nrd.nhtsa.dot.gov/pubs/812115.pdf http://www-nrd.nhtsa.dot.gov/pubs/812115.pdf [2] http://www-nrd.nhtsa.dot.gov/Pubs/812219.pdf http://www-nrd.nhtsa.dot.gov/Pubs/812219.pdf [3] http://www.who.int/violence_injury_prevention/road_safety_status/2015/en/ http://www.who.int/violence_injury_prevention/road_safety_st... [4] http://www.mckinsey.com/industries/automotive-and-assembly/our-insights/ten-ways-autonomous-driving-could-redefine-the-automotive-world http://www.mckinsey.com/industries/automotive-and-assembly/o...
- takeda 10y agoThis is true for a single car, but self driving cars are introducing something that did not happen before. Imagine majority of cars are self driving cars and they all malfunctions due to bug in software update.
- executesorder66 10y ago> and they all malfunctions due to bug in software update. That's assuming everyone with a self driving car is driving the exact same model and they all updated at the exact same time. Chances are there will be many different models and manufactures so an OTA update with a bug will only affect a much smaller percentage of the self driving cars.
- nzoschke 10y agoSeconded. Fast recovery of the problem, fast to publish a postmortem, and a very thurough postmortem. Outages suck, but are inevitable even for Google. With a response like this Google has gained even more trust from me.
- kyrra 10y agoGoogle's take on postmortems is really nice. As the SRE book points out, they are seen as a learning tool for others. Most internal postmortems are available for anyone within the company to see and learn from. As well, they are always blameless. No fingers are pointed at the person who caused the issue in the postmortem. They explain the issue, what happened, and how it can be prevented in the future. Pair this with the outage tracking tools and you can find all the outages that have happened across Google and what caused them. Then there is DiRT[0] testing to try and catch problems in a controlled manner. Having things break randomly through Google's infrastructure and you have to see if your service's setup and oncall people handle it properly is a really awesome exercise. [0] http://queue.acm.org/detail.cfm?id=2371516 http://queue.acm.org/detail.cfm?id=2371516 The opinions stated here are my own, not necessarily those of Google. Edit: Changed from saying "all" to "most" postmortems being available to Googlers to see.
- Thaxll 10y agoself driving cars will be safer than BGP that' s for sure.
- clebio 10y agoOh geez. This comment needs to go into the pantheon of all-time nerdy humor.
- ddispaltro 10y agoAgreed, I think a good postmortem distinguishes great companies from just good companies. The depth on philosophy, reasoning and then action is very digestible.
- fweespee_ch 10y ago> This is also why I am of afraid of a self driving cars and other such life critical software. There are going to be weird edge cases, what prevents you from reaching them? Yes. However, the current failure rate of human drivers being improved on is the standard I care about. http://www.cnbc.com/2015/10/29/crash-data-for-self-driving-cars-may-not-tell-whole-story.html http://www.cnbc.com/2015/10/29/crash-data-for-self-driving-c... > After crunching the data, Schoettle and Sivak concluded there's an average of 9.1 crashes involving self-driving vehicles per million miles traveled. That's more than double the rate of 4.1 crashes per one million miles involving conventional vehicles. That is the only number that matters to me. Google gets that to 4.0 per million miles and I'd say they are good to go.
- brianwawok 10y agoSo crashes like that include drunk drivers, drugged drivers, and texting drivers. What is the crashes per miles for a paying attention driver? If it is 1 per million miles, the self driving car would need to be a lot lower. Now if it was 4am and I am falling asleep at the wheel, I bet any self driving car would beat me. So cool to turn on, but maybe not for a daytime cruise...
- bcook 10y agoWhy are you comparing self-driving cars to exclusively a "paying attention driver"? For self-driving cars to be safer than human drivers, there is no requirement that the self-driving cars should be better/safer than the best human driver... the self-driving car simply needs to be safer than the majority of humans.
- ams6110 10y agoIt points out that a lot depends on the driver. I drive attentively, and moreover I like driving. I would never want a self-driving car for myself. However some seem to think that should be the only choice because they can beat accident averages that include drivers who drink, text, do makeup, masturbate, whatever while driving.
- 10y ago
- jfoster 10y agoAs long as the edge case bugs in self driving cars come up less frequently than human error, it's an overall improvement.
- ocdtrekkie 10y agoCurrently, they don't. Google cars fail every 1,500 miles on average.
- 8note 10y agodo you have a comparison for the number of miles between human errors?
- ocdtrekkie 10y agoIt's hard to get an exact figure, particularly because of unreported accidents, and various sources. But I believe insurance companies have previously stated it's about one in every 250,000 miles. For the sake of giving a wide berth for unreported accidents, and to not give humans the benefit of the doubt, I've been using the rough figure of 150,000 miles between accidents. I don't have a great source for it though, and if anyone finds a good source, it'd be fantastic.
- jfoster 10y agoWhere is that figure coming from? I'm curious what kinds of failures they have.
- ocdtrekkie 10y agohttps://news.ycombinator.com/item?id=11492569 https://news.ycombinator.com/item?id=11492569 <- I detailed it a bit more in this comment, including linking the source report from Google.
- jfoster 10y agoI think when you say they "fail", I think you are referring to the disengagements, right? The report you link to there says: “Immediate manual control” disengage thresholds are set conservatively. Our objective is not to minimize disengages; rather, it is to gather as much data as possible to enable us to improve our self-driving system. Also, table 4 reports the number of disengagements (for any reason) each month, as well as the miles driven each month. In the most recent month in that table, it's actually 16 disengagements over 43275.9 miles. That's approximately one disengagement every 2705 miles; about the distance from Sacramento, CA to Washington, DC. At the start of 2015 it was only 343 miles per disengagement; 53 disengagements over 18192.1 miles. The pace of improvement is incredible, especially considering disengagements are set conservatively. Can a human drive from Sacramento CA to Washington DC without a single close call or mistake along the way? I really doubt it. This technology will be saving lives soon.
- nxzero 10y agoIdea that edge cases in autonomous vehicles would result in 30,000+ deaths a year to me is a stretch. If you dispute this, please explain. If your position is that one death is too many, that is illogical relative to the option of letting people drive cars.
- ocdtrekkie 10y agoCurrently, the self-driving software fails out on a Google Self-Driving Car every 1,500 miles. If the car suddenly stops trying to drive in the road, and the driver isn't attentive (or worse, if Google gets their way and convinces the laws to change so they don't have to have steering wheels) that's a lot of deaths. I'm not saying it won't get better, but pretending self-driving cars is a cure-all right now is hilarious and insane.
- gizmo385 10y agoSource? What kind of failure are you talking about? Minor hiccups or full failures which stop the car entirely?
- ocdtrekkie 10y agoGoogle's report from December 2015: http://static.googleusercontent.com/media/www.google.com/en//selfdrivingcar/files/reports/report-annual-15.pdf http://static.googleusercontent.com/media/www.google.com/en/... Over 424,000 miles driven: 272 times the car had a 'system failure' and immediately returned control to the driver with only a couple seconds of warning. (Approx. every 1,558 miles.) A car mid-traffic spontaneously dropping control of the vehicle would likely create a large number of accidents. 13 car accidents prevented via human intervention (Approx. every 32,615 miles), 10 of which would've been the self-driving car's at-fault (Approx. every 42,400 miles). These virtual accidents were tested with the telemetry recorded during the incident, and it was determined had the human test driver not intervened, an accident would've occurred. Total of these events is 285, which is approximately every 1,487 miles driven. For useful comparison, a rough human average (when you add a large margin to account for unreported accidents) is somewhere around one accident every 150,000 miles driven. (Insurance companies see them every 250,000 miles approximately, I believe.)
- ksou32 10y agoHumans have bugs all the time, you can just faint for no reason while driving . ..
- llamataboot 10y agoMy consolation in that fact is that weird edge cases happen with human driven cars as well. Someone has a seizure and crashes, or more commonly reaches for a cigarette, the radio, their phone. People hit ice or water and overcorrect their spin. People drive too fast. Etc etc etc. Not even all edge cases, many common modes of failure. I except self-driving cars that kill people will be a huge emotional issue for a lot of people in accepting them, but for me, i just want them to be safer than human drivers, which isn't THAT high of a bar to cross.
- fizzbatter 10y agoYup. People seem to be overly critical with automated car failures. Personally, i think automated cars are going to easily be better than humans in the working cases (both human and ai are concious). Next, i expect to see fully operational backup systems. Eg, if a monitoring system decides that the primary system is failing for whatever reason, be it bug or unhandled road condition (tree/etc), the backup system takes over and solely attempts to get the driver off the road, and into a safe location. Humans often fail, but often can attempt to recover. And, as bad as we may be at even recovering, we know to try and avoid oncoming traffic. Computers (currently) are very bad at recovering when they fail. I feel like having a computer driving, in the event of failure, is akin to a narcoleptic driver - when it goes wrong, it goes really wrong. Hence why i hope to see a backup system, completely isolated, and fully intent on safely changing lanes or finding a suitable area to pull over.
- nostalgiac 10y agoSounds good in theory. Until the bug that causes failure is also present in the monitoring system, and as such doesn't fail over to the backup system. AKA exactly what happened here to Google.
- msellout 10y agoHumans also make mistakes in corner cases. I'm no more afraid of an auto-steering car than a human-steered car.
- ben_jones 10y agoSelf driving cars don't have to be perfect. They just have to be safer then driving is today [1]. The real question is if society can handle the unfairness that is death by random software error vs. death by negligent driving. It's easy to blame negligent driving on the driver, we're clearly not negligent so it really doesn't effect us right? But a software error might as well be an act of god, it's something that might actually happen to me! [1]: https://en.wikipedia.org/wiki/List_of_motor_vehicle_deaths_in_U.S._by_year https://en.wikipedia.org/wiki/List_of_motor_vehicle_deaths_i...
- llamataboot 10y agoNo because negligent driving doesn't just put the driver at risk - it puts everyone else on the road plus pedestrians at risk
- wefarrell 10y agoExactly, negligent drivers don't just kill themselves and there is very little you can do to prevent one from killing you.
- 794CD01 10y agoThey don't have to be safer than driving is today. They can be significantly less safe while still being an improvement for society because drivers will be able to focus on other activities while travelling instead of wasting that time focusing on driving the car.
- Grambo 10y agoActually they have to be significantly safer than driving today. People would rather be unsafe and in control than not in control and a tiny bit safer. I know personally if a self driving car could only drive as well as I could then I'd still want to be the one driving.
- oldmanjay 10y ago
- stcredzero 10y agoThis is also why I am of afraid of a self driving cars and other such life critical software. There are going to be weird edge cases, what prevents you from reaching them? Formal systems?
- tremon 10y agoFormal systems are still built on a model of the outside world, not on the world itself FAFAIK. Even if your formal coverage is 100%, you can not anticipate all weird edge cases the real world can come up with.
- eloff 10y agoYes but how many people drive stoned, drunk, or distracted? How many people drive aggressively, speeding, or erratically? How many people do dumb things on the road? As a software engineer I know that there will be bugs and some will likely kill people. But as a driver who has driven many years in less civilized countries, I know that human beings are terrible drivers. Who would you rather share the road with, computer drivers that drive like your grandma, or a bunch of humans? It's a no-brainer right?
- pveierland 10y agoIt also showcases the great thing about self-driving cars. Even though accidents will happen, when it does there will be plenty of sensor data and logs which can be examined to find the exact cause in a post-mortem. An improvement to the software can then be made, and millions of cars deployed can all effectively learn from a single accident. With humans, the amount of knowledge gained and the collective improvement of driving behavior from a single accident is low, and each accident mostly provides some data points to tracked statistics. With machines, great systematic improvements are made possible over time such that the remaining edge cases will become increasingly improbable.
- marcosdumay 10y agoJust like planes. I'll have to point that this is necessary, but not sufficient for enabling an ever improving, extremely safe activity. Aviation also have a just right amount of blame running in the system that is hard to replicate on any other area.
- xiphias 10y agoWhat does post-mortem mean in this context? The software one (after an accident) or the human one (after death)? I think It's crazy that the word gets back the original meaning..
- yeukhon 10y agoAirplanes are equipped with software and many pilots would turn on auto-pilots after a long take off. Bugs are everywhere, and it's just a matter of time before one is so critical and kill people. So our best bet is better quality assurance through proof and overtesting (do this incrementally!)
- Dylan16807 10y ago> As I assumed it was kind of a corner case bug meet corner case bug met corner case bug. I wouldn't really agree with that. There were two pieces of code designed to perform checks on new configs and cancel them. They both failed. Neither of those checks is a corner case. If you had a spec sheet for the system that manages IP blocks, that functionality would be listed as a feature right up front.
- godgod 10y agoImagine the day when the software powering that self driving car chooses to avoid the pedestrian by driving your car off a cliff. Software can kill.
- Donzo 10y agoA 1-in-one-million occurrence will happen a thousand times each day when you perform a billion operations.
- erichocean 10y agoAre BGP updates for Google's own router configurations really so frequent that they can't pay an engineer to at least monitor the propagation of configuration changes? In this case, a human would have instantly seen that the update was a) rejected (as explained in the postmortem), and b) holy shit, WHY DID THE ROUTER CHANGE ITS OWN CONFIGURATION TO BLOW AWAY ALL OF THE GCE ROUTES!?! I'm all for automation, but WTF? Insert even a semi-competent engineer in the loop to monitor the configuration change as it propagates around and the entire problem could have been addressed almost trivially, as the human engineers eventually decided to do.
- jvolkman 10y agoThe human part of any process will also eventually fail, and it's much more difficult to fix human bugs. Better to shoot for full automation.
- heisnotanalien 10y agoBut a computer could do the same thing. It would be possible to alert on all the routes being blown away?
- DanielDent 10y agoAny sufficiently large system quickly reaches a point where a human has difficulty tracking what the system should look like. Google has at least tens of data center locations, each of which will have multiple physical failure domains. There are also many discontiguous routes being announced at all of their network PoPs. They have substantially more PoPs than data centers. It very quickly gets too much to reasonably expect people to be able to keep track of what the system should look like, let alone grasping what it does look like.
- dsl 10y agoFirst of all, BGP is core to Google's load balancing architecture. So within a single datacenter you probably have at least a few dozen devices down stream from each edge router. Secondly, I'm seeing just shy of 500 individual prefixes, 282 directly connected peers (other networks), and a presence at over 100 physical internet exchanges, just for one of Google's four ASes. Would you be able to read over that configuration and tell me if it has errors?
- mikx007 10y agoAssuming bugs are never intentional and mostly random... Maybe instead of one autopilot software, self driving cars of the future will have several, developed by completely different teams. Then a self driving car can take some sort of average or most common output instruction (thus minimizing the risk of random bugs/edge cases...etc.)
- profeta 10y agowhy do you think that will wait for cars? https://en.wikipedia.org/wiki/Therac-25 https://en.wikipedia.org/wiki/Therac-25
- draw_down 10y agoWell, humans driving cars is already a disaster, so.
- lugg 10y agoThis wasn't an edge case. It was two bugs in two sections of code both designed to recover from a serious problem. It sounds like both sections of code were not tested properly at the very least. Sounds to me like someone just didn't bother to test the failsafe part of the code.
- packetslave 10y ago...and you're basing this on what, exactly? It's easy to pontificate about what people "didn't bother to test" based on zero information.
- lugg 10y agoFirst bug: In a failure case, it should remove the failing config, not all of them. Pretty hard thing to miss if you test for it with any level of basic unit test or similar. Second bug: canary failure should prevent further propogation of the bad config. A little more difficult to test with automated tests due to requiring a connection. It sounds like this was in fact tested, but the usage between the two bits of software was not tested. A good integration test would have caught this. But I wouldn't call that required. I would at least however think it was required that the use case of that particular code to be at least manually checked because, you know it's a feature for disaster prevention / recovery. There was enough information to deduce this pretty easily. Although they did tend to glaze over it in the write-up, almost purposefully. For all those spouting that this was a good postmortem, not really, it's a good covering of ones ass, a good spin, sidestepping the real root cause. What has slas and "here take credits" got to do with a postmortem? I'm not really sure why I got downvoted for this. The post mortem was good but it wasn't something I'd aim to strive for. I like gcloud and I'll keep using it but I find the response to this thing a little bit hard to swallow.
- magicalist 10y ago> I'm not really sure why I got downvoted for this. Because you have an apparently incredibly simple mental model for the system and so of course tests for it seem simple?
- robmcm 10y agoThe key thing to remember is a bug in autonomous driving doesn't mean the car swerves off the road at 100mph. If the software crashes or fails the car can come to a stop quite quickly without harm and allowing for human intervention. Having said that I am still scared, I'm not sure how well Tesla auto pilot will handle a tire blowout at 70mph. Perhaps better than I would, but I would much rather I was in control.
- wyldfire 10y ago> . Internal monitors generated dozens of alerts in the seconds after the traffic loss became visible at 19:08 ... revert the most recent configuration changes ... the time from detection to decision to revert to the end of the outage was thus just 18 minutes. It's certainly good that they detected it as fast as they did. But I wonder if the fix time could be improved upon? Was the majority of that time spent discussing the corrective action to be taken? Or does it take that much time to replicate the fix?
- toomuchtodo 10y ago> But I wonder if the fix time could be improved upon? Rushing to enact a solution can sometimes exacerbate the problem.
- Sanddancer 10y agoFrom the rest of the post, it sounds like replication time. Datacenters started dropping an hour beforehand one by one, and they had all fallen over by 19:08. Given that you have to push the rollback to routers around the world, and that peer routers have to propagate the changes from there, 18 minutes for a change like this sounds about right.
- jlgaddis 10y ago... although once the first datacenter once again announced the prefixes into BGP, those networks would have been reachable again, from everywhere. I imagine this is what happened at 19:27 -- the first datacenter came back online. Of course, the traffic load might have overwhelmed that single datacenter but that would be alleviated as soon as additional datacenters came back online ("announced the prefixes"). A portion of the traffic load would shift to each new datacenter as it came back online. It could have been hours later before they were all operational again but, as far as the users were concerned, the service was up and running and back to normal as soon as the first one or two datacenters came back up.
- VLM 10y agoHaving worked in ISP operations on BGP stuff (admittedly more than 10 years ago), it was both too slow and too fast. If the rollout took 12 hours instead of 4 or the VPN failure to total failure was multiple hours instead of minutes, they'd have had enough time to noodle it out. Eventually at a slow enough deploy rate they'd have figured it out. It only took 18 hours to make the final report after all, so an even slower 24 hour deploy would have been slow enough, if enough resources were allocated. On the opposite side, most of the time when you screw up routing the punishment is extremely brutal and fast. If the whole thing croaked in five minutes, "OK who hit enter within the last ten minutes..." and five minutes later its all undone. What happened instead was dude hit enter, all is well hours later although average latency was increasing very slowly as anycast sites shut down. Maybe there's even shift change in the middle. Finally hours later it finally all hit the fan meanwhile the guy who hit enter is thinking "it can't be me, I hit enter over four hours ago followed by three hours of normal operation... must be someone else's change or a memory leak or novel cyberattack or ..." Theoretically if you're going to deploy anycast you could deploy a monitoring tool to traceroute to see that each site is up, however you deploy anycast precisely so that it never drops... Its the titanic effect, why this is unsinkable, why would you bother checking to see if its sinking? And just like the titanic if you break em all in the same accident, that sucker is eventually going down, even if it takes hours to sink.
- ikeboy 10y ago>However, in this instance a previously-unseen software bug was triggered, and instead of retaining the previous known good configuration, the management software instead removed all GCE IP blocks from the new configuration and began to push this new, incomplete configuration to the network. >Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout. I assume the software was originally tested to make sure it works in case of failure. It would be interesting to know exactly what the bug was and why it didn't show in tests.
- djfergus 10y agoNetwork management software complexity is supposed to be one of things that SDN was built to solve (by introducing more modularity and defined interfaces). But in this case the fault was at the edge with BGP route updates, which the internet has been doing for decades. I share your curiosity in the specific bug. However, this is a great detailed post-mortem from a service provider. Your Telco or ISP will never provide this much detail...
- balls187 10y agoNice post mortem. That outtage gives GCE at best a four 9's reliability for 2016.
- daveguy 10y agoBased on the higher level status page: https://status.cloud.google.com/summary https://status.cloud.google.com/summary It looks like GCE uptime is well below four 9's reliability for a sliding 1 year timeframe.
- dgacmu 10y agoTraynor was quoted in a networkworld article last year saying they aim for three and a half nines (99.95%). But you need to read into the incidents more carefully -- figuring out actual "uptime" is quite hard. Consider the longest-lasting incident: "On Tuesday 23 February 2016, for a duration of 10 hours and 6 minutes, 7.8% of Google Compute Engine projects had reduced quotas. ... Any resources that were already created were unaffected by this issue." I'm not sure off the top of my head how I'd try to compute the overall availability #s from that one. One can possibly try to determine and sum the effects on the individual customers, but we can't from the information provided. But it's certainly less overall downtime than just counting it as a 7 hour failure.
- daveguy 10y agoAgreed. It is difficult to tell. But if the bug is preventing you from processing (because you can't save the existing results) then it's essentially down time for new processing. There are also connectivity issues by region and DNS issues. It is difficult to get exact downtime considering partial failures. That said, this is the second major asia-east1 downtime in 90 days: https://status.cloud.google.com/incident/compute/16002 https://status.cloud.google.com/incident/compute/16002
- balls187 10y agoApril's incident is unique, This was the only case (listed) that was a service outtage, which impacted all of GCE. The other incidents (as far as I can tell), were service disruptions at the AZ/regional level. Those disruptions don't impact the 9's, as GCE was available for other regions.
- obulpathi 10y ago> Finally, to underscore how seriously we are taking this event, we are offering GCE and VPN service credits to all impacted GCP applications equal to (respectively) 10% and 25% of their monthly charges for GCE and VPN. These credits exceed what is promised by Google Cloud in their SLA's for Compute Engine and VPN service!
- duskwuff 10y ago... which is precisely (almost word-for-word) what the post-mortem goes on to say. Is there something specific you're trying to call attention to here?
- deleted 10y ago[deleted]
- obulpathi 10y agoNop. Probably did too much copy-pasting :( Mearly wanted to highlight the point.
- icebraining 10y agoOnly barely. They're down to 2.5 minutes of downtime left for the next 30 days if they want to keep the 99.95% level.
- platz 10y ago> configuration file configuration files strike again - remember knight capital?
- pbreit 10y agoDo SLAs even matter in the slightest? Or are they just sort of "feel-good" things or ways for negotiators to demonstrate their worth?
- duskwuff 10y agoSLAs aren't about guaranteeing uptime. They're about setting consequences for downtime.
- cbr 10y agoBut once there are strong consequences for downtime the service provider is going to set up training, monitoring, oncall, etc to make sure things stay within the SLA limits. So you are effectively negotiating uptime.
- qaq 10y agoThe only SLAs that matter are the ones where service provider will suffer serious $ penalties on braking the SLA. Which rules out basically all major cloud providers that will simply issue credit for the downtime.
- cjbprime 10y agoIt looks like there were at least three catastrophic bugs present: 1. Evaluated a configuration change before the change had finished syncing across all configuration files, resulting in rejecting the change. 2. So it tried to reject the change, but actually just deleted everything instead. 3. Something was supposed to catch changes that break everything, and it detected that everything was broken but its attempt to do anything to fix it failed. It is hard to imagine that this system has good test coverage.
- mjibson 10y agoI'm attempting to even imagine how one would build a useful way to test this. Would they have to have a secondary, world-wide datacenter network with all their various services behind it?
- ikeboy 10y agoYou could have it send messages to the actual servers, but with an added flag that says "fake", which makes the servers ignore the message/send back a message saying pass/fail/whatever (testing the flag could happen first, one server at a time manually). Then check whether the program continued to push updates.
- maxander 10y agoYou may be able to build an elaborate system of dummy network operations to test with, but this system may wind up with bugs that mask what would be errors in the real system. And how to you test against that? A dummy network to test the dummy network operations on? What if the dummy network contains bugs that make it behave significantly different from the real network, in error cases? How do you test for that? Its turtles all the way down!
- ikeboy 10y agoYou can test it against the actual network; if something goes wrong, you'll have downtime, but you'll be prepared to get it all back up. Or, to test whether the "prevent errors from going to new places" works, temporarily configure the new places to ignore new configs; if the system works, no messages will be sent there; if the system doesn't work, they ignore the message and you learn about a bug.
- qaq 10y agoDRY "The inconsistency was triggered by a timing quirk in the IP block removal - the IP block had been removed from one configuration file, but this change had not yet propagated to a second configuration file also used in network configuration management."
- deleted 10y ago[deleted]
- fixermark 10y agoDRY is tougher when for practical reasons data must be physically cached locally.
- senderista 10y agoYes, DNS was clearly designed by idiots who had never heard of DRY.
- teraflop 10y ago> There are a number of lessons to be learned from this event -- for example, that the safeguard of a progressive rollout can be undone by a system designed to mask partial failures -- ... This is a really important point that should be more generally known. To quote Google's own "Paxos Made Live" paper, from 2007: > In closing we point out a challenge that we faced in testing our system for which we have no systematic solution. By their very nature, fault-tolerant systems try to mask problems. Thus they can mask bugs or configuration problems while insidiously lowering their own fault-tolerance. As developers we can try to bear this principle in mind, but as Monday's incident demonstrated, mistakes can still happen. So, has anyone managed to make progress toward a "systematic solution" in the last 9 years?
- atomic77 10y agoThis is an interesting question, and seems to get to the core of Nassim Taleb's ideas [1] about fragility and the limits of what we can understand, and how many of our attempts to create artificial stability ultimately bring about the opposite. That said, based on this post-mortem, I think Google, and our industry as a whole, is doing a pretty good job. Periodic failures like this are inevitable, and if they serve to make it less likely that a similar failure occurs in the future, then that is a system as a whole that could be described as "anti-fragile". [1] At least my interpretation of them
- woodman 10y ago> So, has anyone managed to make progress toward a "systematic solution" in the last 9 years? That depends on how you define "solution". If development time isn't a concern, then formal verification is a pretty solid solution. AWS has used TLA+ on a subset of its systems. [0] [0] https://en.wikipedia.org/wiki/TLA%2B https://en.wikipedia.org/wiki/TLA%2B
- deleted 10y ago[deleted]
- strictfp 10y agoDegraded modes of operation is one example of how to visualize masked errors. Another is to trigger an alarm on fallbacks. As a general reflection, many distributed system leave out the cause of their changes and only log actions. Instead of logging "new membership, new members are b,c,d" you are better of logging "node a has not responded to heartbeat in the last 30 seconds, considering it faulty". Following such a principle makes it much easier to spot masked bugs, since you can reason about the behaviour much better. Aggregating logs to a central location and being able to analyze global behaviour in retrospect is also a great feature.
- eranation 10y agoThis is very interesting. From the little I understand (sorry for using AWS terms as I am more versed with AWS than GCE) this can happen to AWS as well right? even if your software is deployed to multiple AZs / multiple regions, if bad routing / network configuration makes it through the various protection mechanisms then basically no amount of redundancy can help if your service is part of the non functional IP block. I mean it seems no matter how redundant you are, there will always be somewhere along the line a single point of failure, even if it has multiple mechanism to prevent it from happening, if all of these mechanisms fail, then it's still a single point. What prevents this from happening at Azure / AWS? Is there anything that general internet routing protocols need to change to prevent it from happening? e.g. I'm sure that we will never hear that Bank of X has transferred a billion dollar to an account but because of propagation errors it published only the credit but didn't finish the debit and now we have two billionaires. This two or more phase commit is pretty much bulletproof in banking as far as I know, and banks are not known to be technologically more advanced than Google, how come internet routing is so prone to errors that can an entire cloud service unavailable for even a small period of time? I'm far from knowing much about networking (although I took some graduate networking courses, I still feel I know practically nothing about it...) So I would appreciate if someone versed in this ELI5 whether it can happen in AWS and Azure regardless of how redundant you are, (which leads to a notion of cross cloud provider redundancy which I'm sure is used in some places) and whether the banking analogy is fair and relevant, and if there are any RFCs to make world-blackout routing nightmares less likely to happen.
- poooogles 10y agoI'm not sure the AWS network follows the same setup, AWS has very distinct blocks between the US/EU/APAC compared to GCP where you can inherit the same IP if you quickly delete/recreate instances in different regions?
- Swannie 10y agoI was going to post the same comment too. My understanding, from the odd bits and bobs of information I have, is that AWS regions are typically managed somewhat independently.
- pjlegato 10y agoAttention startups: this is what incident post-mortems should look like.
- cosud 10y agoGreat writeup! PS: "To make error is human. To propagate error to all server in automatic way is devops." -DevOps Borat
- deleted 10y ago[deleted]
- deleted 10y ago[deleted]
- awinter-py 10y agochaos monkey?
- hsod 10y ago> Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout. Perhaps the progressive rollout should wait for an affirmative conclusion instead of assuming no news is good news? I'm not being snarky, there may be some reason they don't do this.
- windwake12 10y agoPresumably it received a false positive (or it was interpreted as such). This really seems like the root cause, and I suspect a case of happy path engineering striking again.
- nickysielicki 10y agoWhat does Google use for BGP? Quagga, OpenBGPD, BIRD, their own? Also, does anyone have a link to statistics on global BGP software usage? I'm curious what the marketshare looks like.
- kijiki 10y agoGoogle has contributed ISIS and BGP code to Quagga in the past, as well as funding some testing at the OSRF. Presumably they use it in at least some parts of their operations.
- herrvogel- 10y agoA bit of topic, but it really bugs me, that the banner on the top so pixilated.
- stcredzero 10y agoIn this event, the canary step correctly identified that the new configuration was unsafe. Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout. Classic Two Generals. "No news is good news," generally isn't a good design philosophy for systems designed to detect trouble. How do we know that stealthy ninjas haven't assassinated our sentries? Well, we haven't heard anything wrong...
- avs733 10y agothis was my exact thought...it would seem both feasable and reasonable to have a more active canary process i.e.... anycast "canary test in progress" edge routers store new configs anycast "canary test PASS" edge routers activate new config edge routers canary test new config (and pass or revert) edge routers report home that all is well
- fixermark 10y agoIt may not be good design, but it might be necessary / practical design. If you have enough machines that some percentage of them are down or unreachable at any given time, you can't wait for full go-ahead before proceeding; you'll never get full go-ahead. So you're left with probabilistic solutions, and as T approaches infinity the expectation of more than zero false-positives approaches 1.
- stcredzero 10y agoThe whole point of the canary sub-population, though is that 1) It's not your whole population. 2) You want to find out empirically if something's wrong.
- heisenbit 10y ago"Lessons learned from reading post-mortems" http://danluu.com/postmortem-lessons/ http://danluu.com/postmortem-lessons/ is a good place to dig deeper The first graph quoted from a survey paper is a classic fitting the GCE outage well: Initial error --92%--> Incorrect handling of errors explicitly signaled in software
- anoncept 10y agohttps://mitpress.mit.edu/books/engineering-safer-world https://mitpress.mit.edu/books/engineering-safer-world is also an excellent resource that more people who care about post-mortems should read. (As background, the author, MIT Prof. Nancy Leveson, summarizes decades of work in the field, offers groundbreaking new theoretical tools that scale up to some of the world's most complex accidents, and has the experience and evidence to back up their relevance e.g. via work on Therac-25, the Columbia Space Shuttle, and Deepwater Horizon to name just a few...)
- simonebrunozzi 10y agoI love his signature: "Benjamin Treynor Sloss | VP 24x7".
- mjevans 10y agoI hope that one of their solutions is the obvious one; make change control testing a closed loop instead of an open loop. (Watch for /success/ reported instead of failure notification.)
- swills 10y agoThe thing that stood out for me was: "...team...worked in shifts overnight..."
- delroth 10y ago(Usual disclaimer: I speak for myself, not for my employer, etc.) The team in charge of solving this particular problem is located in two sites in two different timezones. This is true of most critical SRE teams at Google, and it is precisely to be able to have 24h coverage in these time sensitive situations. In the 2+ years I have spent in SRE I have never heard of a single instance of an SRE being asked or even encouraged to stay after hours (let alone overnight) for incident remediation. There is quite a lot of emphasis being put on work/life balance.
- senderista 10y agoWow, that's amazing to read, having served as a de-facto SRE (like every other SDE) at an unnamed competitor to GCE, where I was expected to stay up all night if necessary to resolve an issue (relatively few teams had follow-the-sun coverage). I swore I would never carry a pager again after that, but maybe Google really is different.
- ndesaulniers 10y agoAt Google, they do these really awesome post-mortems when there's a major failure. It provides a point of reflection, and are usually well written entertaining reads. Didn't know they made (some?) public. They're a good learning exercise writing one, and is more of a learning exercise than a punishment.
- advisedwang 10y agoIt's worth noting that the publicly posted postmortem is not the same as the internal postmortems (which include much more detail, specific action items, timelines etc). The SRE book (https://landing.google.com/sre/book.html https://landing.google.com/sre/book.html) has a whole chapter on our internal postmortems, which is probably a better learning exercise in how to write one. Source: I work on the team that writes these external postmortems.
- jpatokal 10y agoGoogle publishes a public incident report for all service outages (code red) in the Cloud status dashboard. You can see some in the History page: https://status.cloud.google.com/summary https://status.cloud.google.com/summary Sample: https://status.cloud.google.com/incident/appengine/16002 https://status.cloud.google.com/incident/appengine/16002 Note that the length of the report tends to correlate with the severity of the outages and that disruptions (code orange) disruptions do not get reports. Disclaimer: I work in Cloud Support and write some of these.
- Gravityloss 10y agoI'm waiting for the time when they push over the air updates to airplanes in flight. "You can fly safely, we have canaries and staged deployment" A year forward: "Unfortunately because the canary verification as well as the staged deployment code was broken, instead of one crash and 300 dead, an update was pushed to all aircraft, which subsequently caused them to crash, killing 70,000 people." I'm not 100% sure why they don't do the staged deployment for google scale server networking over a few days (or even weeks in some cases) instead of a few hours, but I don't know the details here... It's good that they had manually triggerable configuration rollback possibility and a pre-set policy so it was solved so quickly.
- nxzero 10y agoComparing the risk of a live update to a system lives depend on to the risk of some Google services going down is irrational. At some point, delaying the deployment of updates system wide would cause more, not less risks.
- Gravityloss 10y agoThere are businesses that fit somewhere between Boeing and Spotify where failures still have some kind of steeper than casual cost. On Hacker News the "move fast and break things" ethos is probably making sense for many of the people submitting and commenting, since their business is closer to casual usage anyway. But that's not the whole audience.
- nxzero 10y agoShit happens, when it comes to engineering, I'd trust Google more than even likely Boeing to manage systemic risk. As for cars, it's a real risk, but not the same as the bugs Google experienced; I personally have experienced a "bug" driving a car at high speeds, which resulted in a number of major electronic systems failing due to custom systems installed by a well known US startup.
- Gravityloss 10y ago
- huula 10y agoI always like Google's serious attitude towards engineering, even after they have made some mistakes, they never try to hide anything.
- rdtsc 10y ago> However, in this instance a previously-unseen software bug was triggered, and instead of retaining the previous known good configuration, the management software instead removed all GCE IP blocks from the new configuration and began to push this new, incomplete configuration to the network. Always test your crash / exception handling / special case termination+recovery code in production. I have seen this too often. Most often in in "every day" cases when service has a "nice" catch way of stopping and recovering. Then has a separate "if killed by SIGKILL/immediate power failure" crash and recovery. This last bit never gets tests and run in production. One day power failure happens, service restart and tries to recover. Code that almost never runs, now runs and the whole thing goes into an unknown broken state.
- senderista 10y agoSee https://en.wikipedia.org/wiki/Crash-only_software https://en.wikipedia.org/wiki/Crash-only_software
- halayli 10y agoThis isn't the first time a config system at Google causes a major outage. https://googleblog.blogspot.com/2014/01/todays-outage-for-several-google.html https://googleblog.blogspot.com/2014/01/todays-outage-for-se...
- rrdharan 10y agoThat's entirely unsurprising. The recent major Facebook outage was also caused by bad configuration IIRC. See: http://danluu.com/postmortem-lessons/ http://danluu.com/postmortem-lessons/ > Configuration > > Configuration bugs, not code bugs, are the most common cause > I’ve seen of really bad outages. When I looked at publicly available > postmortems, searching for “global outage postmortem” returned > about 50% outages caused by configuration changes. Publicly > available postmortems aren’t a representative sample of all > outages, but a random sampling of postmortem databases also > reveals that config changes are responsible for a disproportionate > fraction of extremely bad outages. As with error handling, I’m > often told that it’s obvious that config changes are scary, but > it’s not so obvious that most companies test and stage config > changes like they do code changes.
- contingencies 10y agoGreat link there! Also check out his list of public postmortems at https://github.com/danluu/post-mortems https://github.com/danluu/post-mortems PS. On HN you should use asterisks to italicize instead of > for quoting.
- trhway 10y agoas devops Borat was saying all along, automated propagation of a error as the main root cause here. A error (new configuration) should be rolled out site by site - ok us-east1, move onto us-west1 ... ok, move onto ... . A canary site may be the first in sequence, yet success ("no failure reported") can't be a big "ok" for automated push to all sites at the same time.
- contingencies 10y agoTLDR; they simply didn't test their (global!) custom route announcement management software. An edge case was triggered in production, and they gee-whiz-automatically went offline. Epic fail. PS. To the downvoters, truth hurts.
- mmel 10y agoI think you're getting downvoted due to the snarky tone more than any "truth" you are stating.
- contingencies 10y agoWell, how to phrase the same thing briefly without sounding snarky?
- Estragon 10y agoYou only need to change a few words: "In other words, they simply didn't test their (global!) custom route announcement management software. An edge case was triggered in production, and unsurprisingly they automatically went offline."
- contingencies 10y agoThere is no accounting for taste.
- chj 10y agoUpvoted. I think they should put a soft version of this right on the first line, instead of burying it in an ocean of "harmless", "previously unseen" text dances.
- totally 10y ago> However, in this instance a previously-unseen software bug was triggered, and instead of retaining the previous known good configuration, the management software instead removed all GCE IP blocks from the new configuration > Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process I'm sure the devil is in the details, but generally speaking, these are 2 instances of critical code that gets exercised infrequently, which is a good place for bugs to hide.
- DanielDent 10y agoMy post yesterday seems even more relevant today: https://news.ycombinator.com/item?id=11477552 https://news.ycombinator.com/item?id=11477552 It's a shame it's not easier or more common for people to create clones of (most|all) of their infrastructure for testing purposes. Something like half of outages are caused by configuration oopsies. If you accept that configuration is code, then you also come to the following disturbing conclusion: the usual test environment for critical network-related code in most environments is the production environment.
- aiiane 10y agoThe main issue there is that "environments" are defined by configuration, so if you try to set up a configuration test environment, you run into a direct logical impass: either your configs are production configs, and thus not a separate environment, or they're different from production configs, and thus may provide different test results from production.
- DanielDent 10y agoWhile I agree with you, I think we could get closer to "production" than is common right now. In an AWS environment, imagine a setup where all that differs is the API keys used (the API keys of the production vs test environment). What gets tricky is dealing with external dependencies, user data, and simulating traffic. For an example more relevant to today's issue: imagine a second simulated "internet" in a globally distributed lab environment. With BGP configs, fake external BGP sessions, etc, servers receiving production traffic, etc. I get that it's a lot of work to setup and would require ongoing work to maintain - and that it's hard/impossible to have it correctly simulate the many nuances of real world traffic - and yet I also think in many cases it would be sufficient to prevent issues from making it into production.
- deleted 10y ago[deleted]
- NetStrikeForce 10y agoI think most people are missing the main failure point: Why does one change propagate automatically to all regions? All this could have been contained if they deployed changes on different regions at different times. That would also help with screwing less your overseas users by running a maintenance at 10am their local time :-)
- aiiane 10y ago> These safeguards include a canary step where the configuration is deployed at a single site and that site is verified to still be working correctly, and a progressive rollout which makes changes to only a fraction of sites at a time, so that a novel failure can be caught at an early stage before it becomes widespread. In this event, the canary step correctly identified that the new configuration was unsafe. Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout. The system does do progressive rollouts, which are essentially what you are referring to (albeit perhaps at a different pace). The number of changes being rolled out means that it's not really feasible to hand roll out configurations to different regions, so the checks are automated. In this case, the automated checks failed as well.
- NetStrikeForce 10y agoI'm not sure you really understand what I've tried to say, but it's probably my fault because of my poor grasp of the English language. You are just confirming my previous comment. Your rollouts are automated, so pushing a change automatically configures every region, instead of configuring just one and maybe waiting for a prudential time in human scale before the next one because, surprise!, shit happens. I understand your colleagues probably make lots of changes, but if that introduces risks of global outages IMHO you should reconsider your strategy. And I'm not sure why you downvoted my previous comment. It's a perfectly valid observation, based on the published information.
- senderista 10y agoWaiting a longer time between regional rollouts (so monitoring systems would have time to detect serious failures) would sacrifice deployment latency, but not deployment throughput (assuming deployments can be made in parallel). For continuous deployment, throughput really matters more than latency.
- itaifrenkel 10y agoWhat is the reason different GCE regions use the same IP blocks?
- sengork 10y agoNetworking issues in either the storage or communication subsystems of any platform normally result in wide-spread disruptions.
- hvass 10y agoWhat is defense in depth? It is mentioned as a core principle.
- koalaman 10y agohttps://en.wikipedia.org/wiki/Defense_in_depth_(computing) https://en.wikipedia.org/wiki/Defense_in_depth_(computing)
- JustUhThought 10y agoJust a thought. Maybe change the name from 'post-mortem' to, anything else before the event actually is a post-mortem.
- zaroth 10y agoFor the amount this cost them, they should have bought CloudFlare. If you play with [global BGP anycast] you are bound to get burned. This is not the first time that BGP took out your entire routing. This is probably not the last time that BPG will take out your entire routing. Whoever's job it was to watch the routing, I am sorry. Pulling your own worldwide routes because you have too much automation; it will make a good story once it's filtered down a bit! Icarus was barely up in the air, too early for a fall.
- grogers 10y agoHow important for redundancy/quality of service is the feature of advertising each region's IP blocks from multiple points in Google's network? It seems like region isolation is the most important quality that Google's network could provide, and their current design is what made something like this possible, not just the bugs in the configuration propagation. They mention the ability of the internet to route around failures, so why not rely on that instead?
- trufflepiggames 10y agoI know what bastard did this shit! http://www.bridezilla-game.com http://www.bridezilla-game.com :O
- Tistel 10y agoThe postmortem used the word "quirk." They might consider drilling down on the specifics there. Especially if that is the heart of the bug/accident.
- dylanz 10y agoCompletely off topic, but this thread is an example of why I (and a lot of people) want collapsible comments native to HN. I'm on my phone, in Safari, and I had to scroll for over 20 seconds just to reach the second comment. The first comment was a tangent about self-driving cars, which while relevant, I didn't want to read about.
- OldSchoolJohnny 10y agoEspecially considering that nearly every post on HN features an often tangential first comment that goes on and on and on...
- naturalethic 10y agoAnyone know why this outage caused all my kubernetes pods to restart?