18 ms·
Google Cloud region currently down due to water intrusion
- kotaKat 3y agoAh, a rainy day in the cloud.
- kotaKat 3y ago"Please don't fulminate. Please don't sneer, including at the rest of the community." @dang, since this got hidden and I can't reply to the toppost. We're allowed a little humor, damnit.
- rickette 3y agoAnyone know how this could affect multiple zones? "Customers can failover to zones in other regions". Unless a whole area got flooded.
- deleted 3y ago[deleted]
- gst 3y agoThere's an interesting Twitter thread about that topic here: https://twitter.com/GergelyOrosz/status/1651256082424012806 https://twitter.com/GergelyOrosz/status/1651256082424012806 Based on that thread it sounds like only AWS guarantees that their AZs are in physically separate DCs, while for Google and Microsoft AZs could be in separate buildings of the same DC facility.
- outworlder 3y agoI would really like to see the physical DC separation at "The Dalles, Oregon".
- blacksmith_tb 3y agoLooks like there are three buildings[1] to me, not entirely sure what goes where, obviously. 1: https://goo.gl/maps/Tfw5UpSsoYiN3YMVA https://goo.gl/maps/Tfw5UpSsoYiN3YMVA
- packetslave 3y agoOr even in the same building, just with a different power/network domain.
- rickette 3y agoAh I see, I know Azure and GCP in NL are in separate buildings but indeed on the same site. But that's not guaranteed for other regions, good to know.
- mcast 3y agoAWS treats its availability zones very seriously, each zone has its own independent power substation, air conditioning, and fiber lines. It's incredibly rare for multiple AZs to go down at once, especially since they are more than a few miles apart from each other.
- sokoloff 3y agoUnless they’re in us-east-1 and it’s an Amazon software/service fault.
- local_crmdgeon 3y agoThis. Don't use us-east-1, it's by far the flakiest. PDX is also a bit rough, but Ohio is golden.
- joelrwilliams1 3y agoOhio has tons of problems, no one should ever put their infra in us-east-2 (shhhhhh...don't let the secret out )
- kevan 3y agoFunnily enough floods (GCP) and fires (OVH) are two of the 3 things AWS explicitly mentions in the Well Architected docs. For a lot of companies an AZ going down is an annoyance or bad day but a whole region going down could be a real continuity risk. > Each Availability Zone is separated by a meaningful physical distance from other zones to avoid correlated failure scenarios due to environmental hazards like fires, floods, and tornadoes. https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_fault_isolation_multiaz_region_system.html https://docs.aws.amazon.com/wellarchitected/latest/reliabili...
- jamesfinlayson 3y ago> but a whole region going down could be a real continuity risk Very much so - Australia only got a second region this year, so if your work required data to remain in Australia, you just had to hope that ap-southeast-2 didn't have a major issues. I'm sure there are plenty of other countries with only a single region.
- xyst 3y agoYou would think that the company that literally wrote the book on “Site Reliability Engineering” would actually follow their own recommendations.
- ddol 3y agoDo Google host their own products on Google Cloud, or are there different sets of data centres for Search/Drive/Gmail vs Google Cloud Customers?
- jeffbee 3y agoThis is a leased facility, the kind of place Google rents for cloud customers but doesn't need for itself. Google's own datacenters are https://www.google.com/about/datacenters/locations/ https://www.google.com/about/datacenters/locations/
- packetslave 3y agoOther way around. Google Cloud runs on the same underlying datacenter, compute, and network infrastructure that Search/Drive/Gmail does. [edit: at least in regions where Google HAS its own datacenters, e.g. "us-central-1? yes. europe-west9? maybe not"]. That does not imply that Search / Drive / Gmail runs on top of Google Cloud.
- kevincox 3y agoThey do internally. But when customers want 3 zones in Indonesia they cut corners.
- CydeWeys 3y agoThe recommendations are to run in multiple regions if you need this kind of redundancy. Run everything in a single region and you can be affected by an event like this.
- londons_explore 3y agoGoogles advice is not to rely on uptime in every region. Instead aim for uptime in a few regions, and load balance your users to regions that are healthy. That design is far cheaper for both google and for you - and, in the typical case, users still get nice low latency to a local datacenter, and only in the rare failure case might they have to wait for latency to some other region.
- JCM9 3y agoYes. Azure and GCPs numbers on the size of their AZs and such are more marketing spin than hard engineering. AWS keeps these in separate physical locations to provide true separation. While there have been tech related regional incidents at AWS a physical event disabling multiple AZs would be extremely unlikely given their much more robust and geographically distributed design. If such a physical event had happened in AWS it would have been a non-event with things just failing over to other AZs. Other cloud providers mostly just vaguely put things in another part of the building and say it’s “a separate AZ” but as GCPs woes highlighted that’s corner cutting that bites badly when the whole building has a problem.
- outworlder 3y ago> If such a physical event had happened in AWS it would have been a non-event with things just failing over to other AZs. In many cases in AWS an availability zone is actually composed of multiple datacenters, each with their own redundancies. This may not be true for smaller regions, but in large ones it definitely is. In those cases, losing an entire datacenter would maybe take out a percentage of instances in that AZ. This has happened before and our production systems barely noticed other than provisioning new nodes to replace the failed health checks.
- kyrra 3y agoGoogler, opinions are my own. I think you misunderstand Google's infrastructure. I'm guessing that each GCP zone is actually a Borg Cell (see: https://storage.googleapis.com/pub-tools-public-publication-data/pdf/43438.pdf https://storage.googleapis.com/pub-tools-public-publication-... ). Borg cells tend to be isolated from eachother in many ways in the physical layer (networking and management being a big one, not sure about power). So networking or machine management for an entire zone could go down and not affect other cells. Changes also tend to get pushed on a per-cell basis when they are Google wide rollouts. I believe GCP recommends to replicate data cross regions (https://cloud.google.com/architecture/framework/reliability/design-scale-high-availability#replicate_data_across_regions_for_disaster_recovery https://cloud.google.com/architecture/framework/reliability/...). Also see: https://cloud.google.com/architecture/disaster-recovery#regions_and_zones https://cloud.google.com/architecture/disaster-recovery#regi...
- yegle 3y agoFor physical zone separation you need to check the `supportsPzs` attribute when listing the zones (e.g. https://cloud.google.com/compute/docs/reference/rest/v1/zones https://cloud.google.com/compute/docs/reference/rest/v1/zone..., but you should be able to find many other places where this attribute is surfaced). It says "reserved for future use" but other docs mentioned "physical zone separation": https://googleapis.dev/java/google-api-services-compute/alpha-rev20210525-1.31.5/com/google/api/services/compute/model/InterconnectLocation.html#setSupportsPzs-java.lang.Boolean- https://googleapis.dev/java/google-api-services-compute/alph...
- outworlder 3y agoRandom datacenters should start advertising availability zones since they should have different fault domains anyway. Google can get away with this, why can't smaller companies?
- migf 3y agoIt amazes me that in every market they serve, Amazon has no actual competitors from a feature perspective. Like, Target does not compete with Amazon. They have a totally different home delivery model that is not in the same category of reliability or service. It's annoying.
- londons_explore 3y agoI think it's because lots of amazons services are in 'winner takes all' markets. No random online eshop can offer next day delivery across half the world unless they already have a logistics chain of 100,000 truck drivers spread across the world. But Amazon can. Likewise, no cloud provider has enough data centers to offer multiple separate data centers in the same city, for hundreds of cities around the world. But Amazon does. Any competitor can't offer amazons level of service until they get to amazon scale... Which they never will.
- deleted 3y ago[deleted]
- manojr13 3y agoLet's the servers cool down for sometime. Might have been working very hard.
- jacquesm 3y agoThis seems to significantly under-report what's going on, see: https://www.theregister.com/2023/04/26/google_cloud_outage/ https://www.theregister.com/2023/04/26/google_cloud_outage/ There is mention of a fire as well.
- jonatron 3y agoThis doesn't sound as bad as OVH's 2021 fire.
- nik736 3y agoWell, we had pictures very quickly of the OVH fire. Google seems to be not very transparent on what is exactly happening...
- stingraycharles 3y agoThe linked article says that there was a leak in a water cooling system, which in turn ended up in the battery system which caused a fire. But yeah it’s not coming from Google but second hand reports.
- mike_hearn 3y agoIt's not second hand, it's from the colo provider they're using in Paris.
- sschueller 3y agoIf you trench a fire in water in a DC it might be just as bad.
- jacquesm 3y agoI wouldn't draw any conclusions just yet.
- DebtDeflation 3y agoPlot twist: the server racks were made out of sodium.
- Jgrubb 3y agoeu-west-9 is Paris
- tpmx 3y ago[flagged]
- CydeWeys 3y ago> Due to software bugs other zones in the same region were also previously down. I don't think it was a software bug. I think they were taken down as a precautionary measure, to not risk flood damage or causing additional fires.
- packetslave 3y agoand you're basing that on... what, exactly?
- outworlder 3y agoThe headline is not incorrect. "Multiple Google Cloud services in the europe-west9 region are impacted. Description: Water intrusion in europe-west9-a has caused a multi-cluster failure and has led to an emergency shutdown of multiple zones. We expect general unavailability of the europe-west9 region. There is no current ETA for recovery of operations in the europe-west9 region at this time, but it is expected to be an extended outage" Emergency shutdown of multiple zones. As of a few hours ago they changed the status to report just on us-west9-a.
- tpmx 3y ago> As of a few hours ago they changed the status to report just on us-west9-a. The headline is 1h20m old and says a region is currently down.
- outworlder 3y agoEven the article linked mentions that Google initially reported that only zone A was affected, then got changed to report that the whole region was affected, now it changed to a single zone again. Do you expect realtime updates whenever Google changes the story?
- lukax 3y agoGoogle now has a "data lake" in Paris.
- deleted 3y ago[deleted]
- perrohunter 3y ago[flagged]
- nixcraft 3y agoI hope whoever is hosting data in that zone has thoroughly tested and verified backups offline or with another cloud provider. Of course, you can complete DC failover, depending upon service needs, but it costs more resources. Either way, timely tested backups are the only way to survive natural or manufactured disasters. Good luck to Google OPs team and everyone else involved with the GCP region in the EU.
- IntelMiner 3y agoGoogle's DC is underwater OVH's caught fire What's next, us-east-1 gets hit by Godzilla?
- cgb223 3y agoLol us-east-1 already went down for a day back in 2017 when an intern accidentally took down the whole DC. We could call him “Godzilla” Source: my startup (stupidly) hosted our entire infra in us-east-1 at the time. Was a …tough day
- throwawaaarrgh 3y agous-east-1 is a great place to test your application resiliency :) it's like they threw in chaosmonkey for free!
- mjr00 3y agoIt's funny because AWS, at least for the services I knew of when I was there, did rolling deploys to each region over several days. us-east-1 was always the final day because it was the biggest region, so you'd think it'd the safest region since everything getting deployed was well-tested. But while I was there I remember at least 2 COEs where the root cause was basically, "us-east-1 had some hacky legacy configuration that no other region has and that wasn't known/accounted for."
- kortex 3y agoA perfect example of when "cloud just means someone else's computers". It's literally a leaky abstraction.
- doubled112 3y agoSure is leaky. And a cloud is a bunch of water vapour that eventually comes crashing down to earth. I'll never understand how we decided it was a good metaphor for a place we run our services. Startlingly accurate in this case.
- 88913527 3y agoAfter spending 5 minutes engaging with Product Managers, I am not at all surprised they landed on calling it 'the Cloud'.
- dijit 3y ago“the cloud” comes from old network diagrams that used cloud to mean “internet” or “unknown network”. I think “unknown network” definitely accurately captures what hyperscalers are selling. :)
- paulmd 3y agohttps://www.youtube.com/watch?v=AnxrJiS5uKU https://www.youtube.com/watch?v=AnxrJiS5uKU
- benatkin 3y ago1) Pay for stuff 2) Not be able to use it 3) Company continues to pretend this doesn't happen on the regular
- worldsavior 3y ago> Company continues to pretend this doesn't happen on the regular What do you want them to say? "Hey we have X breakdowns but please, pay!"
- 3y ago
- deleted 3y ago[deleted]
- pclmulqdq 3y ago[flagged]
- fnordpiglet 3y agoLast time they let Bard pick a data center location and design.
- redindian75 3y agoThis is the problem storing data in the cloud - whenever it rains you may have a big data problem.
- burnt_toast 3y agoIt's okay because once the water evaporates its backup in the cloud.
- timack 3y agoReally? Are you cirrus?
- CobrastanJorji 3y agoA series of tubes would have helped with this.
- geocrasher 3y agoThank you for your input, Senator Stevens.
- moffkalast 3y agoOr a big truck to bring some tarps
- aruggirello 3y agoUnfortunately, it appears Google Plumber was discontinued by Alphabet Inc.
- iJohnDoe 3y agoThe plumber was laid off.
- t0mas88 3y agoI can't ignore the feeling that Google Cloud is sub par compared to AWS. How did this again cause a multi zone failure. Why haven't they fixed those dependencies the last few times they had a full region failure.
- dekhn 3y agozones and regions have different definition in google cloud than AWS. Multiple zones are physically co-located and are not truly availability zones because the physical proximity causes shared fates even when they have independent systems (network, power) that should allow one to fail while another doesn't. Even two datacenters in the same city are prey to the same meteor.
- styren 3y agoIn the same building is quite a bit different than 10km apart, even if a huge meteor would lead to the same conclusion for both scenarios.
- ericpauley 3y agoWow, this is a big issue! Is there any way to guarantee AWS-level physical redundancy in GCP without paying the latency/inter-region data transfer (higher than inter-zone outside US: https://cloud.google.com/vpc/network-pricing#egress-within-gcp https://cloud.google.com/vpc/network-pricing#egress-within-g...) pricing? A nice thing about EC2 is that you're getting a pretty dumb, predictable service. There have been multi-zone or global control plane issues but the physical metal has bona fide redundancy between zones/regions.
- lamontcg 3y agoI don't know what it is like these days but us-east AZs used to be in different datacenters that were on different flood plains and power companies. They were just on a very high capacity (for the day) fiber ring. You still had a couple miles of light-delay latency in between them. A sufficiently big enough meteor, or a powerful enough massive hurricane, could probably take out multiple ones at the same time.
- radicaldreamer 3y agoNot sure what kind of fire there was there, but once those automatic sprinkler systems get going, they are very difficult to stop. Someone in my freshman college dorm decided to use one as a clothes hanger hook and broke the thermometer in there. The sprinkler damaged the entire floor with water and the floor below had spotty rain as well. The fire department came and was mainly concerned about evacuating everyone rather than shutting the water off. The water is typically chemically treated and has been sitting there for years as well -- very nasty stuff.
- mvanbaak 3y agodatacenters dont use sprinkler systems (or at least they should not).
- glogla 3y agoYeah, I always thought datacenters would use Halon. It of course has the problem of suffocating everyone.
- packetslave 3y agoHalon has been banned for years because 1) it's bad for the ozone layer and 2) it'll kill you. Newer systems (FM-200, Inergen, etc.) fight the fire by removing heat instead of removing oxygen.
- alwayslikethis 3y agoHalon is still used. Unfortunately the same properties that makes it effective also makes it harm the ozone layer. It does not just remove heat or oxygen, it directly interferes with the reaction involved in combustion, making things stop burning.
- dx034 3y agoWasn't there also this technique of lowering oxygen levels so much that humans can still survive but fire won't spread as fast? Or did this turn out to be too expensive?
- throwawaaarrgh 3y agoI know I said our pipeline abstraction was leaky but this is ridiculous
- jqpabc123 3y ago[flagged]
- wkjagt 3y ago[flagged]
- lamontcg 3y agoNobody here with any thoughts for the operations/datacenter engineers trying to deal with stopping and cleaning up the disaster, just customers complaining...
- okdood64 3y agoNope. Big company bad. <Insert snarky overeactionary comment based on armchair knowledge> Literally no concern here for anyone's safety or sanity in dealing with this.
- rurp 3y ago"Thoughts and Prayers" type comments don't make for particularly interesting reading. I think it's safe to assume that most people feel empathy for others struggling, whether or not they type it out regularly. Then again, some AI evangelists have had me questioning that assumption lately.
- BlackjackCF 3y agoI think people who are complaining are stressed out about their own services being down. If you’ve only ever used the cloud, you’re not necessarily aware of everything that’s involved at data centers. If you’re not familiar with them, I don’t think you’d know how many things can (literally) blow up in your face. If someone sees flooding, they generally aren’t thinking that it’ll lead to fires. Anyway, just want to think that everyone generally has good intentions and just don’t know what’s ACTUALLY happening in the DC, or how much work it will be for the folks working in the DC to restore services. Hopefully all the failsafes kicked in and worked and nobody was injured.
- gauravphoenix 3y ago[flagged]
- trollingagain 3y ago[dead]
- 1970-01-01 3y ago[flagged]
- sgt 3y agoAnd ironed
- effdee 3y agoThis post (in french) has some more details: https://www.mail-archive.com/frnog@frnog.org/msg72320.html https://www.mail-archive.com/frnog@frnog.org/msg72320.html
- richardw 3y agoI can imagine clients who used one DC being impacted. But Google’s services would be designed for a single DC going down, right? Data would be eventually consistent (once they find and plug the hard drives in) but isn’t this the promise of the cloud and they’re (approximately) the best at using it. I have to assume it’s a fault that not even distributed services can paper over. Eg lots of crucial data in flight and they’re reluctant to drop it. Can an expert weigh in? I love Google’s post-mortems. This one will be epic.
- kccqzy 3y ago> But Google’s services would be designed for a single DC going down, right? Right. But nobody forces GCP's customers to design their services to be tolerant of a single DC failure. In fact as a business, actively not designing for such tolerance is an attractive cost-cutting measure.
- outworlder 3y agoCloud customers have no control on which or how many 'datacenters' are used. That's not something that's even advertised or easily available to customers. The logical units are regions and availability zones or the equivalent nomenclature in each cloud. One availability zone is expected to be one or more datacenters. We have thousands of instances in AWS. I do not know - or care - where they are physically located(other than the region name, say, Oregon). I expect at most one availability zone to get impacted if a datacenter goes up in flames (and sometimes, just a portion of one). I mention in another comment that AWS has had issues before and production systems barely got impacted. And recovered with zero intervention - instances with failed health checks get replaced by brand new ones in whatever AZs are still operational. > Data would be eventually consistent (once they find and plug the hard drives in) At the level of abstractions cloud operates, no-one is plugging drives in – someone is, but you can never see it. Most cloud workloads use network attached storage - when you can even see the logical drives (SaaS offerings may not even have that abstraction). We don't know (or care) how many physical hard drives exist, or where they are. Latency requirements probably dictate that they are close to the actual instances, but there's usually data replication going on even across DCs. In addition to that, at least in AWS, if you have saved any volume snapshots at all, they will be in S3. This data will be replicated and underlying systems can even use it to restore lost or corrupted data without you even noticing and sometimes even without a recent snapshot, as storage keeps track of what blocks have been rewritten since the last snapshot. In a particularly bad case you might have to do a restore. In almost a decade and number of volumes in the 6 digits (no clue how many drives that is!) we never had a single volume fail on AWS. Some got into a 'degraded' state and then recovered. We haven't had any failures on GCP either. In the case of GCP, even faulty hypervisors are transparently worked around - we never notice other than some audit logs saying the VM was moved. They even preserve the network connections. AWS requires a stop/start to do the same, but your VM will be up and running in a different hypervisor (sometimes a different datacenter) in a couple of minutes, with all the storage. Mind you, AWS promises eleven nines(!) of durability for S3. When you do have locally attached storage, it's treated as ephemeral and it's gone if the instance restarts. > I have to assume it’s a fault that not even distributed services can paper over. If a single datacenter fails, since it _should_ be at most one AZ(this case seems to be different) that will depend on how the application is architected. Requests in flight will obviously fail, how big of a deal depends on the problem domain. For most web apps, this will cause a retry and that's the end of the story, others will be specifically engineered to deal with receiving multiple messages or dropping messages. For example, if you need at most once delivery guarantees, you need to take extra measures Not all applications can survive an entire region going down. Some can, but that usually raises costs if you are continuously replicating data across regions. If you do that, then you should be able to steer traffic to the surviving regions. You can do that old-school by changing DNS records, or you could have fanciers solutions such as global anycast loadbalancers and have a single IP worldwide that still goes to the closest healthy region.
- dang 3y agoAll: please don't post low-effort comments that merely react to the first association you have. We're trying for curious conversation here, which is something else. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- asymptotic 3y agoWhen I worked at AWS there was a similar scenario in eu-west-2. There was a fire in one of the availability zones (AZs). The fire suppression system kicked in and flooded the data center up to ankle or knee height. All the racks were powered off and the building was evacuated for hours (I don't remember the duration of the evacuation) until the water was pumped out. But for the service team I worked for, our AZ-evacuation story wasn't great at the time and it took us tens of minutes to manually move out of the AZ, but at least there wasn't a customer-visible availability impact. Once we did it was just monitoring and baby-sitting until we got the word to move back in, I think it was 1-2 days later. If you operate on AWS you work with the assumption that an AZ is a failure domain, and can die at any time. Surprisingly many service teams at AWS still operate services that don't handle AZ failure that well (at the time). But if you operate services in the cloud you have to know what the failure domain is.
- jamesfinlayson 3y ago> urprisingly many service teams at AWS still operate services that don't handle AZ failure that well (at the time) Ouch, hopefully none of the major services? I recently had to look into this for work (for disaster recovery preparation) and it seemed like ECS, Lambda, S3, DynamoDB and Aurora Serverless (and probably CloudWatch and IAM) all said they handled availability zone failures transparently enough.
- asymptotic 3y agoI’m familiar with Lambda and DynamoDB. When I left in 2022 they both had strong automated or semi-automated AZ evacuation stories. I’m not that familiar with S3, but I never noticed any concerns with S3 during an AZ outage. I’m not at all familiar with Aurora Serverless or ECS. For all AWS services you can always ask AWS Support pointed, specific questions about availability. They usually defer to the service team and they’ll give you their perspective. Also keep in mind that AWS teams differentiate between the availability of their control and data planes. During an AZ outage you may struggle to create/delete resources until an AZ evacuation is completed internally, but already created resources should always meet the public SLA. That’s why especially for DR I recommend active-active or active-“pilot light”, have everything created in all AZs/regions and don’t need to create resources in your DR plan.
- palcu 3y ago[disclaimer: SRE @ Google, I was involved with the incident, obvious conflicts of interest] Hey Dang, thanks for cleaning up the thread. One thing to note is that the title is not correct. The entire region is not currently down, as the regional impact was mitigated as of 06:39 PDT, per the support dashboard (though I think it was earlier). The impact is currently zonal (europe-west9-a), so having zone in the title as opposed to region would reflect reality closer. Finally, there's lots of good feedback on this thread and on the previous one (https://news.ycombinator.com/item?id=35711349 https://news.ycombinator.com/item?id=35711349), so we obviously have a lot of lessons to learn.
- Waterluvian 3y agoWould you be able to comment a bit on the emotional (perhaps there’s a better word) aspect of the response? Was there a lot of anxiety? Panic? Or was it just a “woof that sucks. Time to follow a checklist and then do a bunch of paper work” ? What I’m curious about is what it feels like on a team at a company like Google when there is a major system failure.
- palcu 3y agoThere's not much emotion as the core team working on the huge outages is more like an "SRE for SRE". They are all people who've been with the company for a long time and they've been in the secondary seat for at least one previous big rodeo. Not to mention that we're all running a checklist that has been exercised multiple times and there's always somebody on the call who could help if a step fails. Personally, I wasn't part this time for the actual mitigation of the overall Paris DC recovery, as I was busy with an unfortunate[0] side effect of the outage. These generate more anxiety, as being woken up at 6am and being told that nobody understands exactly why the system is acting this way is not great. But then again, we're trained for this situation and there are always at least several ways of fixing the issue. Finally, it's worth repeating that incident management is just a part of the SRE job and after several years I've understood that it is not the most important one. The best SREs I know are not great when it comes to a huge incident. But, they're work has avoided the other 99 outages that could have appeared on the front page of Hacker News. [0]: https://news.ycombinator.com/item?id=35734224 https://news.ycombinator.com/item?id=35734224
- faangiq 3y agoBut I thought Google only hired geniuses?
- iJohnDoe 3y agoYes, and we often hear genius insights and anecdotes from their SRE employees who have their blog links posted to HN.