39 ms·
Cooling related failure (in Google London DC)
- dgl 4y agoOracle also had issues around the same time: https://ocistatus.oraclecloud.com/#/incidents/ocid1.oraclecloudincident.oc1.phx.amaaaaaavwew44aa7zoskanlspjh4ll6wxhwxrbkbed4d4cnupxexzqzvlyq https://ocistatus.oraclecloud.com/#/incidents/ocid1.oraclecl... Obviously the weather is easy to blame, but I wonder if the underlying cause is the same datacenter. It's kind of annoying when clouds use availability zones as such an opaque thing, it's not possible to map a zone from one cloud to another and potentially your failure domains overlap. Shameless self promotion: I have a map of all the cloud regions (but not going into as much detail as availability zones): https://cloud-regions.bodge.cloud https://cloud-regions.bodge.cloud -- the clouds just don't publish the data.
- threeseed 4y agoIt is possible to infer the locations if you had a lot of time. AWS at least in the past use to co-locate with other third party servers in the data centre. And so you if you had such a server you could ping AWS endpoints to triangulate where physically those servers might be.
- jhugo 4y agoI'm pretty sure a majority of AWS AZs globally are still not in AWS-owned-and-operated facilities.
- sneak 4y agoAt this point I would be quite surprised if they were not majority AWS owned and operated.
- jhugo 4y agoOf the ones that I know details about in Europe and Asia, not a single one of them are AWS owned and operated. I'm sure the situation is different in the US, and wouldn't be surprised if they own and operate all of their US DCs.
- tambre 4y agoWhy does your map have AWS's eu-north-1 in Estonia? Officially it's said to be in Stockholm and I can find no references nor know of it existing there.
- dgl 4y agoStrange, I'm not sure where I got that from, will fix.
- sneak 4y agoYou have some locations simply pinned to the central squares of cities when that is definitely not where the facilities are. Maybe include a +/- error range on the provided data?
- cjg_ 4y agoThe eu-north-1 AZ’s are in Västerås, Eskilstuna and Katrineholm, all ~100 km from Stockholm and ~40-80 km from each other.
- jeffbee 4y agoI also think it would be nice if Google prominently stated which cloud regions are hosted in their real datacenters and which are in third party facilities. It's not just that you run the risk of some off-brand piece of junk overheating, but also because they so loudly trumpet their PUE and carbon neutrality, but then resell cloud capacity in other people's facilities where you have no idea what the carbon impact is.
- lupire 4y agoThat's not how carbon impact works. You don't buy "clean" products. You buy products from vendors who have an overall average mix of production that meats a "clean" threshold, and/or buys credits from someone else who exceeds threshold.
- sofixa 4y agoIn Google's (and for instance Scaleway too), the point is that some of their DCs use renewable energy only, thus are "clean". If however you use the "wrong" Google region which is hosted in Equinix's DCs, which happen to be powered by coal, it's not even remotely close. However Google make the distinction and you can easily check which region is "clean" and which isn't.
- dgl 4y agoThey do tell you the carbon impact right on the region selector: https://cloud.google.com/blog/topics/sustainability/pick-the-google-cloud-region-with-the-lowest-co2 https://cloud.google.com/blog/topics/sustainability/pick-the... and also https://cloud.withgoogle.com/region-picker/ https://cloud.withgoogle.com/region-picker/ is quite nice.
- mike_d 4y agoGoogle doesn't really have a concept of datacenters internally. In simplified terms everything runs on a borg cell, which usually but not always exists in a single building. Within a region one of your GCP services could be running on one cell, with another service running on another in a different building. If a failure happens or maintenance needs to be done, a cell could be drained to another one in the same region.
- deleted 4y ago[deleted]
- angch 4y agoYou have some of the empty lat, lon data (e.g. AWS GovCloud (US)) interpreted as at (lat,lon) = (0,0), putting them incorrectly on Null Island.
- dgl 4y agoThanks, hidden.
- sofixa 4y agoFor the AZs that's slightly complicated by the fact that at least for AWS, they are randomised (your eu-west-1a isn't necessarily the same as mine), so it might be pretty much impossible to actually check for actual overlap even if you know one DC is used by both Google and Oracle. Spreading out regions seems more appropriate to increase redundancy.
- BillinghamJ 4y agoYour own mappings are available in the console and on the API. So definitely still possible to tell
- deleted 4y ago[deleted]
- dgl 4y agoAmazon does provide a way to map to the real availability zone: https://docs.aws.amazon.com/ram/latest/userguide/working-with-az-ids.html https://docs.aws.amazon.com/ram/latest/userguide/working-wit... It's also not appropriate in all cases to spread out, for example the nearest zones to London are around ~7ms away in mainland Europe.
- klohto 4y agoThe provided info is only relative to AWS. If you wanna map geographical destination across clouds you’re out of luck unless you have reports with precise location available. Generally, AWS already provides certain assurances that AZs are spread out.
- formerly_proven 4y agoI didn't know this, but it makes perfect sense and randomizing the a/b/c is a tried and true solution - e.g. in the electrical grid, L1/2/3 are rotated for each customer, because people tend to connect more stuff to L1 on average.
- thunderrabbit 4y ago
- TheOtherHobbes 4y agoUK infra - physical and electronic - is typically rated to around 30C. So higher temps are going to cause failures in multiple locations. If yesterday's temps lasted for a week instead of a day or two a lot of essential stuff would stop working.
- nsteel 4y agoCarrier-grade (i.e. critical) network equipment should be rated for 40C ambient air and then up to 55C for some short length of time (3 days??) to allow for air conditioning failure. This is what we design and test for. Cheap stuff, including Google's software solutions running on 'commodity hardware', won't handle that. You get what you pay for.
- ben_w 4y agoDoes the UK even design with the assumption of air conditioning being necessary and present, let alone designing for it failing?
- FartyMcFarter 4y agoDatacenters always have A/C I thought?
- ben_w 4y agoThanks :)
- Relestio 4y agoGoogle and other companies do make risk assessments including temperature scenarios.
- nsteel 4y agoYes. You must be able to control the temperature of a building densely packed full of hot radiators. You might be able to avoid active AC within the Arctic circle but you'd m still want to filter that incoming (very) cold air, at which point maybe you might as well basically have an AC system. And yes, you must ensure you can survive a failure because these systems do fail (normally when you need them most) and then otherwise everything inside cooks.
- exikyut 4y agoWow, this map is really cool. I'm (very) idly curious how accurate the reporting for the Sydney region is, because I just looked up the lat/lon, and found myself in the middle of an urban mall environment that I've walked past many times when I've been in the city. Being able to look up at the buildings there and know there are indeed 5 different clouds somewhere up above my head in that specific location would be really cool. Being able to point at a specific building would be even cooler. I do of course (sadly) appreciate the flip side of this coin which is one of (many of) the reasons precise data is not published. So I guess I'm just wondering out loud, probably rhetorically :), how I might find out one day. "<-- That building" is enough resolution for me :) EDIT: A quick Google found baxtel.com (among other websites) has address-level locations for most providers. The buildings are all so boring, haha! (Understandably so though.)
- larrybud 4y agoIt’s pretty far off, at least for azure. It shows a number of azure regions as being located in downtown urban areas, which is typically not the case
- ggm 4y agoElectricity generation is also affected by weather in two ways. Firstly, by the droop in the wires which are built to weather tolerances much as railways are, and if you get outside the tolerances then the risks of problems increase. Secondly, within a limit of my understanding, the efficiency of a turbine system hot-to-cold is affected by the climate it operates in. The ambient temperature, humidity and pressure affects the final stage. Both things might mean that in times of high heat and humidity, the electricity supply system is least able to cope with increased demands for cooling systems, which will themselves draw more power fighting the weather. Separately the HVAC systems for the DCs will have been designed for a specific climate, with margins. I guess the sustained change in night and daytime temps and humidity has hurt their efficiency too, in this window of time. They'll be fine when the weather system passes through, as will the supply network. Met Office says both overnight and daytime peak temps for the inland south have been records. Thats where a lot of ICT infrastructure is.
- deleted 4y ago[deleted]
- Scoundreller 4y agoWe’re also just under a month from the summer solstice (in N. Hemisphere: June 21). London is at 51 degrees North, so the day is pretty long at 16 hours right now. IE: less of that relieving overnight low.
- mnd999 4y agoFortunately it rained last night and the temperature dropped pretty quickly.
- kumarvvr 4y ago> Secondly, within a limit of my understanding, the efficiency of a turbine system hot-to-cold is affected by the climate it operates in. The ambient temperature, humidity and pressure affects the final stage. Usually this is only partially true. Systems have enough margin to account for such conditions (cooling pumps have to pump more water to compensate for higher temps). Also to note that the output of a turbine can be kept constant. Only the efficiency will come down slightly.
- deleted 4y ago[deleted]
- jvolkman 4y agoThis incident page really needs a diff mode or just less boilerplate. It's really difficult to tell at a glance what's changed with each update.
- jlmorton 4y agoCame here to say this. Each update almost subtracts value. Updates should only contain information that has changed. Very frustrating when you're anxiously awaiting new information, and you have to do a word-by-word mental diff.
- willhackett 4y agoI really like this. I think a change to the input questions could solve this — clearer, more specific questions like "What's changed?", "Is it worse/better?".
- scary-size 4y agoAdd to that all timestamps being US/Pacific for an outage in Europe...
- davidkuennen 4y agoIt's funny how HN always complains about every status page. I think Google Cloud has one of the only status pages that is always up to date and very forthcoming in giving as much detail as possible. Personally I couldn't ask for more.
- the_sleaze9 4y agoIt's casual feedback, sure. But it is (by and large) specific and actionable. I think "you couldn't ask for more" is disingenuous at best, and an actively harmful outlook at worst.
- POPOSYS 4y agoIf you are really that interested in the content of a status page of one cloud service provider you should redesign your infra. The value proposition of "cloud" for is not "I can haz cloud of big corp as my own" but "I can haz many cloudz to make resilient infra!". If you rely on one cloud provider you are doing it wrong.
- jonatron 4y agoI had a portable 14000 BTU unit running flat out, and it couldn't keep up. 40c is hard to deal with here.
- sumanthvepa 4y agoI’m a little confused. 40C is fairly common in India where I live. My air conditioning works fine here even in relatively high humidity. Is there something special about how a/c units work in the UK? Are they rated for lower ambient temperature or something?
- xdfgh1112 4y agoOP probably has a small AC unit with a pipe that you send out of the window. We don't have proper AC over here in most houses, because it's normally only hot a few weeks out of the whole year.
- tjoff 4y agoPortable units have terrible efficiency. Unfortunately that is often the only option in apartments.
- jonatron 4y agoIt's less about the AC units, and more about the buildings. In places like Spain and Portugal, white painted outside walls, shading, and shutters on the outside of windows all help keep the heat out. My house has none of that, it just absorbs most of the heat from the sun.
- sumanthvepa 4y agoAh! That explains it. I'm sure if my house were in the UK, I would freeze to death in winter.
- ajdegol 4y agoUK houses are totally fucking useless for cold weather too.
- sandGorgon 4y agoI would like to present AWS and Google Cloud Mumbai. heat wave and covid wave simultaneously! 40 degrees centigrade is warm here.
- zinekeller 4y agoAnd I would actually expect that it is equipped to handle 40-degree outside temperatures year-long (same in zones in southwestern US - I meant 105-degree heat because freedom units!) London didn't historically experience these kind of heat though, only peaking to 37-degree in an hour.
- xxs 4y agowho the heck measures electronics temps/ambient in freedom units? All the datasheets are exclusively in C. [105F is precisely 40C, though]
- knorker 4y agoEverywhere in the world you design for what the local conditions are. I bet Mumbai doesn't mandate winter tyres in winter, right? Sweden does. If suddenly one winter you see -10℃ in Mumbai, would you appreciate Swedes mocking you for not being able to drive in a car not designed for it, with tyres not designed for it, on roads not designed for it, etc…? Nothing in the UK is designed for 40℃. Buildings, the type of steel made for train tracks, ventilation in tube tunnels, the asphalt, the walls in the building, the windows (no double glazing in Mumbai, I assume?). I would expect everything in Mumbai is designed to handle high temperatures. But not cold.
- Linda703 4y ago[dead]
- benjaminwootton 4y agoCooling seems to be one of the first things to creak in hot weather. Our fridges at home were struggling, and some of the local supermarkets had fridges out. Maybe that’s an obvious observation but I would have expected they had a little more operating range right at the point you need them.
- londons_explore 4y agoWorth noting that most home fridges have no kind of indication that they can't keep up. Your fridge, rather than being the 5C it should be to keep your meat safe to eat, might have been up at 12C. You wouldn't be aware (it still feels cold), but you'd end up eating possibly dangerous food. I really wish fridges had an alarm in that case (ie. The fridge has an indicator saying 'too hot. Food is now unsafe to eat').
- emrvb 4y agoThat alarm is available on the more decent models. There are also stickers for inside your fridge that can indicate the temperature. There is also a variant for specific temperatures, like 0, 5 or 7 degrees, that colorize if the temperature has risen, giving a (non-reversible) indication your fridge has been too warm. Edit: Did a quick search for you: https://www.tiptemp.com/Products/Rising-Time-Temperature-Indicating-Labels/ https://www.tiptemp.com/Products/Rising-Time-Temperature-Ind...
- benjaminwootton 4y agoOurs was making strange noise, and we noticed things were wet when we took them out which indicated it was struggling to keep temperature. Agree some kind of alarm or indicator is probably a good idea if its not keeping up!
- londons_explore 4y agoWet things inside is actually an indication that the door isn't properly closed and sealing.
- 4y ago
- kp8h 4y ago
- clhodapp 4y agoGCP sure does seem to have a lot of outages that span across multiple availability zones (and occasionally across multiple regions). It sure does seem like there is a disconnect between the expectations of isolation that they set and what they are able to deliver. It's also interesting that the status page's attempt to spin the scope of impact actually makes it seem worse that it was full-region outage (they said, "There is a cooling related failure in one of our buildings that hosts a portion of capacity for zone europe-west2-a for region europe-west2 ...")
- johndfsgdgdfg 4y agoThe amount of outages happening on GCP is mind boggling. I don't know how do people trust their business on Google. It's an ad company, not an infrastructure company. If you trust an ad company with your business, I guess it's on you.
- internetting3 4y agoGreat point, let's trust a retail company instead.
- ccbccccbbcccbb 4y agoRest assured, if the agenda ever u-turns into "global cooling", google will follow with "heating related failure" reports.
- lloydatkinson 4y agoWhat?
- LAC-Tech 4y agoUnfortunate I guess, but I still remember having to pull an overnighter at work because there was a leak in the sever room :)
- pojzon 4y agoIm curious whether all of those companies migrating to „the cloud” are thinking about those issues. It will be only more and more impactful also with power outages on the horizon. Company I work now in completely does not care at least based on me rising those concerns to mng.
- monkeydust 4y agoFrom a business point of view, if your looking to move your infrastructure to cloud do you now should you be factoring in global warming? If so, This will only play into the hands of those setting up data centres in the far north of the northern hemisphere (e.g. Iceland).
- POPOSYS 4y agoThis reads like: "Oh, tanks by that foreign army are rolling into our cities - should we start thinking about a defense system?" Global warming will affect every aspect not only of your business, but also of your life, maybe a little bit later when you are rich and can afford to live in a self-created bubble, but it will. Yes, you should think about it. Hard. Now.
- the_sleaze9 4y agoVery much - YES.
- danpalmer 4y agoI don't think this really has anything to do with the cloud. On-premise hardware still needs cooling, and is arguably harder to cool as there are fewer economies of scale on cooling infrastructure. Dedicated "bare-metal" machines are just in regular data centres so no difference to the cloud there. I think data centre locations will still be chosen on two factors: distance to customers, and cost of energy. It's just that operators will be looking for cheap energy. Iceland is good because they have a lot of geothermal energy, not because it's cold.
- avidphantasm 4y agoI originally read this title as “Cooking related failure”. It reminded me of when I worked for a US research university in the late 90s and one if the data center operators accidentally heated up a hot dog wrapped in aluminum foil in a microwave. A microwave that was stupidly plugged into a circuit share by some of the network equipment and servers. The microwave allegedly blew up or something, causing a multi-hour outage of the campus network.
- mloughran 4y agoThere was a concurrent incident affecting a large number of GCP services: https://status.cloud.google.com/incidents/fmEL9i2fArADKawkZAa2 https://status.cloud.google.com/incidents/fmEL9i2fArADKawkZA.... It would have been rather useful had GCP linked these (presumably linked) incidents. The first mention of cooling in the concurrent incident was at 14:39 PDT (over 5 hours after first status update, and 4 hours after the cooling incident was created)... This is what was said: > Description: A cooling related failure in one of our buildings that hosts zone europe-west2-a for region europe-west2 is affecting multiple Cloud services.
- dozzman 4y agoMy main confusion with this downtime is that neither their Cloud SQL nor Redis offerings managed to complete fail over despite my org having high availability enabled on both of those plans. Is there something I'm missing here? I would've suspected that failover would kick in for high availability instances and cause minimal downtime however its been almost 24 hours and our Cloud SQL instance is still stuck on attempting to fail over, not to mention that it comes at a premium. Wondering if anyone can help me understand what I'm missing or if the failover behaviour is not working. We've made our own workarounds in the mean time. Relevant docs I've checked for behaviour: https://cloud.google.com/memorystore/docs/redis/high-availability https://cloud.google.com/memorystore/docs/redis/high-availab... https://cloud.google.com/sql/docs/mysql/high-availability https://cloud.google.com/sql/docs/mysql/high-availability EDIT: Have found out from our ops team that the SQL instance recovered around 3am so it was down for approximately 9 hours -- which is still totally useless for something deemed HA.
- onphonenow 4y agoThat’s pretty terrible!
- cshou 4y agoIIUC, HA setting only failover across *zones in the same region*. If the whole region is down, HA won’t be helpful. In this case, the London data center is the region.
- giggsey 4y agoThe region wasn't down though. Only one zone was down? From earlier in the incident history: > Cloud SQL: > Impact/Diagnosis: Non-HA instances backed by europe-west2-a are hard-down in europe-west2-a. HA instances that were in europe-west2-a when the incident started, are down with stuck failovers.
- nogbit 4y agoThat’s expected, Cloud SQL is not multi region. Clouds define HA as being multizonal, which you were. Try Spanner if one region is not enough.
- johnklos 4y agoI'm a huge fan of wondering openly whether what people commonly do is best, or if it's just best for some people and everyone else does it because they haven't really thought about it. For instance, the idea that an expensive, thousand plus dollar rackmount server is only able to run in a special place where the temperature is just right and might fail if a fan fails or the temperature is a wee bit higher than usual is utter bollocks. I build my own rackmount servers that can run at 100º temperatures, even with fan failures. I know this because that's how I test them. I have the OS aggressively throttle on the most egregious failures, but fan failure is much less common in general when you're using 80mm Noctua fans instead of 40mm fans that have to run at many thousands of RPM to keep their zones cool. So maybe people need to rethink the idea that datacenters have to be kept at 70º or below, and instead should insist on better thought out hardware.
- pclmulqdq 4y agoGoogle datacenters run hot, as do many other cloud providers. They are happy to use high temperatures in their DCs to improve efficiency. The problem is that they have also removed the equipment to handle super-hot days for "efficiency."
- knorker 4y agoYour servers can run at 100℃ ambient temperature? Oh, I guess you mean F? Yeah, servers can run just fine at 40℃. Well… unless they fail because of it. :-) That is: The ones that don't fail work just fine. If you have a DC with 10'000 of your servers, and maybe 30'000 hard drives. What percentage of them will fail on any given day at 25℃ vs 40℃? But it's not just your servers. Can your AC equipment/evaporators work at 40℃? And if your ACs start failing you could be looking at a cascading failure where it's actually more like 60-70℃, or just plain "a fire", in your DC. Can your generators work? And the answer also isn't "every component in my DC must be milspec extended temperature range". It's actually fine to build a DC in Iceland that's not specced for outside temperature of 50℃. In fact it would be a ridiculous waste to do so. Google of course measures this. E.g. https://www.techrepublic.com/article/google-research-temperature-not-a-major-factor-in-hard-disk-failures/ https://www.techrepublic.com/article/google-research-tempera... But do keep in mind that 40℃ outside may mean 50℃ inside. Or indeed just 60℃ hotspots on the DC floor. Hell, your network cables may not even be rated for 60℃. They usually aren't. Your server may be fine with 40℃ inlet, but going out the air may disconnect it due to melting the network cable.