6 ms·
AWS North Virginia data center outage – resolved
https://health.aws.amazon.com/health/status?t=2026-05-07 https://health.aws.amazon.com/health/status?t=2026-05-07
https://www.theregister.com/off-prem/2026/05/08/aws-warns-of-ec2-impairment-as-power-loss-hits-notorious-us-east-1-region/5235509 https://www.theregister.com/off-prem/2026/05/08/aws-warns-of...
https://www.reuters.com/business/retail-consumer/amazon-cloud-unit-says-data-center-overheating-north-virginia-disrupts-services-2026-05-08/ https://www.reuters.com/business/retail-consumer/amazon-clou...
- merek 5mo agoRelated: AWS EC2 outage in use1-az4 (us-east-1) https://news.ycombinator.com/item?id=48057294 https://news.ycombinator.com/item?id=48057294
- tcp_handshaker 5mo agoI bet post-mortem will say vibe coding confused fahrenheit and celsius, we run too hot...
- geodel 5mo agoNow I totally understand the issue. I will set temperature 70K. Using SI unit of temperature is the best practice.
- fabian2k 5mo agoI thought cooling was pretty much pre-planned in any data center, and you simply don't install more stuff than you can cool? So did some cooling equipment fail here or was there an external reason for the overheating? Or does Amazon overbook the cooling in their data centers?
- DevelopingElk 5mo agoOne of the data center's cooling loops broke.
- bdangubic 5mo agoNo backups?
- bradgessler 5mo agoWhat happens when the backup breaks?
- noir_lord 5mo agoYou have a back up for the back up backup. Turtles all the way down. At AWS scale even unlikely hardware events become more common I guess.
- odyssey7 5mo agoEach turtle gives them another 9. How many 9s are they down due to incidents over the past year?
- tardedmeme 5mo agoThey're definitely more than half a day now, which is only two and a half nines.
- oldmanrahul 5mo agoAt a certain point earth is a single point of failure.
- minimaltom 5mo agoThey absolutely have backups, I presume they were ineffective or also down for _reasons_.
- michaelt 5mo agoI once worked at a company that had a wealth of backups. A backup generator, backup batteries as the generator takes a few seconds to start, a contract for emergency fuel deliveries, a complete failover data centre full of hot standby hardware, 24/7 ops presence, UPSes on the ops PCs just in case, weekly checks that the generators start, quarterly checks by turning off the breakers to the data centre, and so on. It wasn't until a real incident that we learned: (a) the system wasn't resilient to the utility power going on-off-on-off-on-off as each 'off' drained the batteries while the generator started, and each 'on' made the generator shut down again; (b) the ops PCs were on UPSes but their monitors weren't (C13 vs C5 power connector) and (c) the generator couldn't be refuelled while running. Even if you've got backup systems and you test them - you can never be 100% sure.
- AdamJacobMuller 5mo agoThis is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC components and require mechanical repairs that require specialized labor and long lead times. In a bigger facility its not uncommon to have cooling be N+3 or more when N becomes a bigger number because you're effectively always servicing something or have something down waiting for a blower assembly which needs to be literally made by a machinist with a lathe because that part doesn't exist anymore but that's still cheaper than replacing the whole unit. The system are also under-provisioned in the sense that if every compute capacity in the facility suddenly went from average power draw to 100% power draw you would overload the cooling capacity, you would also commonly overload things in the electrical and other paths too. Over provisioning is just the nature of the industry. In general neither of these things poses a real problem because compute loads don't spike to 100% of capacity and when they do spike they don't spike for terribly long and nobody builds facilities on a knife-edge of cooling or power capacity. The problem comes when you have the intersection of multiple events. You designed your cooling system to handle 200% of average load which is great because you have lots of headroom for maintenance/outages. Repair guy comes on Tuesday to do work on a unit and finds a bad bearing, has to get it from the next state over so he leaves the unit off overnight to not risk damaging the whole fan assembly (which would take weeks to fabricate). The two adjacent cooling units are now working JUST A BIT harder to compensate and one of them also had a motor which was just slightly imbalanced or a fuse which was loose and warming up a bit and now with an increased duty cycle that thing which worked fine for years goes pop. Now you're minus two units in an N+2 facility. Not really terrible, remember you designed for 200% of average load. That 3rd unit on the other side of the first failed unit, now under way more load, also has a fault. You're now minus 3 in a N+2 facility. Still, not catastrophic because really you designed for 200% of average load. The thing is, it's now 4AM, the onsite ops guy can't fix these faults and needs to call the vendor who doesn't wake up till 7AM and won't be onsite till 9. Your load starts ramping up. Everything up above happens daily in some datacenter in the USA. It happens in every datacenter probably once a year. What happens next is the confluence of events which puts you in the news. One of your bigger customers decides now is a great time to start a huge batch processing job. Some fintech wants to run a huge model before market open or some oil firm wants to do some quick analysis of a new field. They spin up 10000 new VMs. Normally, this is fine, you have the spare capacity. But, remember, you planned for 200% of AVERAGE cooling capacity and this is not nodes which are busy but not terribly busy, these are nodes doing intense optimized number crunching work which means they draw max power and thus expel max waste heat. Not only has your load in terms of aggregate number of machines spiked but their waste heat impact is also greater on average. Boom, cascading failure, your cooling is now N-4. Server fans start ramping up faster which consumes more power. Your cooling is now N-5. Alarms are blaring all over the place. Safeties on the cooling units start to trip as they exceed their load and refrigerant pressures rise. Your cooling is now N-6. Your cooling is now N-7. Your cooling is now 0.
- Andys 5mo agoI worked in a DC that had multiple redundant chillers on the roof, and multiple redundant coolers on each floor, but the whole building's cooling failed at once when the water lines failed somehow. They didn't say how, but apparently the pipes between each floor and the roof were not redundant. It took almost 24 hours to fix.
- el_benhameen 5mo agoGood listen on similar topics here: https://signalsandthreads.com/the-thermodynamics-of-trading/ https://signalsandthreads.com/the-thermodynamics-of-trading/
- Havoc 5mo agoCould someone explain to me why they don't build these things near oceans? Like nuclear plants that need plenty cooling capacity too Two loop cycle with heat exchanger to get rid of the heat
- sheept 5mo agoThis is just a guess, but land near oceans is more expensive/populated, and water is comparatively cheap
- kinow 5mo agoI had a class in my masters about data centers (HPC Infrastructures). The professor was using some data centers somewhere in the middle of USA, in an area with hot weather as example. He compared that with ideal scenario (weather, power source, etc.). In one of the slides, there were factors that influence the decision of where to build a data center, and several of the items involved finding a place with enough space and skilled people to work at this data center. He also commented sometimes there is politics involved on choosing the place for a next data center.
- ikr678 5mo agoOff the top of my head: Ocean levels of salt in a water system are much more expensive to maintain (even the secondary loop). Coastal land much more expensive. If you go to a remote coastal site, you probably won't have as good access to power. Coastal sites usually exposed to more severe weather events. Other fun unpredicatble things eg-Diablo Canyon nuclear facility has had issues with debris and jellyfish migration blocking their saltwater cooling intake. https://www.nbcnews.com/news/world/diablo-canyon-nuclear-plant-california-knocked-offline-jellyfish-creature-called-flna739606 https://www.nbcnews.com/news/world/diablo-canyon-nuclear-pla...
- idiotsecant 5mo agoAnd oysters / mussels / clams / every other creature that starts small and turns calcium into brick finds your cooling system to be a delightful place to raise a family, especially in delicate heat exchangers with small easily blockable passages.
- tailscaler2026 5mo agous-east-1 is down? shocking! stop putting SPOF services there. this location has had frequent issues for the past 15 years.
- unethical_ban 5mo agoThis is correct... unless there is a specific requirement to be in that location for some kind of IXP or ultra low latency, I can't imagine putting mission-critical things in only that region.
- deleted 5mo ago[deleted]
- cmiles8 5mo agoAWS’s US-East 1 continues to be the Achilles heel of the Internet. And while yes building across multiple regions and AZs is a thing, AWS has had a string of issues where US-East 1 has broader impacts, which makes things far less redundant and resilient than AWS implies.
- keeganpoppen 5mo agoanecdotally (well, more "second-hand-ly i heard that..." it sounds like there were some carry-on effects on us-east-2 as a result of people migrating over from us-east-1, so, yeah... kinda hilarious how the multiple region / AZ thing is just so plainly a façade, but yet we all seem to just collectively believe in it as an article of faith in the Cloud Religion... or whatever...
- qaq 5mo agoIt's no magic given the size of us-east-1 there is no spare capacity to absorb all the workloads
- 8organicbits 5mo agoOne of the SRE tricks is to reserve your capacity so when the cloud runs out of capacity you're still covered. It's expensive, but you don't want to get stuck without a server when the on-demand dries up.
- cherioo 5mo agoIs it really failing more, or we just don’t hear about failure happening elsewhere? Last i heard azure outage it wasn’t even on HN frontpage
- stingraycharles 5mo agoIt really is failing more, and it’s well known amongst industry experts. It’s the oldest, largest, and most utilized region of AWS. I’ve heard people say that the underlying physical infrastructure is older, but I think that’s a bit of speculation, although reasonable. The current outage is attributed to a “thermal event”, which does indeed suggest underlying physical hardware. It’s also the most complex region for AWS themselves, as it’s the “control pad” for many of their global services.
- nikcub 5mo agoboth realtime markets where multi-AZ is hard?
- kikimora 5mo agoOrder books had to run on a single server for performance reasons. Similarly a realtime multiplayer game.
- OhMeadhbh 5mo ago[flagged]
- aurareturn 5mo agoThese things are dangerous. Someone who can take AWS down such as an employee can place a bet. These bets aren’t as innocent as they seem because the bettors can often influence or change the outcome.
- shimman 5mo agoIt's a good thing big tech hires for ethical engineers and not ones that only care about money or social status.
- zaphirplane 5mo agoYou forgot the /s
- morgoths_bane 5mo agoThankfully their leadership is leading the way in ethics since inception, so I am confident that no such shenanigans will ever take place. I may even bet on this.
- whatsupdog 5mo ago[flagged]
- 5mo ago
- BugsJustFindMe 5mo ago[flagged]
- aussieguy1234 5mo agoOnce known for having super reliable services, I've heard this company is scrambling to re hire some of the engineers they overconfidently "replaced" with AI. When customers pay for cloud services, they expect them to be maintained by competent engineers. edit: Not sure why the downvotes. If you fire the engineers that have been keeping your systems running reliably for years, what do you expect to happen?
- wmf 5mo agoIt's a cooling equipment failure. Equipment is going to fail.
- corvad 5mo agoIt's always East 1... Jokes aside I don't understand how often east-1 is taken down compared to other regions. Like it should be pretty similar to other regions architecture wise.
- tom1337 5mo agoIsn't east one the "core" datacenter and also the oldest? I'd imagine it has more load than the other regions and also has more tech debt and architectural / engineering debt because they had less experience when they built it. Also iirc some services rely on east-1 as a single point of failure for configuration (like IAM or some S3 stuff?)
- doitLP 5mo agoYes it tends to have the most things running in it, including backbone and internal services that only exist in that region.
- __turbobrew__ 5mo agoWhat I have seen at other companies is that the older datacenters have suboptimal designs which are impossible to fix after the fact.
- JimDabell 5mo agoAmusingly: > AWS in 2025: The Stuff You Think You Know That’s Now Wrong > us-east-1 is no longer a merrily burning dumpster fire of sadness and regret. — https://www.lastweekinaws.com/blog/aws-in-2025-the-stuff-you-think-you-know-thats-now-wrong/ https://www.lastweekinaws.com/blog/aws-in-2025-the-stuff-you... Otherwise a good article!
- tardedmeme 5mo ago> Otherwise a good article! Who is Gell-Mann and why is he so forgetful?
- bandrami 5mo agoIt's the oldest regional system and has some structural importance (e.g. the internal CA resides there I think)
- yomismoaqui 5mo agoHow many nines of are we at this year?
- toast0 5mo agoEight eights!
- fastest963 5mo agoCoinbase claimed multiple AZs were down but the AWS statement was that only a single AZ was affected. Does anyone have more details?
- merek 5mo agoI can't find an official source, but I suspect the blast radius isn't limited to the AZ. I have systems running in us-east-1, and over the course of the incident, I noticed unexplainable intermittent connectivity issues that I've never seen before, even outside of az4.
- b40d-48b2-979e 5mo agoNever trust a crypto company to be honest.
- adamg203 5mo agospent the evening looking at SLI graphs waiting for the region to blow up but it never did. only a few envs across many had some degraded EBS vols in the single AZ. it was absolutely a single az (use-az4).
- bombcar 5mo agoEast-1 going down always takes some things from other AZs, because there's always something dependent on East-1.
- fastest963 5mo agoCoinbase confirmed on X that the exchange only ran in one AZ for latency reasons: https://x.com/i/status/2052855725857329254 https://x.com/i/status/2052855725857329254
- jeffbee 5mo agoI don't see anything on downdetector suggesting this was particularly disruptive.
- sitzkrieg 5mo agousing aws since s3 came out and i’ve yet to see any major company do multi az failover in any capacity whatsoever. default region ftw
- jedberg 5mo agoWe were doing multi-AZ and multi-region failover at Netflix all the way back in 2011: https://netflixtechblog.com/the-netflix-simian-army-16e57fbab116 https://netflixtechblog.com/the-netflix-simian-army-16e57fba...
- sitzkrieg 5mo agofair point. credit where it is due not a user but my understanding is netflix never went down w us-east-1 :-)
- deleted 5mo ago[deleted]
- matt3210 5mo agoRight, cooling.
- fukinstupid 5mo ago[flagged]
- whatever1 5mo ago2/last 365 days down. My Ubuntu nas is 0/last 365days down. Come and give me your cash if you want resilience.
- justincormack 5mo agoBut you couldnt apply security updates because Ubuntu was doen
- tornikeo 5mo agoI wonder if hetzner had better uptime in EU than AWS this year.
- altern8 5mo agoWhy no love for OVH? I find Hetzner's UI to be super-confusing, making it hard to manage things.
- mimischi 5mo agoAt the rate that people claim AWS us-east to go down, folks will argue that OVH has a tendency to go up in flames!
- teitoklien 5mo agoovh rocks, they are also far more customer support friendly than hetzner. Half the time hetzner feels like they are doing you a favor by letting you rent servers from them. Ovh is way simpler, and openstack integration from them works good enough for most of my needs. Kinda insane how atrocious docs are tho. No .md markdown format to let agents read stuff yet -_-
- noAnswer 5mo agoThey once had offerings for dedicated servers without hard drivers. They did network boot from NFS. So the costs where between a full dedicated server and a virtual one. Sadly it was very badly engineered. Small disk IO was so bad that you basically couldn't run MySQL. I did run a MX. For every mail postfix would complained that the filesystem did run a few secondes in the future. At some point they gave up and stuck a USB stick into every server. It was dead by thousand cuts and put a bad taste in my mouth. But I have to admit that was a long time ago and I should probably give them another chance.
- tornikeo 5mo agoWho said anything about ui? I just grab my project write key and Codex handles it all, no UI from idea to production at all.
- ElenaDaibunny 5mo ago[flagged]
- rswail 5mo agoSo in the comments here we have the usual about us-east-1, it's centralized, it's a SPOF for AWS, they should fix it, don't put your stuff there, etc. This was one data centre in one zone of a multi-zone region. Yes IAM/R53 and others are centralized there, yes, reworking those service to be decentralized and cross-region would be a Good Thing. But us-east-1 is already multi-zone (6 with a seventh marked as "coming in 2026") with multi DC within zones. From memory, when a global service like IAM is out, it's more likely to be bugs in the implementation or dependency than a "if this was cross-region it wouldn't have died" issue. But this wasn't an outage of any AWS global service this time. The only one that seemed to have more impact was/is MSK. Which is likely to be more of an issue with Kafka than anything AWS related.
- grendelt 5mo agoThere sure are a lot of eggs in that East basket.
- dnnddidiej 5mo agoI remember someone said friends dont let friends use USE1 last time and I thought that as the slack message saying USE1 and all the stuff we deploy there has gone to shit.
- jasonlotito 5mo agoIn this thread: People confusing AZs, Regions, and why people would want to use single AZs.
- 1vuio0pswjnm7 5mo agoActual CNBC title: AWS data center outage hits trading on FanDuel, Coinbase - recovery to take hours
- ptdorf 5mo agoTotally unrelated: Amazon board: institutional knowledge is the real bottleneck, guys! Let's replace these entitled seniors and staff with juniors armed with kiro. Same difference, right?