27 ms·
AWS multiple services outage in us-east-1
- qrush 1y agoAWS's own management console sign-in isn't even working. This is a huge one. :(
- igleria 1y agofunny that even if we have our app running fine in AWS europe, we are affected as developers because of npm/docker/etc being down. oh well.
- dijit 1y agoAWS has made the internet into a single-point-of failure. What's the point of all the auto-healing node-graph systems that were designed in the 70s and refined over decades: if we're just going to do mainframe development anyway?
- voidUpdate 1y agoTo be fair, there is another point of failure, Cloudflare. It seems like half the internet goes down when Cloudflare has one of their moments
- yubblegum 1y agoClourflare is not merely a single point of failure. They are the official MITM of the internet. They control the flow of information. They probably know more about your surfing habits than Google at this point. There are some sites I can not even connect to using IP addresses anymore. That company is very concerning and not because of an outage. In fact, I wish one day we have a full cloudflare outage and the entire net goes dark and it finally sink in how much control this one f'ing company has over information in our so called free society.
- voidUpdate 1y agoNot every site is controlled by Cloudflare. My site doesn't use it at all (the fact that it is currently not working properly is entirely coincidental), because I don't really see a reason to use it. Whenever they go down, I'm unaffected
- yubblegum 1y agoI'm still trying for figure out if you are being ironic but anyway thanks for the laugh!
- ryanmcdonough 1y agoNow, I may well be naive - but isn't the point of these systems that you fail over gracefully to another data centre and no-one notices?
- codeulike 1y agoI get the impression that this has been thought about to some extent, but its a constantly changing architecture with new layers and new ideas being added, so for every bit of progress there's the chance of new Single Points Of Failure being added. This time it seems to be a DNS problem with DynamoDB
- spwa4 1y agoReddit seems to be having issues too: "upstream connect error or disconnect/reset before headers. retried and the latest reset reason: connection timeout"
- martinheidegger 1y ago[dead]
- solatic 1y agoAnd yet, AMZN is up for the day. The market doesn't care. Crazy.
- al_james 1y agoCant even login via the AWS access portal.
- atymic 1y agoLooks like maybe a DNS issue? https://www.whatsmydns.net/#A/dynamodb.us-east-1.amazonaws.com https://www.whatsmydns.net/#A/dynamodb.us-east-1.amazonaws.c... Resolves to nothing.
- immibis 1y agoIt's plausible that Amazon removes unhealthy servers from all round-robins including DNS. If all servers are unhealthy, no DNS. Alternatively, perhaps their DNS service stopped responding to queries or even removed itself from BGP. It's possible for us mere mortals to tell which of these is the case.
- Nextgrid 1y agoChances are there's some cyclical dependencies. These can creep up unnoticed without regular testing, which is not really possible at AWS scale unless they want to have regular planned outages to guard against that.
- theshrike79 1y agoIt's always DNS.
- Sparkyte 1y agoMaybe they forgot to pay the bills.
- lexandstuff 1y agoYes, we're seeing issues with Dynamo, and potentially other AWS services. Appears to have happened within the last 10-15 minutes.
- atymic 1y agoYep, first alert for us fired @ 2025-10-20T06:55:16Z
- jimrandomh 1y agoThe RDS proxy for our postgres DB went down.
- comp_throw7 1y agoWe're seeing issues with RDS proxy. Wouldn't be surprised if a DNS issue was the cause, but who knows, will wait for the postmortem.
- romanhotsiy 1y agoWe're also seeing issues with Lambda and RDS proxy endpoint.
- richardwardza 1y agoWe changed our db connection settings to go direct to the db and that's working. Try taking the proxy out the loop
- qianli_cs 1y agoWe're seeing issues with multiple AWS services https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status
- atymic 1y agohttps://news.ycombinator.com/item?id=45640754 https://news.ycombinator.com/item?id=45640754
- thundergolfer 1y agoThis is widespread. ECR, EC2, Secrets Manager, Dynamo, IAM are what I've personally seen down.
- mpcoder 1y agoI can't even see my EKS clusters
- mopatches 1y agoSimilar: our EC2s are gone.
- imstil3earning 1y agoeverything is gone :(
- ZeWaka 1y agoAlexa devices are also down.
- nvarsj 1y agoAnd ring! Don’t know why the chime needs an AWS connection. That was surprising.
- Aachen 1y agoSignal is down from several vantage points and accounts in Europe, I'd guess because of this dependence on Amazon overseas We're having fun figuring out how to communicate amongst colleagues now! It's when it's gone when you realise your dependence
- lexandstuff 1y agoThankfully Slack is still holding up.
- netsharc 1y ago[flagged]
- philipallstar 1y agoAnd the post office still works, so ah, at least kidnappers can send ransom demands.
- tubs 1y agoIt’s super broken for me. Random threads no longer appear.
- LostMyLogin 1y agoIt’s acting up for me but wondering if it’s unrelated. Imagines failing to post and threads acting strange.
- tomwojcik 1y agoSame. My Slack mobile app managed to sync the new messages, but it took it about 30 seconds, while usually it's sub 2 seconds.
- rhdunn 1y agoSlack is having issues with huddles, canvas, and messaging per https://slack-status.com/ https://slack-status.com/. Earlier it was just huddles and canvas.
- nik736 1y agoTwilio seems to be affected as well
- richardwardza 1y agoTheir entire status page is red!
- rickette 1y agoCouple of years ago us-east was considered the least stable region here on HN due to its age. Is that still a thing?
- sammy2255 1y agoYes
- motiejus 1y agoWhen I was there at aws (left about a decade ago), us-east-1 was considered least stable, because it was the biggest. I.e. some bottle-necks in new code appearing only _after_ you've deployed there, which is of course too late. It didn't help that some services had their deploy trains (pipelines in amazon lingo) of ~3 weeks, with us-east-1 being the last one. I bet the situation hasn't changed much since.
- shawabawa3 1y ago>It didn't help that some services had their deploy trains (pipelines in amazon lingo) of ~3 weeks, with us-east-1 being the last one. oof, so you're saying this outage could be cause by a change merged 3 weeks ago?
- immibis 1y agoCouple of weeks or months ago the front page was saying how us-east-1 instability was a thing of the past due to <whatever chang of architecture>.
- esskay 1y agoYup, never add anything new to us-east-1. There is never a good reason to willingly use that region.
- roosgit 1y agoCan confirm. I was trying to send the newsletter (with SES) and it didn't work. I was thinking my local boto3 was old, but I figured I should check HN just in case.
- montek01singh 1y agoI cannot create a support ticket with AWS as well.
- starkindustries 1y agoZoom is unable to send screenshots.
- starkindustries 1y agozoom unable to send messages now as well.
- sammy2255 1y agoCan't resolve any records for dynamodb.us-east-1.amazonaws.com However, if you desperately need to access it you can force resolve it to 3.218.182.212. Seems to work for me. DNS through HN curl -v --resolve "dynamodb.us-east-1.amazonaws.com:443:3.218.182.212" https://dynamodb.us-east-1.amazonaws.com/ https://dynamodb.us-east-1.amazonaws.com/
- rstupek 1y agothank you for that info!!!!!
- sam1r 1y agoDude!! Life saver.
- XCSme 1y agoIt's always DNS
- alex_suzuki 1y agoConfirmed. > Based on our investigation, the issue appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1.
- planckscnst 1y agoThere's also dynamodb-fips.us-east-1.amazonaws.com if the main endpoint is having trouble. I'm not sure if this record was affected the same way during this event.
- pageandrew 1y agoCan't even get STS tokens. RDS Proxy is down, SQS, Managed Kafka.
- bpye 1y agoAmazon.ca is degraded, some product pages load but can't see prices. Amusing.
- ares623 1y agoDid someone vibe code a DNS change
- renatovico 1y agodocker hub or github cache internal maybe is affected: Booting builder /usr/bin/docker buildx inspect --bootstrap --builder builder-1c223ad9-e21b-41c7-a28e-69eea59c8dac #1 [internal] booting buildkit #1 pulling image moby/buildkit:buildx-stable-1 #1 pulling image moby/buildkit:buildx-stable-1 9.6s done #1 ERROR: received unexpected HTTP status: 500 Internal Server Error ------ > [internal] booting buildkit: ------ ERROR: received unexpected HTTP status: 500 Internal Server Error
- mopatches 1y agoDockerHub shows full outage: https://www.dockerstatus.com/ https://www.dockerstatus.com/
- BartjeD 1y agoSo much for the peeps claiming amazing Cloud uptime ;)
- JackSlateur 1y agoCould you give us that uptime, in number ?
- BartjeD 1y agoI'm afraid you missed the emoticon at the end of the sentence. A `;)` is normally understoond to mean the author isn't entirely serious, and is making light of something or other. Perhaps you American downvoters were on call and woke up with a fright, and perhaps too much time to browse Hacker News. ;)
- deleted 1y ago[deleted]
- Ekaros 1y agoWasn't the point why AWS is so much premium that you will always get at least 6 nines if not more in availability?
- miohtama 1y ago1. Competitors are not any better, or worse 2. Trusted brand
- abujazar 1y agoLast time I checked the standard SLA is actually 99 % and the only compensation you get for downtime is a refund. Which is why I don't use AWS for anything mission critical.
- systemvoltage 1y agoWhat do you do if not AWS?
- abujazar 1y agoBeen using AWS too, but for a critical service we mirrored across three Hetzner datacenters with master-master replication as well as two additional locations for cluster node voting.
- esskay 1y agoTheres literally thousands of options. 99% of people on AWS do not need to be on AWS. VPS servers or load balanced cloud instances from providers like Hetzner are more than enough for most people. It still baffles me how we ended up in this situation where you can almost hear peoples disapproval over the internet when you say AWS / Cloud isn't needed and you're throwing money away for no reason.
- Nextgrid 1y agoThere's nothing particularly wrong with AWS, other than the pricing premium. The key is that you need to understand no provider will actually put their ass on the line and compensate you for anything beyond their own profit margin, and plan accordingly. For most companies, doing nothing is absolutely fine, they just need to plan for and accept the occasional downtime. Every company CEO wants to feel like their thing is mission-critical but the truth is that despite everything being down the whole thing will be forgotten in a week. For those that actually do need guaranteed uptime, they need to build it themselves using a mixture of providers and test it regularly. They should be responsible for it themselves, because the providers will not. The stuff that is actually mission-critical already does that, which is why it didn't go down.
- hexbin010 1y agoWhy after all these years is us-east-1 such a SPOF?
- nikolay 1y agoChoosing us-east-1 as your primary region is good, because when you're down, everybody's down, too. You don't get this luxury with other US regions!
- tokioyoyo 1y agoDoing pretty well up here in Tokyo region for now! Just can't log into console and some other stuff.
- happymellon 1y agoCheck the URL, we had an issue a couple of years ago with the Workspaces. US East was down but all of our stuff was in EU. Turns out the default URL was hardcoded to use the us east interface and just by going to workspaces and then editing your URL to be the local region got everyone working again. Unless you mean nothing is working for you at the moment.
- thdhhghgbhy 1y agoDoesn't this mean you are not regionally isolated from us-east-1?
- Sparkyte 1y agoI am down with that lets all build in US-East-1.
- sam1r 1y agoSometimes we all need a tech shutdown.
- sunrunner 1y agoAs they say, every cloud outage has a silver lining. * Give the computers a rest, they probably need it. Heck, maybe the Internet should just shut down in the evening so everyone can go to bed (ignoring those pesky timezone differences) * Free chaos engineering at the cloud provider region scale, except you didn't opt in to this one and know about in advance, making it extra effective * Quickly figure out a map which of the things you use have a dependency on a single AWS region without no capability to change or re-route
- systemvoltage 1y agoworkos is down too, timing is highly correlated with AWS outage: https://status.workos.com/ https://status.workos.com/ That means Cursor is down, can't login.
- rdm_blackhole 1y agoVercel functions are down as well.
- gianpaj 1y agoYes https://www.vercel-status.com/ https://www.vercel-status.com/
- __coder__ 1y agoPerplexity also have outage. https://status.perplexity.ai https://status.perplexity.ai
- abujazar 1y agoI find it interesting that AWS services appear to be so tightly integrated that when there's an issue in a region, it affects most or all services. Kind of defeats the purported resiliency of cloud services.
- tokioyoyo 1y agoYou know how people say X startup is ChatGPT wrapper? A significant chunk of AWS services are wrappers of main services (DynamoDB, EC2, S3 and etc).
- abujazar 1y agoYes, and that's exactly the problem. It's like choosing a microservice architecture for resiliency and building all the services on top of the same database or message queue without underlying redundancy.
- pm90 1y agoafaik they have a tiered service architecture, where tier 1 services are allowed to rely on tier 0 services but not vice-versa, and have a bunch of reliability guarantees on tier 0 services that are higher than tier 1. It is kinda cool that the worst aws outages are still within a single region and not global.
- UltraSane 1y agoThere IS a huge amount of redundancy built into the core services but nothing is perfect.
- Aperocky 1y agoDNS is always the single point of failure. But I think what wasn't well considered was the async effect - If something is gone for 5 minutes, maybe it will be just fine, but when things are properly asynchronous, then the workflows that have piled up during that time becomes a problem in itself. Worst case, they turn into poison pills which then break the system again.
- yuvadam 1y agoDuring the last us-east-1 apocalypse 14 years ago, I started awsdowntime.com - don't make me regsiter it again and revive the page.
- Sparkyte 1y agoI remember the one where some contractor accidentally cut the trunk between AZes.
- deleted 1y ago[deleted]
- oasisbob 1y agoIn us-east-1? That doesn't sound that impactful, have always heard that us-east-1's network is a ring. Back before AWS provided transparency into AZ assignments, it was pretty common to use latency measurements to try and infer relative locality and mappings of AZs available to an account.
- Sparkyte 11mo agoIt was a long time ago. It was very impactful at the time. You could still reach the other AZes but that one. I think this was US-East1c.
- padjo 1y agoFriends don’t let friends use us-east-1
- seviu 1y agoI can't log in to my AWS account, in Germany, on top of that it is not possible to order anything or change payment options from amazon.de. No landing page explaining services are down, just scary error pages. I thought account was compromised. Thanks HN for, as always, being the first to clarify what's happening. Scary to see that in order to order from Amazon Germany, us-east1 must be up. Everything else works flawlessly but payments are a no go.
- that_guy_iain 1y agoI just ordered stuff from Amazon.de. And I highly any Amazon site can go down because of one region. Just like Netflix are rarely affected.
- seviu 1y agoI can’t even login, I get the internal error treatment. This is on Amazon.de
- that_guy_iain 1y agoI'm on Amazon.de and I literally ordered stuff seconds before posting the comment. They took the money and everything. The order is in my order history list.
- serial_dev 1y agoI wanted to log into my Audible account after a long time on my phone, I couldn't, started getting annoyed, maybe my password is not saved correctly, maybe my account was banned, ... Then checking desktop, still errors, checking my Amazon.de, no profile info... That's when I started suspecting that it's not me, it's you, Amazon! Anyway, I guess, I'll listen to my book in a couple of hours, hopefully. Btw, most parts of the amazon.de is working fine, but I can't load profiles, and can't login.
- echelon_musk 1y agoYou might be interested in Libation [0]. I use it to de-DRM my Audible library and generate a cue sheet with chapters in for offline listening. [0] https://getlibation.com/ https://getlibation.com/
- jumploops 1y ago"Never choose us-east-1"
- ta1243 1y agoNever choose a single point of failure. Or rather Ensure your single point of failure risk is appropriate for your business. I don't have full resilience for my companies AS going down, but we do have limited DR capability. Same with the loss of a major city or two. I'm not 100% confident in a Thames Barrier flood situation, as I suspect some of our providers don't have the resilience levels we do, but we'd still be able to provide some minimal capability.
- ctbellmar 1y agoVarious AI services (e.g. Perplexity) are down as well
- rvz 1y agoJust tried Perplexity and it has no answer. Damn, this is really bad. Looking forward to the postmortem.
- bartvk 1y agoI don't like how they phrased it. From the Verge: “Perplexity is down right now,” Perplexity CEO Aravind Srinivas said on X. “The root cause is an AWS issue. We’re working on resolving it.” What he should have said, IMHO, is "The root cause is that Perplexity fully depends on AWS." I wonder if they're actually working on resolving that, or that they're just waiting for AWS to come back up.
- tosh 1y agoseeing issues with SES in us-east-1 as well
- grebc 1y agoGood thing hyperscalers provide 100% uptime.
- empressplay 1y agoCan't check out on Amazon.com.au, gives error page
- deleted 1y ago[deleted]
- kondro 1y agoThis link works fine from Australia for me.
- sam1r 1y agoChime has completely been down for almost 12 hours. Impacting all banking series with red status error. Oddly enough, only their direct deposits are functioning without issues. https://status.chime.com/ https://status.chime.com/
- nodesocket 1y agoAffecting Coinbase[1] as well, which is ridiculous. Can't access the web UI at all. At their scale and importance they should be multi-region if not multi-cloud. [1] https://status.coinbase.com https://status.coinbase.com
- bradhe 1y agoSeems the underlying issue is with DynamoDB, according to the status page, which will have a big blast radius in other services. AWS' services form a really complicated graph and there's likely some dependency, potentially hidden, on us-east-1 in there.
- Splizard 1y agoThe issue appears to be cascading internationally due to internal dependencies on us-east-1
- port3000 1y agoEven railway's status page is down (guess they use Vercel): https://railway.instatus.com/ https://railway.instatus.com/
- ta1243 1y agoMeanwhile my pair of 12 year old raspberry pi's hangling my home services like DNS survive their 3rd AWS us-east-1 outage. "But you can't do webscale uptime on your own" Sure. I suspect even a single pi with auto-updates on has less downtime.
- deleted 1y ago[deleted]
- __alexs 1y agoIs there any data on which AWS regions are most reliable? I feel like every time I hear about an AWS outage it's in us-east-1.
- nodesocket 1y agoI don't recommend to my clients they use us-east-1. It's the oldest and most prone to outages. I usually always recommend us-east-2 (Ohio) unless they require West Coast.
- dr-smooth 1y agoand if they need West Coast, it's us-west-2. I consider us-west-1 to be a failed region. They don't get some of the new instance types, you can't get three AZs for your VPCs, and they're more expensive than the other US regions.
- bradhe 1y agous-east-1 was, probably still is, AWS' most massive deployment. Huge percentage of traffic goes through that region. Also, lots of services backhaul to that region, especially S3 and CloudFront. So even if your compute is in a different region (at Tower.dev we use eu-central-1 mostly), outages in us-east-1 can have some halo effect. This outage seems really to be DynamoDB related, so the blast radius in services affected is going to be big. Seems they're still triaging.
- philipp-gayret 1y agoOur Alexa's stopped responding and my girl couldn't log in to myfitness pal anymore.. Let me check HN for a major outage and here we are :^) At least when us-east is down, everything is down.
- kryptn 1y agoWonder if this is related https://www.dockerstatus.com/pages/533c6539221ae15e3f000031 https://www.dockerstatus.com/pages/533c6539221ae15e3f000031
- Titan2189 1y agoYup > We have identified the underlying issue with one of our cloud service providers.
- ngruhn 1y agoCan't login to Jira/Confluence either.
- Xenoamorphous 1y agoSeems to work fine for me. I'm in Europe so maybe connecting to some deployment over here.
- danias 1y agoYou are already logged in. If you try to access your account settings, for example, you will be disappointed...
- roschdal 1y ago[flagged]
- tietjens 1y agoThis is such an HN response. Oh, no problem, I'll just avoid the internet for all of my important things!
- jjcob 1y agoDoor locks, heating and household appliances should probably not depend on Internet services being available.
- tommit 1y agoDo you not have a self-hosted instance of every single service you use? :/
- deleted 1y ago[deleted]
- pinkgolem 1y agoNo. But for the important ones, yes I do. Everything in and around my house is working fully offline
- dukeyukey 1y agoThey are probably being sarcastic.
- rirze 1y agoNot very helpful. I wanted to make a very profitable trade but can’t login to my brokerage. I’m losing about ~100k right now.
- fragmede 1y agowhat's the trade?
- dude250711 1y agoThey are amazing at LeetCode though.
- tokioyoyo 1y agoLet's be nice. I'm sure devs and ops are on fire right now, trying to fix the problems. Given the audience of HN, most of us could have been (have already been?) in that position.
- dude250711 1y agoThey choose their hiring-retention practices and they choose to provide global infrastructure, when is the good time to criticise them? Granted, they are not as drunk on LLM as Google and Microsoft. So, at least we can say this outage had not been vibe-coded (yet).
- fragmede 1y agohugops ftw
- rirze 1y agoNo we wouldn’t because there’s like a 50/50 chance of being a H1B/L1 at AWS. They should rethink their hiring and retention strategies.
- rwky 1y agoTo everyone that got paged (like me), grab a coffee and ride it out, the week can only get better!
- esskay 1y agoTo everyone who was supposed to get paged but didn't, do a postmortem, chances are your service is running via Twilio and needs migrating elsewhere.
- avian 1y ago> grab a coffee and ride it out The way things are today I'm thankful the coffee machine still works without AWS.
- rwky 1y agoFeel sorry for anyone with a "smart" coffee machine.
- greenavocado 1y agoWith how long it may last, pour some cold water on the coffee and come back to drink the supernatant in a few hr
- hshdhdhehd 1y agoBurn that SLO, next persons problem eh!
- amadeoeoeo 1y agoOh no... may be LaLiga found out pirates hosting on AWS?
- agos 1y agothis is how I discover that is not just Serie A doing this shenanigans. I'm not really surprised
- sofixa 1y agoAll the big leagues take "piracy" very seriously and constantly try to clamp down on it. TV rights is one of their main revenue sources, and it's expected to always go up, so they see "piracy" as a fundamental threat. IMO, it's a fundamental misunderstanding on their side, because people "pirating" usually don't have a choice - either there is no option for them to pay for the content (e.g. UK's 3pm blackout), or it's too expensive and/or spread out. People in the UK have to pay 3-4 different subscriptions to access all local games. The best solution, by far, is what France's Ligue 1 just did (out of necessity though, nobody was paying them what they wanted for the rights after the previous debacles). Ligue 1+ streaming service, owned and operated by them which you can get access through a variety of different ways (regular old TV paid channel, on Amazon Prime, on DAZN, via Bein Sport), whichever suits you the best. Same acceptable price for all games.
- slumberlust 1y agoMore and more ads at every level every year, when will it be enough?
- miamibre 1y agoMLB in the US does the same thing for the regular season, it's awesome despite the blackouts which prevent you from watching your local team but you can get around that with a simple VPN. But alas I believe that they will be making the service part of ESPN which will undoubtedly make the product worse just like they will do with NFL Red Zone. The problem is that leagues miss out on billions of dollars of revenue when they do this AND they also have to maintain the streaming service which is way outside their technical wheelhouse. MLS also has a pretty straightforward streaming service through AppleTV which I also enjoy. What i find weird is that people complain (at least in the case of the MLS deal) that it's a BAD thing, that somehow having an easily accessible service that you just pay for and get access to without a contract or cable is diminishing popularity / discoverability of the product?
- donmb 1y agoAsana down Postman workspaces don't load Slack affected And the worst: heroku scheduler just refused to trigger our jobs
- SeanAnderson 1y agoLooks like it affected Vercel, too. https://www.vercel-status.com/ https://www.vercel-status.com/ My website is down :( (EDIT: website is back up, hooray)
- maximefourny 1y agoHave you done anything for it to be back up? Looks like mines are still down.
- LostMyLogin 1y agoLooks as if they are rerouting to a different region.
- hugh-avherald 1y agomines are generally down
- l5870uoo9y 1y agoStatic content resolves correctly but data fetching is still not functional.
- TiredOfLife 1y agoService that runs on aws is down when aws is down. Who knew.
- jellyfishbeaver 1y agoI had a chuckle on my way home yesterday. Standing on the train platform and seeing "Next departure in: (Vercel Connection Error)" on the screen. :P
- TechDebtDevin 1y agoImagine using vercel, a company that literally contributes to the starvation of children and is proud of it. Also, literally just learn to use a Dockerfile and a vps, like why do these PaaS even exist, you're paying 3x for the same AWS services.
- magnio 1y agonpm and pnpm are badly affected as well. Many packages are returning 502 when fetched. Such a bad time...
- samsepia 1y agoYup, was releasing something to prod and can't even build a react app. I wonder if there is some sort of archive that isn't affected?
- JCharante 1y agoAWS CodeArtifact can act as a proxy and fetch new packages from npm when needed. A bit late for that though but sharing if you want to future proof against the yearly us-east-1 outage
- JCharante 1y agoOh damn that ruins all our builds for regions I thought would be unaffected
- alex_suzuki 1y agoPaddle (payment provider) is down as well: https://paddlestatus.com/ https://paddlestatus.com/
- mcintyre1994 1y agoPresumably the root cause of the major Vercel outage too: https://www.vercel-status.com/ https://www.vercel-status.com/
- hyruo 1y agoNo wonder, when I opened Vercel it showed a 502 error.
- colesantiago 1y agoIt seems that all the sites that ask for distributed systems in their interview and has their website down wouldn't even pass their own interview. This is why distributed systems is an extremely important discipline.
- mangamadaiyan 1y agoMaybe actually making the interviews less of a hazing ritual would help. Hell, maybe making today's tech workplace more about getting work done instead of the series of ritualistic performances that the average tech workday has degenerated to might help too. Ergo, your conclusion doesn't follow from your initial statements, because interviews and workplaces are both far more broken than most people, even people in the tech industry, would think.
- colesantiago 1y agoWell it looks like if companies and startups did their job in hiring the proper distributed systems skills more rather than hazing for the wrong skills we wouldn't be in this outage mess. Many companies on Vercel don't think to have a strategy to be resilient to these outages. I rarely see Google, Ably and others serious about distributed systems being down.
- the_mitsuhiko 1y ago> Many companies on Vercel don't think to have a strategy to be resilient to these outages. But that's the job of Vercel and it looks like they did a pretty good job. They rerouted away from the broken region.
- rester324 1y agoThere was a huuuge GCP outage just a few months back: https://news.ycombinator.com/item?id=44260810 https://news.ycombinator.com/item?id=44260810
- dist-epoch 1y agodistributed systems != continuous uptime
- ksajadi 1y agoA lot of status pages hosted by Atlasian StatusPage are down! The irony…
- atonse 1y agoI can't believe this. When status page first created their product, they used to market how they were in multiple providers so that they'd never be affected by downtime. Maybe all that got canned after the acquisition?
- mumber_typhoon 1y ago>Oct 20 12:51 AM PDT We can confirm increased error rates and latencies for multiple AWS Services in the US-EAST-1 Region. This issue may also be affecting Case Creation through the AWS Support Center or the Support API. We are actively engaged and working to both mitigate the issue and understand root cause. We will provide an update in 45 minutes, or sooner if we have additional information to share. Weird that case creation uses the same region as the case you'd like to create for.
- easton 1y agoThe support apis only exist in us-east-1, iirc. It’s a “global” service like IAM, but that usually means modifications to things have to go through us-east-1 even if they let you pull the data out elsewhere.
- transitivebs 1y agocan't log into https://amazon.com https://amazon.com either after logging out; so many downstream issues
- trusche 1y agoBoth Intercom and Twilio are affected, too. - https://status.twilio.com/ https://status.twilio.com/ - https://www.intercomstatus.com/us-hosting https://www.intercomstatus.com/us-hosting I want the web ca. 2001 back, please.
- Xenoamorphous 1y agoSlack now failing for me.
- jcmeyrignac 1y agoImpossible to connect to JIRA here (France).
- baumschubser 1y agoLogin issues with Jira Cloud here in Germany too. Just a week after going from Jira on-prem to cloud
- binsquare 1y agoDon't miss this
- jug 1y agoOf course this happens when I take a day off from work lol Came here after the Internet felt oddly "ill" and even got issues using Medium, and sure enough https://status.medium.com https://status.medium.com
- JCharante 1y agoRing is affected. Why doesn’t Ring have failover to another region?
- voxadam 1y agoThat's understandably bad for anyone who depends on Ring for security but arguably a net positive for the rest of us. Amazon’s Ring to partner with Flock: https://news.ycombinator.com/item?id=45614713 https://news.ycombinator.com/item?id=45614713
- lsllc 1y agoThe Ring (Doorbell) App isn't working, nor is any the MBTA (Transit) Status pages/apps.
- LostMyLogin 1y agoMy apartment uses “SmartRent” for access controls and temps in our unit. It’s down…
- arch-choot 1y agoSo there's no way to get back in if you step out for food?
- mittermayr 1y agoCareful: NPM _says_ they're up (https://status.npmjs.org/ https://status.npmjs.org/) but I am seeing a lot of packages not updating and npm install taking forever or never finishing. So hold off deploying now if you're dependent on that.
- drinchev 1y agoAlso npm audit times out.
- gjvr 1y agoYep. It's the auditing part that is broken. As a (dangerous) workaround use --no-audit
- olex 1y agoThey've acknowledged an issue now on the status page. For me at least, it's completely down, package installation straight up doesn't work. Thankfully current work project uses a pull-through mirror that allows us to continue working.
- tonyhart7 1y ago"Thankfully current work project uses a pull-through mirror that allows us to continue working." so there is no free coffee time???? lmao
- binsquare 1y agoThe internal disruption reviews are going to be fun :)
- Msurrow 1y agoThe fun is really gonna start if the root cause of this somehow implicates an AI as a primary cause.
- karel-3d 1y agoI haven't seen the "90% of our code is AI" nonsense from Amazon.
- gdulli 1y agoTheir business doesn't depend on selling AI so they have the luxury of not needing to be in on that grift.
- some_furry 1y agohttps://archive.is/b6aUD https://archive.is/b6aUD
- aurareturn 1y agoIt's never an AI's fault since it's up to a human to implement the AI and put in a process that prevents this stuff from happening. So blame humans even if an AI wrote some bad code.
- Msurrow 1y agoI agree but then again it’s always a humans fault in the end. So probably a root cause will have a bit more neuance. I was more thinking of the possible headlines and how that would potentially affect the public AI debate. Since this event is big enough to actually get the attention of eg risk management at not-insignificant orgs.
- 1y ago
- codebolt 1y agoAtlassian cloud is having problems as well.
- seanieb 1y agoClearly this is all some sort of mass delusion event, the Amazon Ring status says everything is working. https://status.ring.com/ https://status.ring.com/ (Useless service status pages are incredibly annoying)
- storgaard 1y agoAtlassian is down as well so they probably can't access their Atlassian Statuspage admin panel to update it.
- netdevphoenix 1y agoWhen you know a service is down but the service says it's up: it's either your fault or the service is having a severe issue
- ArcHound 1y agoGood luck to all on-callers today. It might be an interesting exercise to map how many of our services depend on us-east-1 in one way or another. One can only hope that somebody would do something with the intel, even though it's not a feature that brings money in (at least from business perspective).
- hipratham 1y agoStrangely some of our services are scaling up on east-1, and there is downtick on downdetector.com so issue might be resolving.
- cpfleming 1y agoSeems to be upsetting Slack a fair bit, messages taking an age to send and OIDC login doesn't want to play.
- AtNightWeCode 1y agoConsidering the history of east-1 it is fascinating that it still causes so many single point of failure incidents for large enterprises.
- askonomm 1y agoDocker is also down.
- 1659447091 1y agoAlso: Snapchat, Ring, Roblox, Fortnite and more go down in huge internet outage: Latest updates https://www.the-independent.com/tech/snapchat-roblox-duolingo-fortnite-down-not-working-b2848289.html https://www.the-independent.com/tech/snapchat-roblox-duoling... To see more (from the first link): https://downdetector.com https://downdetector.com
- tedk-42 1y agoInternet, out. Very big day for an engineering team indeed. Can't vibe code your way out of this issue...
- rvz 1y ago> Can't vibe code your way out of this issue... Exactly. This time, some LLM providers are also down and can't help vibe coders on this issue.
- fragmede 1y agoQwen3 on lm-studio running fine on my work Mac M3, what's wrong with yours?
- LostMyLogin 1y agoPour one out for everyone on-call right now.
- xvector 1y agoAfter some thankless years preventing outages for a big tech company, I will never take an oncall position again in my life. Most miserable working years I have had. It's wild how normalized working on weekends and evenings becomes in teams with oncall. But it's not normal. Our users not being able to shitpost is simply not worth my weekend or evening. And outside of Google you don't even get paid for oncall at most big tech companies! Company losing millions of dollars an hour, but somehow not willing to pay me a dime to jump in at 3AM? Looks like it's not my problem!
- scns 1y ago> And outside of Google you don't even get paid for oncall at most big tech companies. What the redacted?
- redwall_hp 1y ago
- mslm 1y agoHappened to be updating a bunch of NPM dependencies and then saw `npm i` freeze and I'm like... ugh what did I do. Then npm login wasn't working and started searching here for an outage, and wala.
- skolsuper 1y agovoila
- kalleboo 1y agoIt's fun watching their list of "Affected Services" grow literally in front of your eyes as they figure out how many things have this dependency. It's still missing the one that earned me a phone call from a client.
- zenexer 1y agoIt's seemingly everything. SES was the first one that I noticed, but from what I can tell, all services are impacted.
- deleted 1y ago[deleted]
- hvb2 1y agoIn AWS, if you take out one of dynamo db, S3 or lambda you're going to be in a world of pain. Any architecture will likely use those somewhere including all the other services on top. If in your own datacenter your storage service goes down, how much remains running
- goatking 1y agoAgreed, but you can put EC2 on that list as well
- mlrtime 1y ago
- robertpohl 1y agoLooks like we're back!
- DataDaemon 1y agoBut but this is a cloud, it should exist in the cloud.
- stavros 1y agoAWS truly does stand for "All Web Sites".
- circadian 1y agoBGP (again)?
- klon 1y agoStatuspage.io seems to load (but is slow) but what is the point if you can't post an incident because Atlassian ID service is down.
- edtech_dev 1y agoSignal is also down for me.
- chaidhat 1y agois this why docker is down?
- dolibasija 1y agoyes hub.docker.com. 75 IN CNAME elb-default.us-east-1.aws.dckr.io.
- sph 1y ago10:30 on a Monday morning and already slacking off. Life is good. Time to touch grass, everybody!
- saejox 1y agoAWS has been the backbone of the internet. It is single point of failure most websites. Other hosting services like Vercel, package managers like npm, even the docker registeries are down because of it.
- pmig 1y agoThanks god we built all our infra on top of EKS, so everything works smoothly =)
- whatsupdog 1y agoI can not login to my AWS account. And, the "my account" on regular amazon website is blank on Firefox, but opens on Chrome. Edit: I can login into one of the AWS accounts (I have a few different ones for different companies), but my personal which has a ".edu" email is not logging in.
- fujigawa 1y agoAppears to have also disabled that bot on HN that would be frantically posting [dupe] in all the other AWS outage threads right about now.
- altairprime 1y agoThat’s done by human beings.
- bootsmann 1y agoApparently hiring 1000s of software engineers every month was load bearing
- mrcsharp 1y agoBitbucket seems affected too [1]. Not sure if this status page is regional though. [1] https://bitbucket.status.atlassian.com/incidents/p20f40pt1rgv https://bitbucket.status.atlassian.com/incidents/p20f40pt1rg...
- deleted 1y ago[deleted]
- gramakri2 1y agonpm registry also down
- thomas_witt 1y agoSeems to be really only in us-east-1, DynamoDB is performing fine in production on eu-central-1.
- fairity 1y agoAs this incident unfolds, what’s the best way to estimate how many additional hours it’s likely to last? My intuition is that the expected remaining duration increases the longer the outage persists, but that would ultimately depend on the historical distribution of similar incidents. Is that kind of data available anywhere?
- greybeard69 1y agoTo my understanding the main problem is DynamoDB being down, and DynamoDB is what a lot of AWS services use for their eventing systems behind the scenes. So there's probably like 500 billion unprocessed events that'll need to get processed even when they get everything back online. It's gonna be a long one.
- jewba 1y ago500 billions events. Always blows my mind how many people use aws
- Implicated 1y agoI know nothing. But I'd imagine the number of 'events' generated during this period of downtime will eclipse that number every minute.
- zimpenfish 1y ago"I felt a great disturbance in us-east-1, as if millions of outage events suddenly cried out in terror and were suddenly silenced" (Be interesting to see how many events currently going to DynamoDB are actually outage information.)
- nicce 1y agoI wonder how many companies have properly designed their clients. So that the timing before re-attempt is randomised and the re-attempt timing cycle is logarithmic.
- thomas_witt 1y agoDynamoDB is performing fine in production in eu-central-1. Seems to be really limited to us-east-1 (https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status). I think they host a lot of console and backend stuff there.
- okr 1y agoYet. Everything goes down the ... Bach ;)
- OhioMan2943 1y agoCan't log into tidal for my music
- antihero 1y agoNavidrome seems fine
- codegladiator 1y agoThey haven't listed SES there yet in the affected services on their status page
- OhioMan2943 1y agoIt's weird that we're living in a time where this could be a taste of a prolonged future global internet blackout by adversarial nations. Get used to this feeling I guess :)
- kedihacker 1y agoOnly us east 1 gets new services immediately others might do but not a guarantee. Which regions are a good alternative
- frays 1y agoRobinhood's completely down. Even their main website: https://robinhood.com/ https://robinhood.com/
- deleted 1y ago[deleted]
- mittermayr 1y agoAmazing, I wonder what their interview process is like, probably whiteboarding a next-gen LLM in WASM, meanwhile, their entire website goes down with us-east-1... I mean.
- 1-6 1y agoFriends tell their friends about more mature brokerages once the account goes over $100k.
- zwnow 1y agoI love this to be honest. Validates my anti cloud stance.
- flanked-evergl 1y agoNo service that does not run on cloud has ever had outages.
- Nextgrid 1y agoBut at least a service that doesn't run on cloud doesn't pay the 1000% premium for its supposed "uptime".
- zwnow 1y agoAt least its in my control :)
- speedgoose 1y agoNot having control or not being responsible are perhaps major selling points of cloud solutions. To each their own, I also rather have control than having to deal with a cloud provider support as a tiny insignificant customer. But in this case, we can take a break and come back once it's fixed without stressing.
- zwnow 1y agoBusinesses not taking responsibility for their own business should not exist in the first place...
- flanked-evergl 1y agoNo business is fully integrated. Doing so would be dumb and counter productive.
- zwnow 1y ago
- emrodre 1y agoTheir status page (https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status) says the only disrupted service is DynamoDB, but it's impacting 37 other services. It is amazing to see how big a blast radius a single service can have.
- thmpp 1y agoAWS engineers are trained to use their internal services for each new system. They seem to like using DynamoDB. Dependencies like this should be made transparent.
- Nextgrid 1y agoNot sure why this is downvoted - this is absolutely correct. A lot of AWS services under the hood depend on others, and especially us-east-1 is often used for things that require strong consistency like AWS console logins/etc (where you absolutely don't want a changed password or revoked session to remain valid in other regions because of eventual consistency).
- bsjaux628 1y agoNot "like using", they are mandated from the top to use DynamoDB for any storage. At my org in the retail page, you needed director approval if you wanted to use a relational DB for a production service.
- myroon5 1y ago> Dependencies like this should be made transparent even internally, Amazon's dependency graph became visually+logically incomprehensible a long time ago
- stevepotter 1y agoEx employee here who built an aws service. Dynamo is basically mandated. You need like VP approval to use a relational database because of some scaling stuff they ran into historically. That sucks because we really needed a relational database and had to bend over backwards to use dynamo and all the nonsense associated with not having sql. It was super low traffic too
- glemmaPaul 1y agoLOL making one db service a central point of failure, charge gold for small compute instances. Rage about needing Multi AZ, make the costs come onto the developer/organization. But, now fail on a region level, so are we going to now have multi-country setup for simple small applications?
- philipallstar 1y agoJust don't buy it if you don't want it. No one is forced to buy this stuff.
- benterix 1y ago> No one is forced to buy this stuff. Actually, many companies are de facto forced to do that, for various reasons.
- philipallstar 1y agoHow so?
- jacquesm 1y agoCertification, for one. Governments will mandate 'x, y and/or z' and only the big providers are able to deliver.
- mlrtime 1y agoThat is not the same as mandating AWS, it just means certain levels of redundancy. There are no requirements to be in the cloud.
- jacquesm 1y agoNo, that's not what it means. It means that in order to be certified you have to use providers that in turn are certified or you will have to prove that you have all of your ducks in a row and that goes way beyond certain levels of redundancy, to the point that most companies just give up and use a cloud solution because they have enough headaches just getting their internal processes aligned with various certification requirements. Medical, banking, insurance to name just a couple are heavily regulated and to suggest that it 'just means certain levels of redundancy' is a very uninformed take.
- munchlax 1y agoNowadays when this happens it's always something. "Something went wrong." Even the error message itself is wrong whenever that one appears.
- urbandw311er 1y agoDisplaying and propagating accurate error messages is an entire science unto itself... ...I can see why it's sometimes sensible to invest resource elsewhere and fall back to 'something'.
- munchlax 1y agoIMHO if error handling is rocket science, the error is you
- urbandw311er 1y agoPerhaps you're not handling enough errors ;-)
- foobar1962 1y agoI use the term “unexpected error” because if the code got to this alert it wasn’t caught by any traps I’d made for the “expected” errors.
- zigzag312 1y agoReddit shows: "Too many requests. Your request has been rate limited, please take a break for a couple minutes and try again."
- world2vec 1y agoSlack and Zoom working intermittently for me
- werdl 1y agoLooks like a DNS issue - dynamodb.us-east-1.amazonaws.com is failing to resolve.
- lgats 1y ago"Based on our investigation, the issue appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1." it seems they found your comment
- socalgal2 1y agoAmazon itself apperas to be out for some products. I get a "Sorry, We couldn't find that page" when clicking on products
- andreygubarev 1y agohttps://status.tailscale.com/ https://status.tailscale.com/ clients' auth down :( what a day
- kondro 1y agoThat just says the homepage and knowledge base are down and that admin access specifically isn’t effected.
- andreygubarev 1y agoyep, admin panel works, but in practice my devices are logged out and there is no way to re-authorize them.
- kondro 1y agoI can authenticate my devices just fine.
- andreygubarev 1y agointeresting, which auth provider u are using? browser based auth via Google wasn't working for me. tailscale used as jumphost for private subnets in aws, and... so that was painful incident as access to corp resources is mandatory for me
- kondro 1y agoI was using browser-based auth via Google.
- thecopy 1y agoI did get 500 error from their public ECR too
- goodegg 1y agoHappy Monday People
- XCSme 1y agoYeah, noticed from Zoom: https://www.zoomstatus.com/incidents/yy70hmbp61r9 https://www.zoomstatus.com/incidents/yy70hmbp61r9
- goodegg 1y agoTerraform Cloud is having problem too
- rirze 1y agoWe just had a power outage in Ashburn starting at 10 pm Sunday night. It restored at 3:40am ish, and I know datacenters have redundant power sources but the timing is very suspicious. The AWS outage supposedly started at midnight
- OliverGuy 1y agoTheir latest update on the status page says it's a Dynamodb DNS issue
- shawabawa3 1y agobut the cause of that could be anything, including some kind of config getting wiped due to a temporary power outage
- Hilift 1y agoEven with redundancy, the response time between NYC and Amazon East in Ashburn is something like 10 ms. The impedance mismatch and dropped packets and increased latency would doom most organizations craplications.
- franktankbank 1y ago> craplications LOL
- tdiff 1y agoThat strange feeling of the world getting cleaner for a while without all these dependant services.
- andrewinardeer 1y agoSignal not working here for me in AU
- moribvndvs 1y agoSo, uh, over the weekend I decided to use the fact that my company needs a status checker/page to try out Elixir + Phoenix LiveView, and just now I found out my region is down while tinkering with it and watching Final Destination. That’s a little too on the nose for my comfort.
- lbreakjai 1y agoWell at least you don't have to figure out how to test your setup locally.
- hobo_mark 1y agoWhen did Snapchat move out of GCP?
- dijit 1y agoThey might have an implicit dependency on AWS, even if they're not primarily hosted there.
- freeqaz 1y agoSince I'm 5+ years out from my NDA around this stuff, I'll give some high level details here. Snapchat heavily used Google AppEngine to scale. This was basically a magical Java runtime that would 'hot path split' the monolithic service into lambda-like worker pools. Pretty crazy, but it worked well. Snapchat leaned very heavily on this though and basically let Google build the tech that allowed them to scale up instead of dealing with that problem internally. At one point, Snap was >70% of all GCP usage. And this was almost all concentrated on ONE Java service. Nuts stuff. Anyway, eventually Google was no longer happy with supporting this and the corporate way of breaking up is "hey we're gonna charge you 10x what did last year for this, kay?" (I don't know if it was actually 10x. It was just a LOT more) So began the migration towards Kubernetes and AWS EKS. Snap was one of the pilot customers for EKS before it was generally available, iirc. (I helped work on this migration in 2018/2019) Now, 6+ years later, I don't think Snap heavily uses GCP for traffic unless they migrated back. And this outage basically confirms that :P
- garbthetill 1y agoThats so interesting to me, I always assume companies like google who have "unlimited" dollars will always be happy to eat the cost to keep customers, especially given gcp usage outside googles internal services is way smaller compared to azure and aws. Also interesting to see snapchat had a hacky solution with AppEngine
- makeitdouble 1y agoThe "unlimited dollars" come from somewhere after all. GCP is behind in market share, but has the incredible cheat advantage of just not being Amazon. Most retailers won't touch Amazon services with a ten foot pole, so the choice is GCP or Azure. Azure is way more painful for FOSS stacks, so GCP has its own area with only limited competition.
- antihero 1y agoMy website on the cupboard laptop is fine.
- deleted 1y ago[deleted]
- goinggetthem 1y agoThis is from Amazon's latest earnings call when Andy Jessy was asked why they aren't growing as much as there competitors "I think if you look at what matters to customers, what they care they care a lot about what the operational performance is, you know, what the availability is, what the durability is, what the latency and throughput is of of the various services. And I think we have a pretty significant advantage in that area." also "And, yeah, you could just you just look at what's happened the last couple months. You can just see kind of adventures at some of these players almost every month. And so very big difference, I think, in security."
- JCM9 1y agoWell that aged well
- spprashant 1y agoThat was a bit of ramble. Not something I d expect from a CEO of AWS who probably handles press all the time. It reminds me of that viral clip from a beauty pageant where the contestant went on a geographical ramble while the question was about US education.
- ssehpriest 1y agoAirtable is down as-well. A lot of businesses have all their workflows depending on their data on airtable.
- polaris64 1y agoIt looks like DNS has been restored: dynamodb.us-east-1.amazonaws.com. 5 IN A 3.218.182.189
- miyuru 1y agoI wonder if the new endpoint was affected as well. dynamodb.us-east-1.api.aws
- rafa___ 1y ago"Oct 20 2:01 AM PDT We have identified a potential root cause for error rates for the DynamoDB APIs in the US-EAST-1 Region. Based on our investigation, the issue appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1..." It's always DNS...
- gadders 1y agoSubstack seems to by lying about their status: https://substack.statuspage.io/ https://substack.statuspage.io/
- littlecranky67 1y agoJust a couple of days ago in this HN thread [0] there were quite some users claiming Hetzner is not an options as their uptime isn't as good as AWS, hence the higher AWS pricing is worth the investment. Oh, the irony. [0]: https://news.ycombinator.com/item?id=45614922 https://news.ycombinator.com/item?id=45614922
- bigblind 1y agoI don't have an opinion either way, but for now, this is just anecdotal evidence.
- brazukadev 1y agoLooks fine for pointing an irony
- bigblind 1y agoIn some ways yes. But in some ways this is like saying it's more likely to rain on your wedding day.
- DataDaemon 1y agoFinally IT managers will start understanding that cloud is no difference than Hetzner.
- aembleton 1y agoWhen things go wrong, you can point at a news article and say its not just us that have been affected.
- zimpenfish 1y agoI tried that but Slack is broken and the message hasn't got through yet...
- benterix 1y ago
- stepri 1y ago“Based on our investigation, the issue appears to be related to DNS resolution of the DynamoDB API endpoint in US-EAST-1. We are working on multiple parallel paths to accelerate recovery.” It’s always DNS.
- commandersaki 1y agoSomeone probably failed to lint the zone file.
- huflungdung 1y ago[dead]
- DrewADesign 1y agoDNS strikes me as the kind of solution someone designed thinking “eh, this is good enough for now. We can work out some of the clunkiness when more organizations start using the Internet.” But it just ended up being pretty much the best approach indefinitely.
- cindyllm 1y ago[dead]
- movpasd 1y agoSeems like an example of "worse is better". The worse solution has better survival characteristics (on account of getting actually made).
- DrewADesign 1y agoI wouldn’t say it’s the worst… a largely decentralized worldwide namespace is not an easy thing to tackle and for the most part it totally works.
- ifwinterco 1y ago
- tomaytotomato 1y agoSlack, Jira and Zoom are all sluggish for me in the UK
- kalleboo 1y agoI wonder if that's not due to dependencies on AWS but all-hands-on-deck causing far more traffic than usual
- disposable2020 1y agoI seem to recall other issues around this time in previous years. I wonder if this is some change getting shoe-horned in ahead of some reinvent release deadline...
- devttyeu 1y agoCan't update my selfhosted HomeAssistant because HAOS depends on dockerhub which seems to be still down.
- grk 1y agoDoes anyone know if having Global Accelerator set up would help right now? It's in the list of affected services, I wonder if it's useful in scenarios like this one.
- XorNot 1y agoWell that takes down Docker Hub as well it looks like.
- AshLeece 1y agoYep, was just thinking the same when my Kubernetes failed a HelmRelease due to a pull error…
- killingtime74 1y agoSignal is down for me
- miduil 1y agoYes. https://status.signal.org/ https://status.signal.org/ > Signal is experiencing technical difficulties. We are working hard to restore service as quickly as possible. Edit: Up and running again.
- grenran 1y agoseems like services are slowly recovering
- kitd 1y agoO ffs. I can't even access the NYT puzzles in the meantime ... Seriously disrupted, man
- littlecranky67 1y agoMaybe unrelated, but yesterday I went to pick up my package from an Amazon Locker in Germany, and the display said "Service unavailable". I'll wait until later today before I go and try again.
- HighGoldstein 1y agoI wonder why a package locker needs connectivity to give you a package. Since your package can't be withdrawn again from a different location, partitioning shouldn't be an issue.
- jcgl 1y agoGenerally speaking, it's easier to have computation (logic, state, etc.) centralized. If the designers didn't prioritize scenarios where decentralization helped, then centralization would've been the better option.
- bstsb 1y agoglad all my services are either Hetzner servers or EU region of AWS!
- gbalduzzi 1y agoTwilio is down worldwide: https://status.twilio.com/ https://status.twilio.com/
- croemer 1y agoCoinbase down as well
- croemer 1y agoCoinbase down as well: https://status.coinbase.com/ https://status.coinbase.com/
- littlecranky67 1y agoBest option for a whale to manipulate the price again.
- 00deadbeef 1y agoIt's not DNS There's no way it's DNS It was DNS
- foobar1962 1y agoThat or a Windows update.
- secondcoming 1y agoOr unattended-upgrades
- o1o1o1 1y agoI'm so happy we chose Hetzner instead but unfortunately we also use Supabase (dashboard affected) and Resend (dashboard and email sending affected). Probably makes sense to add "relies on AWS" to the criteria we're using to evaluate 3rd-party services.
- countWSS 1y agoReddit itself breaking down and errors appear. Does reddit itself depends on this?
- alvis 1y agoWhy would us-east-1 cause many UK banks and even UK gov web sites down too!? Shouldn't they operate in the UK region due to GDPR?
- Nextgrid 1y ago2 things: 1) GDPR is never enforced other than token fines based on technicalities. The vast majority of the cookie banners you see around are not compliant, so it the regulation was actually enforced they'd be the first to go... and it would be much easier to go after those (they are visible) rather than audit every company's internal codebases to check if they're sending data to a US-based provider. 2) you could technically build a service that relies on a US-based provider while not sending them any personal data or data that can be correlated with personal data.
- Ylpertnodi 1y ago>GDPR is never enforced Yes it is. >other than token fines based on technicalities. Result!
- Nextgrid 1y agoRead my post again. You can go to any website and see evidence of their non-compliance (you don't have to look very hard - they generally tend to push these in your face in the most obnoxious manner possible). You can't consider a regulation being enforced if everyone gets away with publishing evidence of their non-compliance on their website in a very obnoxious manner.
- GoblinSlayer 1y agoIntegration with USA for your safety :)
- assimpleaspossi 1y agoI'm thinking about that one guy who clicked on "OK" or hit return.
- rvitorper 1y agoSomebody, somewhere tried to rollback something and it failed
- starkindustries 1y agoIt has started recovering now. https://www.whatsmydns.net/#A/dynamodb.us-east-1.amazonaws.com https://www.whatsmydns.net/#A/dynamodb.us-east-1.amazonaws.c... is showing full recovery of dns resolutions.
- mk89 1y agoIt's fun to see SRE jumping left and right when they can do basically nothing at all. "Do we enable DR? Yes/No". That's all you can do. If you do, it's a whole machinery starting, which might take longer than the outage itself. They can't even use Slack to communicate - messages are being dropped/not sent. And then we laugh at the South Koreans for not having backed up their hard drives (which got burnt by actual fire, a statistically way less occurring event than an AWS outage). OK that's a huge screw up, but hey, this is not insignificant either. What will happen now? Nothing, like nothing happened after Crowdstrike's bug last year.
- XorNot 1y agoSignal seems to be dead too though, which is much more of a WTF?
- GoblinSlayer 1y agoA decentralized messenger is Tox.
- assimpleaspossi 1y agoAs of 4:26am Central Time in the USA, it's back up for one of my services.
- croemer 1y agoRelated thread: https://news.ycombinator.com/item?id=45640838 https://news.ycombinator.com/item?id=45640838
- croemer 1y agoRelated thread: https://news.ycombinator.com/item?id=45640772 https://news.ycombinator.com/item?id=45640772
- martinheidegger 1y agoDesigned to provide 99.999% durability and 99.999% availability Still designed, not implemented
- riknos314 1y agoThe real challenge is that implementations aren't static. Just because today's implementation has 4 9s that doesn't mean tomorrow's will...
- gritzko 1y agoidiocracy_window_view.jpg
- the_duke 1y agoSlack now also down: https://slack-status.com/ https://slack-status.com/
- raspasov 1y ago02:34 Pacific: Things seem to be recovering.
- jpfromlondon 1y agoThis will always be a risk when sharecropping.
- shakesbeard 1y agoSlack (canvas and huddles), Circle CI and Bitbucket are also reporting issues due to this.
- karel-3d 1y agoSlack is down. Is that related? Probably is.
- rob 1y agoMy Alexa is hit or miss at responding to queries right now at 5:30 AM EST. Was wondering why it wasn't answering when I woke up.
- comrade1234 1y agoI like that we can advertise to our customers that over the last X years we have better uptime than Amazon, google, etc.
- eska 1y agoJust yesterday I saw another Hetzner thread where someone claimed AWS beats them in uptime and someone else blasted AWS for huge incidents. I bet his coffee tastes better this morning.
- VBprogrammer 1y agoI honestly wonder if there is safety in the herd here. If you have a dedicated server in a rack somewhere that goes down and takes your site with it. Or even the whole data center has connectivity issues. As far as the customer is concerned, you screwed up. If you are on AWS and AWS goes down, that's covered in the news as a bunch of billion dollar companies were also down. Customer probably gives you a pass.
- BrentOzar 1y ago> If you are on AWS and AWS goes down, that's covered in the news as a bunch of billion dollar companies were also down. Customer probably gives you a pass. Exactly - I've had clients say, "We'll pay for hot standbys in the same region, but not in another region. If an entire AWS region goes down, it'll be in the news, and our customers will understand, because we won't be their only service provider that goes down, and our clients might even be down themselves."
- Aldipower 1y agoMy minor 2000 users web app hosted on Hetzner works fyi. :-P
- mlrtime 1y agoBut how are you going to web scale it!? /s
- Aldipower 1y agoWeb scale? It is an _web_ app, so it is already web scaled, hehe. Seriously, this thing runs already on 3 servers. A primary + backup and a secondary in another datacenter/provider at Netcup. DNS with another AnycastDNS provider called ClouDNS. Everything still way cheaper then AWS. The database is already replicated for reads. And I could switch to sharding if necessary. I can easily scale to 5, 7, whatever dedicated servers. But I do not have to right now. The primary is at 1% (sic!) load. There really is no magic behind this. And you have to write your application in a distributable way anyway, you need to understand the concepts of stateless, write-locking, etc. also with AWS.
- aembleton 1y agoRight up until the DNS fails
- Aldipower 1y agoI am using ClouDNS. That is an AnycastDNS provider. My hopes are that they are more reliable. But yeah, it is still DNS and it will fail. ;-)
- seeg 1y agoquay.io was down: https://status.redhat.com https://status.redhat.com
- dddfdfdfdfdf 1y agoHello world
- dddfdfdfdfdf 1y agohello world
- drcongo 1y agoSnow day!
- deleted 1y ago[deleted]
- throw-10-13 1y agothis is why you avoid us-east-1
- philipwhiuk 1y agoAnd yet you're still impacted because it hosts IAM
- danielpetrica 1y agoIn this moments I think devs should invest in vendor independence if they can. While I'm not to that stage yet (cloudlfare dependence) using open technologies like docker (or Kubernetes), Traefik instead of managed services can help in this disaster situations by switching to a different provider in a faster way than having to rebuild from zero. as a disclosure I'm not still to that point on my infrastructure But I'm trying to slowly define one for my self
- deleted 1y ago[deleted]
- pinkgolem 1y agoFyi: traffic also had problems bringing new devices online
- danielpetrica 1y agoI'm speacking of the self hosted version you can install on your own vps not the managwment version. I don't like using managed services if possible.
- shawn_w 1y agoOne of the radio stations I listen to is just dead air tonight. I assume this is the cause.
- myself248 1y agoA physical on-air broadcast station, not a web stream? That likely violates a license; they're required to perform station identification on a regular basis. Of course if they had on-site staff it wouldn't be an issue (worst case, just walk down to the transmitter hut and use the transmitter's aux input, which is there specifically for backup operations like this), but consolidation and enshittification of broadcast media mean there's probably nobody physically present.
- shawn_w 1y agoYeah, real over the air fm radio. This particular station is a Jack one owned by iHeart; they don't have DJs. Probably no techs or staff in the office overnight.
- JCM9 1y agoUS-East-1 is literally the Achilles Heel of the Internet.
- rvitorper 1y agoExactly
- sofixa 1y agoYou would think that after the previous big us-east-1 outages (to be fair there have been like 3 of them in the past decade, but still, that's plenty), companies would have started to move to other AWS regions and/or to spread workloads between them.
- JCM9 1y agoIt’s not that simple. The bit AWS doesn’t talk much about publicly (but will privately if you really push them) is that there’s core dependencies behind the scenes on us-east-1 for running AWS itself. When us-east-1 goes down the blast radius has often impacted things running in other regions. It impacts AWS internally too. For example rather ironically it looks like the outage took out AWS’s support systems so folks couldn’t contact support to get help. Unfortunately it’s not as simple as just deploying in multiple regions with some failover load balancing.
- shawabawa3 1y agoour eu-central-1 services had zero disruption during this incident, the only impact was that if we _had_ had any issues we couldn't log in to the AWS console to fix them So moving stuff out of us-east-1 absolutely does help
- JCM9 1y agoSure it helps. Folks just saying there’s lots of examples where you still get hit by the blast radius of a us-east-1 issue even if you’re using other regions.
- 1y ago
- BaudouinVH 1y agocanva.com was down until a few minutes ago.
- codebolt 1y agoAtlassian cloud is also having issues. Closing in on the 3 hour mark.
- ryanmcdonough 1y agoNow, I may well be naive - but isn't the point of these systems that you fail over gracefully to another data centre and no-one notices?
- spicybright 1y agoIt should be! When I was a complete newbie at AWS my first question was why do you have to pick a region, I thought the whole point was you didn't have to worry about that stuff
- nicce 1y agoAs far as I know, region selection is about regulation and privacy and guarantees on that.
- speedgoose 1y agoThe region labels found within the metadata are very very powerful. They make lawyers happy and they stop intelligence services to access the associated resources. For example, no one would even consider accessing data from a European region without the right paperwork.
- speed_spread 1y agoBecause if they were caught they'd have to pay _thousands_ of dollars in fines and get sternly talked to be high ranking officials.
- joncrane 1y agoIt's also about latency and nearness to users. Also some regions don't have all features so feature set also matters.
- Nifty3929 1y agoOne might hope that this, too, would be handled by the service. Send the traffic to the closest region, and then fallback to other regions as necessary. Basically, send the traffic to the closest region that can successfully serve it. But yeah, that's pretty hard and there are other reasons customers might want to explicitly choose the region.
- tosh 1y agoSES and signal seem to work again
- jacquesm 1y agoEvery week or so we interview a company and ask them if they have a fall-back plan in case AWS goes down or their cloud account disappears. They always have this deer-in-the-headlights look. 'That can't happen, right?' Now imagine for a bit that it will never come back up. See where that leads you. The internet got its main strengths from the fact that it was completely decentralized. We've been systematically eroding that strength.
- hvb2 1y ago> The internet got its main strengths from the fact that it was completely decentralized. Decentralized in terms of many companies making up the internet. Yes we've seen heavy consolidation in now having less than 10 companies make up the bulk of the internet. The problem here isn't caused by companies chosing one cloud provider over the other. It's the economies of scale leading us to few large companies in any sector.
- jacquesm 1y agoI think one reason is that people are just bad at statistics. Chance of materialization * impact = small. Sure. Over a short enough time that's true for any kind of risk. But companies tend to live for years, decades even and sometimes longer than that. If we're going to put all of those precious eggs in one basket, as long as the basket is substantially stronger than the eggs we're fine, right? Until the day someone drops the basket. And over a long enough time span all risks eventually materialize. So we're playing this game, and usually we come out ahead. But trust me, 10 seconds after this outage is solved everybody will have forgotten about the possibility.
- hvb2 1y agoAbsolutely, but the cost of perfection (100% uptime in this case) is infinite. As long as the outages are rare enough and you automatically fail over to a different region, what's the problem?
- 1y ago
- greatgib 1y agoWhen I follow the link, I arrive on a "You broke reddit" page :-o
- JCM9 1y agoHave a meeting today with our AWS account team about how we’re no longer going to be “All in on AWS” as we diversify workloads away. Was mostly about the pace of innovation on core services slowing and AWS being too far behind on AI services so we’re buying those from elsewhere. The AWS team keeps touting the rock solid reliability of AWS as a reason why we shouldn’t diversify our cloud. Should be a fun meeting!
- GoblinSlayer 1y agoBut then you will be affected by outages of every dependency you use.
- caymanjim 1y agoThis is the real problem. Even if you don't run anything in AWS directly, something you integrate with will. And when us-east-1 is down, it doesn't matter if those services are in other availability zones. AWS's own internal services rely heavily on us-east-1, and most third-party services live in us-east-1. It really is a single point of failure for the majority of the Internet.
- parliament32 1y ago> Even if you don't run anything in AWS directly, something you integrate with will. Why would a third-party be in your product's critical path? It's like the old business school thing about "don't build your business on the back of another"
- macintux 1y agoNo man is an island, entire of itself
- thinkindie 1y agoNot necessarily our critical path but today circleci was affected greatly which also affected our capacity to deploy. Luckily it was a Monday morning therefore we didn’t even have to deploy an hot fix.
- TrackerFF 1y agoLots of outage in Norway, started approximately 1 hour ago for me.
- nextaccountic 1y agoIs this why reddit is down? (https://www.redditstatus.com/ https://www.redditstatus.com/ still says it is up but with degraded infrastructure)
- krowek 1y agoShameless from them to make it look like it's a user problem. It was loading fine for me one hour ago, now I refresh the page and their message states I'm doing too many requests and should chill out (1 request per hour is too many for you?)
- anal_reactor 1y agoI remember that I made a website and then I got a report that it doesn't work on newest Safari. Obviously, Safari would crash with a message blaming the website. Bro, no website should ever make your shitty browser outright crash.
- balder1991 1y agoActually I’m just thinking that knowledge about how to crash Safari is valuable.
- anal_reactor 1y agoTrue. At the time though I was just focused on fixing the bug.
- etothet 1y agoNever ascribe to malice that which is adequately explained by incompetence. It’s likely that, like many organizations, this scenario isn’t something Reddit are well prepared for in terms of correct error messaging.
- kaptainscarlet 1y agoI got a rate limit error which didn't make sense since it was my first time opening reddit in hours.
- TrackerFF 1y agoLots of outage happening in Norway, too. So I'm guessing it is a global thing.
- deleted 1y ago[deleted]
- weberer 1y agoLlama-5-beelzebub has escaped containment. A special task force has been deployed to the Virginia data center to pacify it.
- sota_pop 1y agoAmazon devs to ClaudeCode: “That didn’t fix it. The service is down, pls fix. Make no mistakes. pls.”
- cmiles8 1y agoUS-East-1 and its consistent problems are literally the Achilles Heel of the Internet.
- Danborg 1y agor/aws not found There aren't any communities on Reddit with that name. Double-check the community name or start a new community.
- xodice 1y agoMajor us-east-1 outages happened in 2011, 2015, 2017, 2020, 2021, 2023, and now again. I understand that us-east-1, N. VA, was the first DC but for fucks sake they've had HOW LONG to finish AWS and make us-east-1 not be tied to keeping AWS up.
- hvb2 1y agoFirst, not all outages are created equal, so you cannot compare them like that. I believe the 2021 one was especially horrific because of it affecting their dns service (route53) and the outage made writes to that service impossible. This made fail overs not work etcetera so their prescribed multi region setups didn't work. But in the end, some things will have to synchronizes their writes somewhere, right? So for dns I could see how that ends up in a single region. AWS is bound by the same rules as everyone else in the end... The only thing they have going for them that they have a lot of money to make certain services resilient, but I'm not aware of a single system that's resilient to everything.
- xodice 1y agoIf AWS fully decentralized its control planes, they’d essentially be duplicating the cost structure of running multiple independent clouds and I understand that is why they don't however as long as AWS is reliant upon us-east-1 to function, they have not achieved what they claim to me. A single point of failure for IAM? Nah, no thanks. Every AWS “global” service be it IAM, STS, CloudFormation, CloudFront, Route 53, Organizations, they all have deep ties to control systems originally built only in us-east-1/n. va. That's poor design, after all these years. They've had time to fix this. Until AWS fully decouples the control plane from us-east-1, the entire platform has a global dependency. Even if your data plane is fine, you still rely on IAM and STS for authentication and maybe Route 53 for DNS or failover CloudFormation or ECS for orchestration... If any of those choke because us-east-1’s internal control systems are degraded, you’re fucked. That’s not true regional independence.
- hvb2 1y agoYou can only decentralized your control plane if you don't have conflicting requirements? Assuming you cannot alter requirements or SLAs, I could see how their technical solutions are limited. It's possible, just not without breaking their promises. At that point it's no longer a technical problem
- ivad 1y agoSeems to have taken down my router "smart wifi" login page, and there's no backup router-only login option! Brilliant work, linksys....
- ta988 1y agoHappened to lots of commercial routers too (free wifi with sign-in pages in stores for example) and that's way outside us-east-1
- kristopherleads 1y agoWas just on a Lufthansa and then United flight - both of which did not have WiFi. Was wondering if there was something going on at the infrastructure level.
- Alex-C137 1y agoUnfortunately that is also be par for the course
- rvba 1y agoWhat if they use the same router inside AWS and now they cannot login too?
- hmry 1y agoWiFi login portal (Icomera) on the train I'm on doesn't work either.
- shinycode 1y agoIt’s that period of the year when we discover AWS clients that don’t have fallback plans
- waste_monk 1y ago♫ It's the most blunderful time of the year There'll be much admin moaning And servers not glowing and the NOC crew in tears It's the most blunderful time of the year ♫
- okr 1y agoBtw. we had a forced EKS restart last week on thursday due to Kubernetes updates. And something was done with DNS there. We had problems with ndots. Caused some trouble here. Would not be surprised, if it is related, heh.
- 0x002A 1y agoAs Amazon moves from day-1 company as it claimed once, to be the sales company like Oracle focusing on raking money, expect more outages to come, and longer to be resolved. Amazon is burning and driving away the technical talent and knowledge knowing the vendor lock-in will keep bringing the sweet money. You will see more sales people hoovering around your c-suites and executives, while you will face even worse technical support, that seem not knowing what they are talking about, yet alone to fix the support issue you expect to be fixed easily. Mark my words, and if you are putting your eggs in one basket, that basket is now too complex and too interdependent, and the people who built and knew those intricacies are driven away with RTOs, move to hubs. Eventually those services; all others (and also aws services themselves) heavily dependent on, might be more fragile than the public knows.
- zht 1y agoDo you have data suggesting AWS outages are more frequent and/or take longer to resolve?
- 0x002A 1y agoThis is a prediction, not a historical pattern to be observed now. Only future data can verify if this prediction was correct or not.
- MisterSandman 1y agoAWS has existed for like 2 decades now, is that not enough evidence for you?
- 0x002A 1y agoAnd has substantially changed in management in the last couple years. Have you read my first post?
- hopelite 1y agoThat is why technical leaders’ role wouldn’t demand they not only gather data, but also report things like accurate operational, alternative, and scenario cost analysis; financial risks; vendor lock-in; etc. However, as may be apparent just from that small set, it is not exactly something technical people often feel comfortable with doing. It is why at least in some organizations you get the friction of a business type interfacing with technical people in varying ways, but also not really getting along because they don’t understand each other and often there are barriers of openness.
- hubertzhang 1y agoI cannot pull images from docker hub.
- chistev 1y agoWhat is HN hosted on?
- Bender 1y ago2 physical servers, one active and one standby running BSD at M5 internet hosting.
- kkfx 1y agoHonestly anyone do have outages, that's nothing extraordinary, what's wrong is the number of impacted services. We choose (at least almost choose) to ditch mainframes for clusters also for resilience. Now with cheap desktop iron labeled "stable enough to be a serious server" we have seen mainframes re-created sometimes with a cluster of VM on top of a single server, sometimes with cloud services. Ladies and Gentleman's it's about time to learn reshoring in the IT world as well. Owning nothing, renting all means extreme fragility.
- Ygg2 1y agoIronically enough I can't access Reddit due to no healthy upstream.
- me551ah 1y agoWe created a single point of failure on the Internet, so that companies could avoid single points of failure in their data centers.
- nonethewiser 1y agoChina is unaffected.
- Spivak 1y agoIt's actually kinda great. When AWS has issues it makes national news and that's all you need to put on your status page and everyone just nods in understanding. It's a weird kind of holiday in a way.
- deleted 1y ago[deleted]
- retired_account 1y agoLooks like they’re nearly done fixing it. > Oct 20 3:35 AM PDT > The underlying DNS issue has been fully mitigated, and most AWS Service operations are succeeding normally now. Some requests may be throttled while we work toward full resolution. Additionally, some services are continuing to work through a backlog of events such as Cloudtrail and Lambda. While most operations are recovered, requests to launch new EC2 instances (or services that launch EC2 instances such as ECS) in the US-EAST-1 Region are still experiencing increased error rates. We continue to work toward full resolution. If you are still experiencing an issue resolving the DynamoDB service endpoints in US-EAST-1, we recommend flushing your DNS caches. We will provide an update by 4:15 AM, or sooner if we have additional information to share.
- chibea 1y agoIt's a bit funny that they say "most service operations are succeeding normally now" when, in fact, you cannot yet launch or terminate new EC2 instance, which is basically the defining feature of the cloud...
- rswail 1y agoIn that region, other regions are able to launch EC2s and ECS/EKS without a problem.
- jamwil 1y agoIs that material to a conversation about service uptime of existing resources, though? Are there customers out there that are churning through the full lifecycle of ephemeral EC2 instances as part of their day-to-day?
- shawabawa3 1y agoany company of non trivial scale will surely launch ec2 nodes during the day one of the main points of cloud computing is scaling up and down frequently
- 1y ago
- pantulis 1y agoNow I know why the documents I was sending to my Kindle didn't go through.
- CTDOCodebases 1y agoI'm getting rate limit issues on Reddit so it could be related.
- testemailfordg2 1y agoSeems like we need more anti-trust cases on AWS or need to break it down, it is becoming too big. Services used in rest of the world get impacted by issues in one region.
- arielcostas 1y agoBut they aren't abusing their market power, are they? I mean, they are too big and should definitely be regulated but I don't think you can argue they are much of a monopoly when others, at the very least Google, Microsoft, Oracle, Cloudflare (depending on the specific services you want) and smaller providers can offer you the same service and many times with better pricing. Same way we need to regulate companies like Cloudflare essentially being a MITM for ~20% of internet websites, per their 2024 report.
- SergeAx 1y agoHow much longer are we going to tolerate this marketing bullshit about "Designed to provide 99.999999999% durability and 99.99% availability"?
- coleca 1y agoThat is for S3 not AWS as a whole, AWS has never claimed otherwise. AFAIK S3 has never broken the 11 9s of durability.
- sitzkrieg 1y agoit is very funny to me that us-east-1 going down nukes the internet. all those multiple region reliability best practices are for show
- geye1234 1y agoPotentially-ignoramus comment here, apologies in advance, but amazon.com itself appears to be fine right now. Perhaps slower to load pages, by about half a second. Are they not eating (much of) their own dog food?
- bananapub 1y agothose sentences aren't logically connected - the outage this post is about is mostly confined to `us-east-1`, anyone who wanted to build a reliable system on top of AWS would do so across multiple regions, including Amazon itself. `us-east-1` is unfortunately special in some ways but not in ways that should affect well-designed serving systems in other regions.
- julianozen 1y agoThey are 100% fully on AWS and nothing else. (I’m Ex-Amazon) It seems like the outage is only effecting one region so AWS is likely falling back to others. I’m sure parts of the site are down but the main sites are resilient
- rwky 1y agoI was getting 500 errors a few hours ago on amazon.com
- altairprime 1y agoKindle downloads and Amazon orders history were wholly unavailable, with rapid errors from the responsive website.
- amai 1y agoThe internet was once designed to survive a nuclear war. Nowadays it cannot even survive until tuesday.
- aiiizzz 1y agoSlack was acting slower than usual, but did not go down. Color me impressed.
- chibea 1y agoOne main problem that we observed was that big parts of their IAM / auth setup was overloaded / down which led to all kinds of cascading problems. It sounds as if Dynamo was reported to be a root cause, so is IAM dependent on dynamo internally? Of course, such a large control plane system has all kinds of complex dependency chains. Auth/IAM seems like such a potentially (global) SPOF that you'd like to reduce dependencies to an absolute minimum. On the other hand, it's also the place that needs really good scalability, consistency, etc. so you probably like to use the battle proof DB infrastructure you already have in place. Does that mean you will end up with a complex cyclic dependency that needs complex bootstrapping when it goes down? Or how is that handled?
- cowsandmilk 1y agoMany AWS customers have bad retry policies that will overload other systems as part of their retries. DynamoDB being down will cause them to overload IAM.
- joncrane 1y agoWhich is interesting because per their health dashboard, >We recommend customers continue to retry any failed requests.
- veltas 1y agoCan't exactly change existing widespread practice so they're ready for that kind of handling.
- otterley 1y agoThey should continue to retry but with exponential backoff and jitter. Not in a busy loop!
- bcrosby95 1y agoIf the reliability of your system depends upon the competence of your customers then it isn't very reliable.
- bdangubic 1y agoyou put your sh*t in us-east-1 you need to plan for this :)
- amai 1y agoWe are on Azure. But our CI/CD pipelines are failing, because Docker is on AWS.
- kuon 1y agoI realize that my basement servers have better uptime than AWS this year! I think most sysadmin don't plan for AWS outage. And economically it makes sense. But it makes me wonder, is sysadmin a lost art?
- fooqux 1y ago> But it makes me wonder, is sysadmin a lost art? I dunno, let me ask chatgpt. Hmmm, it said yes.
- tripplyons 1y agoChatGPT often says yes to both a question and its inverse. People like to hear yes more than no.
- ninininino 1y agoYou missed their point. They were making a joke about over-reliance on AI.
- tredre3 1y ago> But it makes me wonder, is sysadmin a lost art? Yes. 15-20 years ago when I was still working on network-adjacent stuff I witnessed the shift to the devops movement. To be clear, the fact that devops don't plan for AWS failures isn't an indication that they lack the sysadmin gene. Sysadmins will tell you very similar "X can never go down" or "not worth having a backup for service Y". But deep down devops are developers who just want to get their thing running, so they'll google/serveroverflow their way into production without any desire to learn the intricacies of the underlying system. So when something breaks, they're SOL. "Thankfully" nowadays containers and application hosting abstracts a lot of it back away. So today I'd be willing to say that devops are sufficient for small to medium companies (and dare I say more efficient?).
- Centigonal 1y ago> But deep down devops are developers who just want to get their thing running, so they'll google/serveroverflow their way into production without any desire to learn the intricacies of the underlying system. So when something breaks, they're SOL. Depends on the devops team. I have worked with so many devops engineers who came from network engineering, sysadmin, or SecOps backgrounds. They all bring a different perspective and set of priorities.
- throwaway984393 1y ago[dead]
- LightBug1 1y agoRemember when the "internet will just route around a network problem"? FFS ...
- nla 1y agoI still don't know why anyone would use AWS hosting.
- nemo44x 1y agoSomeone’s got a case of the Monday’s.
- skywhopper 1y agoThere are plenty of ways to address this risk. But the companies impacted would have to be willing to invest in the extra operational cost and complexity. They aren’t.
- amelius 1y agoMedium also.
- randomtoast 1y agoThing is us-east-1 the primary region for many services of AWS. DynamoDB is a very central offering used by many service. And the issue that has happend is very common[^1]. I think no matter how hard you try to avoid it, in the end there's always a massive dependency chain for modern digital infrastructure[^2]. [1]: https://itsfoss.community/uploads/default/optimized/2X/a/ad3e17ff0a5a13b9c6f48716e7d35f965ec8137d_2_750x750.jpeg https://itsfoss.community/uploads/default/optimized/2X/a/ad3... [2]: https://xkcd.com/2347/ https://xkcd.com/2347/
- karel-3d 1y agoSlack was down, so I thought I will send message to my coworkers on Signal. Signal was also down.
- alimbada 1y agoE-mail still exists...
- pjmlp 1y agoIt just goes to show the difference between best practices in cloud computing, and what everyone ends up doing in reality, including well known industry names.
- mrbungie 1y ago> It just goes to show the difference between best practices in cloud computing, and what everyone ends up doing in reality, Well, inter-region DR/HA is a expensive thing to ensure (whether on salaries, infra or both), specially when you are in AWS.
- deleted 1y ago[deleted]
- esafak 1y agoDoes AWS follow its own Well-Architected Framework!?
- spyspy 1y agoEh, the "best practices" that would've prevented this aren't trivial to implement and are definitely far beyond what most engineering teams are capable of, in my experience. It depends on your risk profile. When we had cloud outages at the freemium game company I worked at, we just shrugged and waited for the systems to come back online - nobody dying because they couldn't play a word puzzle. But I've also had management come down and ask what it would take to prevent issues like that from happening again, and then pretend they never asked once it was clear how much engineering effort it would take. I've yet to meet a product manager that would shred their entire roadmap for 6-18 months just to get at an extra 9 of reliability, but I also don't work in industries where that's super important.
- pjmlp 1y agoIndeed, yet one would expect AWS to lead by example, including all of those that are only using a single region.
- mcphage 1y agoIt shouldn’t, but it does. As a civilization, we’ve eliminated resilience wherever we could, because it’s more cost-effective. Resilience is expensive. So everything is resting on a giant pile of single point of failures. Maybe this is the event to get everyone off of piling everything onto us-east-1 and hoping for the best, but the last few outages didn’t, so I don’t expect this one to, either.
- Trasmatta 1y agoThe irony is that true resilience is very complex, and complexity can be a major source of outages in and of itself
- lanstin 1y agoI have enjoyed this paper on such dynamics: https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s11-bronson.pdf https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s... It is kind of the child of what used to be called Catastrophe Theory, which in low dimensions is essentially a classification of folding of manifolds. Now the systems are higher dinemsional and the advice more practical/heuristic.
- mschuster91 1y ago> Maybe this is the event to get everyone off of piling everything onto us-east-1 and hoping for the best, but the last few outages didn’t, so I don’t expect this one to, either. Doesn't help either. us-east-1 hosts the internal control plane of AWS and a bunch of stuff is only available in us-east-1 at all - most importantly, Cloudfront, AWS ACM for Cloudfront and parts of IAM. And the last is the one true big problem. When IAM has a sniffle, everything else collapses because literally everything else depends on IAM. If I were to guess IAM probably handles millions if not billions of requests a second because every action on every AWS service causes at least one request to IAM.
- TheNewsIsHere 1y agoThe last re:Invent presentation I saw from one of the principals working on IAM quoted 500 million requests per second. I expect that’s because IAM also underpins everything inside AWS, too.
- JCM9 1y agoUS-East-1 is more than just a normal region. It also provides the backbone for other services, including those in other regions. Thus simply being in another region doesn’t protect you from the consistent us-east-1 shenanigans. AWS doesn’t talk about that much publicly, but if you press them they will admit in private that there are some pretty nasty single points of failure in the design of AWS that can materialize if us-east-1 has an issue. Most people would say that means AWS isn’t truly multi-region in some areas. Not entirely clear yet if those single points of failure were at play here, but risk mitigation isn’t as simple as just “don’t use us-east-1” or “deploy in multiple regions with load balancing failover.”
- helsinkiandrew 1y ago>US-East-1 is more than just a normal region. It also provides the backbone for other services, including those in other regions I thought that if us-east-1 goes down you might not be able to administer (or bring up new services) in other zones, but if you have services running that can take over from us-east-1, you can maintain your app/website etc. I haven’t had to do this for several years but that was my experience a few years ago on an outage - obviously it depends on the services you’re using. You can’t start cloning things to other zones after us-east-1 is down - you’ve left it too late
- cmiles8 1y agoWell that sounds like exactly the sort of thing that shouldn’t happen when there’s an issue given the usual response is to spin things up elsewhere, especially on lower priority services where instant failover isn’t needed.
- sgarland 1y agoIt depends on the outage. There was one a year or two ago (I think? They run together) that impacted EC2 such that as long as you weren’t trying to scale, or issue any commands, your service would continue to operate. The EKS clusters at my job at the time kept chugging along, but had Karptenter tried to schedule more nodes, we’d have had a bad time.
- helsinkiandrew 1y ago> The incident underscores the risks associated with the heavy reliance on a few major cloud service providers. Perhaps for the internet as a whole, but for each individual service it underscores the risk of not hosting your service in multiple zones or having a backup
- runako 1y agoEven though us-east-1 is the region geographically closest to me, I always choose another region as default due to us-east-1 (seemingly) being more prone to these outages. Obviously, some services are only available in us-east-1, but many applications can gain some resiliency just by making a primary home in any other region.
- joncrane 1y agoWhat services are only available in us-east-1?
- tom1337 1y agoIAM control plane for example: > There is one IAM control plane for all commercial AWS Regions, which is located in the US East (N. Virginia) Region. The IAM system then propagates configuration changes to the IAM data planes in every enabled AWS Region. The IAM data plane is essentially a read-only replica of the IAM control plane configuration data. and I believe some global services (like certificate manager, etc.) also depend on the us-east-1 region https://docs.aws.amazon.com/IAM/latest/UserGuide/disaster-recovery-resiliency.html https://docs.aws.amazon.com/IAM/latest/UserGuide/disaster-re...
- tomchuk 1y agoIAM, Cloudfront, Route53, ACM, Billing...
- nijave 1y agoparts of S3 (although maybe that's better after that major outage years ago)
- runako 1y agoIn addition to those listed in sibling comments, new services often roll out in us-east-1 before being made available in other regions. I recently ran into an issue where some Bedrock functionality was available in us-east-1 but not one of the other US regions.
- 1y ago
- fsto 1y agoIronically, the HTTP request to this article timed out twice before a successful response.
- josefritzishere 1y ago[flagged]
- jmedefind 1y agoPretty sure this is satire and not even remotely true.
- josefritzishere 1y agoI beleived it because of a thread I read 3 months ago about non-specific Amazon layoffs but you are right. It's AI slop, and not accurate. https://www.reddit.com/r/programming/comments/1m6krap/its_really_time_tech_workers_start_talking_about/ https://www.reddit.com/r/programming/comments/1m6krap/its_re...
- aeon_ai 1y agoIt's not DNS There's no way it's DNS It was DNS
- ajross 1y agoIt's "DNS" because the problem is that at the very top of the abstraction hierarchy in any system is a bit of manual configuration. As it happens, that naturally maps to the bootstrapping process on hardware needing to know how to find the external services it needs, which is what "DNS" is for. So "DNS" ends up being the top level of manual configuration. But it's the inevitability of the manual process that's the issue here, not the technology. We're at a spot now where the rest of the system reliability is so good that the only things that bring it down are the spots where human beings make mistakes on the tiny handful of places where human operation is (inevitably!) required.
- allarm 1y ago> hardware needing to know how to find the external services it needs, which is what "DNS" is for. So "DNS" ends up being the top level of manual configuration. Unless DNS configuration propagates over DHCP?
- ajross 1y agoDHCP can only tell you who the local DNS server is. That's not what's failed, nor what needs human configuration. At the top of the stack someone needs to say "This is the cluster that controls boot storage", "This is the IP to ask for auth tokens", etc... You can automatically configure almost everything but there still has to be some way to get started.
- jamesbelchamber 1y agoThis website just seems to be an auto-generated list of "things" with a catchy title: > 5000 Reddit users reported a certain number of problems shortly after a specific time. > 400000 A certain number of reports were made in the UK alone in two hours.
- sinpor1 1y ago[flagged]
- ktosobcy 1y agoUhm... E(U)ropean sovereigny (and in general spreading the hosting as much as possbile) needed ASAP…
- BirAdam 1y agoWell, except for a lot of business leaders saying that they don't care if it's Amazon that goes down, because "the rest of the internet will be down too." Dumb argument imho, but that's how many of them think ime.
- tjwebbnorfolk 1y agobecause... EU clouds don't break? https://news.ycombinator.com/item?id=43749178 https://news.ycombinator.com/item?id=43749178
- Cthulhu_ 1y agoNah, because European services should not be affected by a failure in the US. Whatever systems they have running in us-east-1 should have failovers in all major regions. Today it's an outage in Virginia, tomorrow it could be an attack on undersea cables (which I'm confident are mined and ready to be severed at this point by multiple parties).
- protocolture 1y agoMined? My understanding is that they are maintained too regularly for that, or we would know. Also, lots of the bad guy boogeymen countries have legal and technical methods to do this without property damage. Just blackhole a bunch of routes.
- ktosobcy 1y agoNo, it's called "diversification". Applies both to stock/currency/investments as well as enything else :P
- tjwebbnorfolk 1y ago
- d_burfoot 1y agoI think AWS should use, and provide as an offering to big customers, a Chaos Monkey tool that randomly brings down specific services in specific AZs. Example: DynamoDB is down in us-east-1b. IAM is down in us-west-2a. Other AWS services should be able to survive this kind of interruption by rerouting requests to other AZs. Big company clients might also want to test against these kinds of scenarios.
- davidrupp 1y agoAWS Fault Injection Service: https://docs.aws.amazon.com/fis/latest/userguide/what-is.html https://docs.aws.amazon.com/fis/latest/userguide/what-is.htm...
- jrochkind1 1y agoAt some point AWS has so many services it's subject to a version of xkcd Rule 34 -- if you can imagine it, there's an AWS service for it.
- davidrupp 1y agoI used to tell people there that my favorite development technique was to sit down and think about the system I wanted to build, then wait for it to be announced at that year's re:Invent. I called it "re:Invent and Simplify". "I" built my best stuff that way.
- bob1029 1y agoOne thing has become quite clear to me over the years. Much of the thinking around uptime of information systems has become hyperbolic and self-serving. There are very few businesses that genuinely cannot handle an outage like this. The only examples I've personally experienced are payment processing and semiconductor manufacturing. A severe IT outage in either of these businesses is an actual crisis. Contrast with the South Korean government who seems largely unaffected by the recent loss of an entire building full of machines with no backups. I've worked in a retail store that had a total electricity outage and saw virtually no reduction in sales numbers for the day. I have seen a bank operate with a broken core system for weeks. I have never heard of someone actually cancelling a subscription over a transient outage in YouTube, Spotify, Netflix, Steam, etc. The takeaway I always have from these events is that you should engineer your business to be resilient to the real tradeoff that AWS offers. If you don't overreact to the occasional outage and have reasonable measures to work around for a day or 2, it's almost certainly easier and cheaper than building a multi cloud complexity hellscape or dragging it all back on prem. Thinking in terms of competition and game theory, you'll probably win even if your competitor has a perfect failover strategy. The cost of maintaining a flawless eject button for an entire cloud is like an anvil around your neck. Every IT decision has to be filtered through this axis. When you can just slap another EC2 on the pile, you can run laps around your peers.
- vidarh 1y ago> The takeaway I always have from these events is that you should engineer your business to be resilient An enduring image that stays with me was when I was a child and the local supermarket lost electricity. Within seconds the people working the tills had pulled out hand cranks by which the tills could be operated. I'm getting old, but this was the 1980's, not the 1800's. In other words, to agree with your point about resilience: A lot of the time even some really janky fallbacks will be enough. But to somewhat disagree with your apparent support for AWS: While it is true this attitude means you can deal with AWS falling over now and again, it also strips away one of the main reasons people tend to give me for why they're in AWS in the first place - namely a belief in buying peace of mind and less devops complexity (a belief I'd argue is pure fiction, but that's a separate issue). If you accept that you in fact can survive just fine without absurd levels of uptime, you also gain a lot more flexibility in which options are viable to you. The cost of maintaining a flawless eject button is indeed high, but so is the cost of picking a provider based on the notion that you don't need one if you're with them out of a misplaced belief in the availability they can provide, rather than based on how cost effectively they can deliver what you actually need.
- bilekas 1y agoThese things happen when profits are the measure everything. Change your provider, but if their number doesn't go up, they wont be reliable. So your complaints matter nothing because "number go up". I remember the good old days of everyone starting a hosting company. We never should have left.
- epistasis 1y agoIt certainly doesn't take profit for bad planning to happen... And everybody starting a hosting company is definitely a profit driven activity.
- bilekas 1y ago> And everybody starting a hosting company is definitely a profit driven activity. Absolutely, nobody was doing it out of charity, but there is more diversity in the market and thus more innovation and then the market decides. Right now we have 3 major providers, and that makes up the lion's share. That's consolidation of a service. I believe that's not good for the market or the internet as a whole.
- DanHulton 1y agoI forget where I read it originally, but I strongly feel that AWS should offer a `us-chaos-1` region, where every 3-4 days, one or two services blow up. Host your staging stack there and you build real resiliency over time. (The counter joke is, of course, "but that's `us-east-1` already! But I mean deliberately and frequently.)
- time0ut 1y agoInteresting day. I've been on an incident bridge since 3AM. Our systems have mostly recovered now with a few back office stragglers fighting for compute. The biggest miss on our side is that, although we designed a multi-region capable application, we could not run the failover process because our security org migrated us to Identity Center and only put it in us-east-1, hard locking the entire company out of the AWS control plane. By the time we'd gotten the root credentials out of the vault, things were coming back up. Good reminder that you are only as strong as your weakest link.
- shawabawa3 1y agofor what it's worth, we were unable to login with root credentials anyway i don't think any method of auth was working for accessing the AWS console
- kondro 1y agoSure it was, you just needed to login to the console via a different regional endpoint. No problems accessing systems from ap-southeast-2 for us during this entire event, just couldn’t access the management planes that are hosted exclusively in us-east-1.
- nijave 1y agoLike the other poster said, you need to use a different region. The default region (of course) sends you to us-east-1 e.x. https://us-east-2.console.aws.amazon.com/console/home https://us-east-2.console.aws.amazon.com/console/home
- 1970-01-01 1y agoI remember Facebook had a similar story when they botched their BGP update and couldn't even access the vault. If you have circular auth, you don't have anything when somebody breaks DNS.
- crote 1y agoWasn't there an issue where they required physical access to the data center to fix the network, which meant having to tap in with a keycard to get in, which didn't work because the keycard server was down, due to the network being down?
- mrbluecoat 1y ago> due to an "operational issue" related to DNS Always DNS..
- rose-knuckle17 1y agoaws had an outage. Many companies were impacted. Headlines around the world blame AWS. the real news is how easy it is to identify companies that have put cost management ahead of service resiliency. Lots of orgs operating wholly in AWS and sometimes only within us-east-1 had no operational problems last night. Some that is design (not using the impacted services). Some of that is good resiliency in design. And some of that was dumb luck (accidentally good design). Overall, those companies that had operational problems likely wouldn't have invested in resiliancy expenses in any other deployment strategy either. It could have happened to them in Azure, GCP or even a home rolled datacenter.
- aurareturn 1y agoRedundancy is insanely expensive especially for SaaS companies where the biggest cost is cloud. Are customers willing to pay companies for that redundancy? I think not. Once every few years outage for 3 hours is fine for non critical services.
- skopje 1y ago>> Redundancy is insanely expensive especially for SaaS companies That right there means the business model is fucked to begin with. If you can't have a resilient service, then you should not be offering that service. Period. Solution: we were fine before the cloud, just a little slower. No problem going back to that for some things. Not everything has to be just in time at lowest possible cost.
- AtNightWeCode 1y agoIn general it is not expensive. In most cases you can either load balance across two regions all the time or have a fallback region that you scale out/up and switch to if needed.
- cheeze 1y agoQuite expensive to build though. Many of these companies don't have the sharpest engineers building multi-cloud. IMO, going multi AZ or multi-cloud adds a good amount of complexity. TBH I don't care if last.fm doesn't work for 8 hours a year, that isn't a big deal. My bank? Yeah that should work.
- dangoodmanUT 1y agoReminder that AZs don't go down Entire regions go down Don't pay for intra-az traffic friends
- avi_vallarapu 1y agoThis is the reason why it is important to plan Disaster recovery and also plan Multi-Cloud architectures. Our applications and databases must have ultra high availability. It can be achieved with applications and data platforms hosted on different regions for failover. Critical businesses should also plan for replication across multiple cloud platforms. You may use some of the existing solutions out there that can help with such implementations for data platforms. - Qlik replicate - HexaRocket and some more. Or rather implement native replication solutions available with data platforms.
- zoklet-enjoyer 1y agoIs this why Wordle logged me out and my 2 guesses don't seem to have been recorded? I am worried about losing my streak.
- mentalgear 1y ago> Amazon Alexa: routines like pre-set alarms were not functioning. It's ridiculous how everything is being stored in the cloud, even simple timers. It's past high time to move functionality back on-device, which would come with the advantage of making it easier to de-connect from big tech's capitalist surveillance state as well.
- Itanu 1y agoexactly why it won't happen :)
- freedomben 1y agoExactly. I half-seriously like to say things like, "I'm excited for a time when we have powerful enough computers to actually run applications on them instead of being limited to only thin clients." Only problem is most of the younger people don't get the reference anymore, so it's mainly the olds that get it
- ibejoeb 1y agoThis is just a silly anecdote, but every time a cloud provider blips, I'm reminded. The worst architecture I've ever encountered was a system that was distributed across AWS, Azure, and GCP. Whenever any one of them had a problem, the system went down. It also cost 3x more than it should.
- manishsharan 1y agoYou mean multi-cloud strategy ! You wanna know how you got here ? See the sales team from Google flew out an executive to NBA Finals, Azure Sales team flew out another executive to NFL superBowl and the AWS team flew out yet another executive to Wimbledon finals. And thats how you end up with multi-cloud strategy.
- kevstev 1y agoEh, businesses want to stay resilient to a single vendor going down. My least favorite question in interviews this past year was around multi-cloud. Because imho it just isn't worth it- the increased complexity, the trying to like-like services across different clouds that aren't always really the same, and then just the ongoing costs of chaos monkeying and testing that this all actually works, especially in the face of a partial outage like this vs something "easy" like a complete loss of network connectivity... but that is almost certainly not what CEOs want to hear (mostly who I am dealing with here going for VPE or CTO level jobs). I could care less about having more vendor dinners when I know I am promising a falsehood that is extremely expensive and likely going to cost me my job or my credibility at some point.
- pluto_modadic 1y agosticker shock / looking at alternative vendors
- ibejoeb 1y agoIn this particular case, it was resume-oriented architecture (ROAr!) The original team really wanted to use all the hottest new tech. The management was actually rather unhappy, so the job was to pare that down to something more reliable.
- bgwalter 1y agoProbably related: https://www.nytimes.com/2025/05/25/business/amazon-ai-coders.html https://www.nytimes.com/2025/05/25/business/amazon-ai-coders... "Pushed to use artificial intelligence, software developers at the e-commerce giant say they must work faster and have less time to think." Every bit of thinking time spent on a dysfunctional, lying "AI" agent could be spent on understanding the system. Even if you don't move your mouse all the time in order to please a dumb middle manager.
- add-sub-mul-div 1y agoKeep going
- renegade-otter 1y agoIf we see more of this, it would not be crazy to assume that all this compelling of engineers to "use AI" and the flood of Looks Good To Me code is coming home.
- Cthulhu_ 1y agoBig if, major outages like this aren't unheard of, and so far, fairly uncommon. Definitely hit harder than their SLAs promise though. I hope they do an honest postmortem, but I doubt they would blame AI even if it was somehow involved. Not to mention you can't blame AI unless you go completely hands-off - but that's like blaming an outsourcing partner, which also never happens.
- vivzkestrel 1y agostupid question: is buying a server rack and running it at home subject to more downtimes in a year than this? has anyone done an actual SLA analysis?
- alphabettsy 1y agoThat depends on a lot of factors, but for me personally, yes it is. Much worse. Assuming we’re talking about hosting things for Internet users. My fiber internet connection has gone down multiple times, though relatively quickly restored. My power has gone out several times in the last year, with one storm having it out for nearly 24 hrs. I was sleep when it went out and I didn’t start the generator until it was out for 3-4 hours already, far longer than my UPSes could hold up. I’ve had to do maintenance and updates both physical and software. All of those things contribute to a downtime significantly higher than I see with my stuff running on Linode, Fly.io or AWS. I run Proxmox and K3s at home and it makes things far more reliable, but it’s also extra overhead for me to maintain. Most or all of those things could be mitigated at home, but at what cost?
- amelius 1y agoMaybe if you use a UPS and Starlink then ...
- dboreham 1y agoUnanswerable question. Better to perform a failure mode analysis. That rack in your basement would need redundant power (two power companies or one power company and a diesel generator which typically won't be legal to have at your home), then redundant internet service (actually redundant - not the cable company vs the phone company that underneath use the same backhaul fiber).
- pluto_modadic 1y agoso, funny story, my fiber got cut (backhoe) and it took then 12 hours to restore it. If you had /two/ houses, in separate towns, you'd have better luck. Or, if you had cell as a backup. Or: if you don't care about it being down for 12 hours.
- moralestapia 1y agoCurious to know how much does an outage like this cost to others. Lost data, revenue, etc. I'm not talking about AWS but whoever's downstream. Is it like 100M, like 1B?
- nivekney 1y agoWait a second, Snapchat impacted AGAIN? It was impacted during the last GCP outage.
- menomatter 1y agoWhat are the design best practices and industry standards for building on-premise fallback capabilities for critical infrastructure? Say for health care/banking ..etc
- skopje 1y agoA relative of mine lived and worked in the US for Oppenheimer Funds in the 1990's and they had their own datacenters all over the US, multiple redundancy for weather or war. But every millionaire feels entitled to be a billionaire now, so all of that cost was rolled into a single point of cloud failure. Reminds me of a great Onion tagline: "Plowshare hastily beaten back into sword."
- saltyoldman 1y agoGCP is underrated
- thinkindie 1y agoToday’s reminder: multi-region is so hard even AWS can’t get it right.
- draxil 1y agoI was just about to post that it didn't affect us (heavy AWS users, in eu-west-1). Buut, I stopped myself because that was just massively tempting fate :)
- mannyv 1y agoThis is why we use us-east-2.
- muttantt 1y agous-east-2 remains the best kept secret
- jdlyga 1y agoTime to start calling BS on the 9's of reliability
- esskay 1y agoEr...They appear to have just gone down again.
- lexandstuff 1y agoYep. I don't think they ever fully recovered, but status page is still reporting a lot of issues.
- jrochkind1 1y agoMy systems didn't actually seem to be affected until what I think was probably a SECOND spike of outages at about the time you posted. The retrospective will be very interesting reading! (Obviously the category of outages caused by many restored systems "thundering" at once to get back up is known, so that'd be my guess, but the details are always good reading either way).
- indoordin0saur 1y agoMine are more messed up now (12:30 ET) than they were this morning. AWS is lying that they've fixed the issue.
- 1970-01-01 1y agoSomeone, somewhere, had to report that doorbells went down because the very big cloud did not stay up. I think we're doing the 21st century wrong.
- Johnny555 1y agoMy Ring doorbell works just fine without an internet connection (or during a cloud outage). The video storage and app notifications are another matter, but the doorbell itself continues to ring when someone pushes the button.
- lapetitejort 1y agoSomeone, somewhere, had to report that rock throwers went down because the very big cloud did not stay up. I think we're doing the 16th century wrong.
- wartywhoa23 1y agoExcept we're not doing the 16th century right now.
- lapetitejort 1y agoBlack powder is still used by firearm enthusiasts, and just like in the 16th century I'm sure they don't appreciate it getting wet when it rains
- nokeya 1y agoServerless is down because servers are down. What an irony.
- thebruce87m 1y agoServerless is just someone elses server, or something
- bicepjai 1y agoIs this the outage that took Medium down ?
- mmmlinux 1y agoOhno, not Fortnite! oh, the humanity.
- TriangleEdge 1y agoSeverity - Degraded... https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status https://downdetector.com/ https://downdetector.com/
- mlhpdx 1y agoCool, building in resilience seems to have worked. Our static site has origins in multiple regions via CloudFront and didn’t seem to be impacted (not sure if it would have been anyway). My control plane is native multi-region, so while it depends on many impacted services it stayed available. Each region runs in isolation. There is data replication at play but failing to replicate to us-east-1 had no impact on other regions. The service itself is also native multi-region and has multiple layers where failover happens (DNS, routing, destination selection). Nothing’s perfect and there are many ways this setup could fail. It’s just cool that it worked this time - great to see. Nothing I’ve done is rocket science or expensive, but it does require doing things differently. Happy to answer questions about it.
- SteveNuts 1y ago> Our static site has origins in multiple regions via CloudFront and didn’t seem to be impacted This seems like such a low bar for 2025, but here we are.
- immibis 1y agoYou're also betting that CloudFront isn't one of the several AWS services that only works when us-east-1 is up.
- mlhpdx 1y agoYeah, it's not clear how resilient CloudFront is but it seems good. Since content is copied to the points of presence and cached it's the lightly used stuff that can break (we don't do writes through CloudFront, which in IMHO is an anti-pattern). We setup multiple "origins" for the content so hopefully that provides some resiliency -- not sure if it contributed positively in this case since CF is such a black box. I might setup some metadata for the different origins so we can tell which is in use.
- x3n0ph3n3 1y agoCloudFront isn't just for CDN, but also for DDoS protection. Writes through CloudFront are not an anti-pattern.
- ta1243 1y agoPaying for resilience is expensive. not as expensive as AWS, but it's not free. Modern companies live life on the edge. Just in time, no resilience, no flexibility. We see the disaster this causes whenever something unexpected happens - the Evergiven blocking Suez for example, let alone something like Covid However increasingly what should be minor loss of resilience, like an AWS outage or a Crowdstrike incident, turns into major failures. This fragility is something government needs to legislate to prevent. When one supermarket is out that's fine - people can go elsewhere, the damage is contained. When all fail, that's a major problem. On top of that, the attitude that the entire sector has is also bad. People thing IT should tail once or twice a year and it's not a problem. If that attitude affect truly important systems it will lead to major civil projects. Any civilitsation is 3 good meals away from anarchy. There's no profit motive to avoid this, companies don't care about being offline for the day, as long as all their mates are also offline.
- webdoodle 1y agoI in-housed an EMR for a local clinic because of latency and other network issues taking the system offline several times a month (usually at least once a week). We had zero downtime the whole first year after bringing it all in house, and I got employee of the month for several months in a row.
- fogzen 1y agoGreat. Hope they’re down for a few more days and we can get some time off.
- twistedpair 1y agoWow, about 9 hours later and 21 of 24 Atlassian services are still showing up as impacted on their status page. Even @ 9:30am ET this morning, after this supposedly was clearing up, my doctor's office's practice management software was still hosed. Quite the long tail here. https://status.atlassian.com/ https://status.atlassian.com/
- BiraIgnacio 1y agoIt's scary to think about how much power and perhaps influence the AWS platform has. (albeit it shouldn't be surprising)
- megous 1y agoI didn't even notice anything was wrong today. :) Looks like we're well disconnected from the US internet infra quasi-hegemony.
- indoordin0saur 1y agoSeems like major issues are still ongoing. If anything it seems worse than it did ~4 hours ago. For reference I'm a data engineer and it's Redshift and Airflow (AWS managed) that is FUBAR for me.
- wavemode 1y agoSEV-0 for my company this morning. We can't connect to RDS anymore.
- assholesRppl2 1y agoYep, confirmed worse - DynamoDB now returning "ServiceUnavailableException"
- perching_aix 1y agoIn addition to those, Sagemaker also fails for me with an internal auth error specifically in Virginia. Fun times. Hope they recover by tomorrow.
- autophagian 1y agoYeah. We had a brief window where everything resolved and worked and now we're running into really mysterious flakey networking issues where pods in our EKS clusters timeout talking to the k8s API.
- cj00 1y agoYeah, networking issues cleared up for a few hours but now seem to be as bad as before.
- ivape 1y agoThere are entire apps like Reddit that are still not working. What the fuck is going on?
- htrp 1y agothundering herd problems.... every time they say they fix it something else breaks
- altbdoor 1y agoHad a meeting where developers were discussing the infrastructure for an application. A crucial part of the whole flow was completely dependant on an AWS service. I asked if it was a single point of failure. The whole room laughed, I rest my case.
- Elidrake24 1y agoIf you were dependent upon a single distribution (region) of that Service, yes it would be a massive single point of failure in this case. If you weren't dependent upon a particular region, you'd be fine.
- zimbu668 1y agoOf course lots of AWS services have hidden dependencies on us-east-1. During a previous outage we needed to update a Route53(DNS) record in us-west-2, but couldn't because of the outage in us-east-1.
- ineedasername 1y agoSo, AWS's redundant availability goes something like "Don't worry, if nothing is working in us-east-1, it will trigger failover to another regions" ... "Okay, where's that trigger located?" ... "In the us-east-1 region also" ... "Doens't that seem a problem to you?" ... "You'd think it might be! But our logs say it's never been used."
- wubrr 1y agoSome 'regional' AWS services still rely on other services (some internal) that are only in us-east-1.
- antinomicus 1y agoEven Amazon’s own services (ie ring) were affected by this outage
- ta1243 1y agoRelying on AWS is a single point of failure. Not as much as relying on a single AWS region, but it's still a single point. It's fairly difficult to avoid single points of failure completely, and if you do it's likely your suppliers and customers haven't managed to. It's about how much your risk level is. AWS us-east-1 fails constantly, it has terrible uptime, and you should expect it to go. A cyberattack which destroyed AWSs entire infrastructure would be less likely. BGP hijacks across multiple AWS nodes are quite plausible though, but that can be mitigated to an extent with direct connects. Sadly it seems people in charge of critical infrastructure don't even bother thinking about these things, because next quarters numbers are more important. I can avoid London as a single point of failure, but the loss of Docklands would cause so much damage to the UK's infrastructure I can't confidently predict that my servers in Manchester connected to peering points such as IXman will be able to reach my customer in Norwich. I'm not even sure how much international connectivity I could rely on. In theory Starlink will continue to work, but in practice I'm not confident. When we had power issues in Washington DC a couple of months ago, three of our four independent ISPS failed, as they all had undeclared requirements on active equipment in the area. That wasn't even a major outage, just a local substation failure. The one circuit which survived was clearly just fibre from our (UPS/generator backed) equipment room to a data centre towards Baltimore (not Ashburn).
- bigbuppo 1y agoWhose idea was it to make the whole world dependent on us-east-1?
- nemomarx 1y agoThe NSA might be happy everything runs through a local data center to their Virginia offices
- kondro 1y agoThe most recent public count of datacenters for AWS in us-east-1 is 159. I suspect that’s even an unwieldy number for NSA to spy on.
- dijit 1y ago11 years ago: https://news.ycombinator.com/item?id=8448894 https://news.ycombinator.com/item?id=8448894 US$70 billion in spend aggregating data, back then, this number has only increased. https://journals.sagepub.com/doi/pdf/10.1177/2053951714541861 https://journals.sagepub.com/doi/pdf/10.1177/205395171454186...
- lenerdenator 1y ago1) People in us-east-1. 2) People who thought that just having stuff "in the cloud" meant that it was automatically spread across regions. Hint, it's not; you have to deploy it in different regions and architect/maintain around that. 3) Accounting.
- wting 1y agoEh, us-east-1 is the oldest AWS region and if you get some AWS old timers talking over some beers they'll point out the legacy SPOFs that still exist in us-east-1.
- rickmode 1y agoIsn’t it the cheapest AWS region? Or at least among the cheapest. If I’m correct, this incentivizes users to start there.
- redwood 1y agoSurprising and sad to see how many folks are using DynamoDB There are more full featured multi-cloud options that don't lock you in and that don't have the single point of failure problems. And they give you a much better developer experience... Sigh
- pardner 1y agoDarn, on Heroku even the "maintenance mode" (redirects all routes to a static url) won't kick in.
- deleted 1y ago[deleted]
- JCM9 1y agoAt 3:03 AM PT AWS posted that things are recovering and sounded like issue was resolved. Then things got worse. At 9:13 AM PT it sounds like they’re back to troubleshooting. Honestly sounds like AWS doesn’t even really know what’s going on. Not good.
- vishnugupta 1y agoThis is exacerbated by the fact that this is Diwali week which means the most of Indian engineers will be out on leave. Tough luck.
- dingnuts 1y ago[flagged]
- jeffrallen 1y agoRight, like vacation?
- jedimastert 1y agoI'm unaware of any organization that doesn't give the same number of vacation days based on religion and region- or company-wide holidays...
- Waterluvian 1y agoI know there's a lot of anecdotal evidence and some fairly clear explanations for why `us-east-1` can be less reliable. But are there any empirical studies that demonstrate this? Like if I wanted to back up this assumption/claim with data, is there a good link for that, showing that us-east-1 is down a lot more often?
- everfrustrated 1y agoThe unreliability claim is driven by two factors. 1. When aws deploys changes they run through a pipeline which pushes change to regions one at a time. Most services start with us-east-1 first. 2. us-east-1 is MASSIVE and considerably larger than the next largest region. There's no public numbers but I wouldn't be surprised if it was 50% of their global capacity. An outage in any other region never hits the news.
- iLoveOncall 1y ago> 1. When aws deploys changes they run through a pipeline which pushes change to regions one at a time. This is true. > Most services start with us-east-1 first. This is absolutely false. Almost every service will FINISH with the largest and most impactful regions.
- captainkrtek 1y agoAgreed. Most services start deployments on a small number of hosts in single AZs in small less-known regions, ramping up from there. In all my years there I don’t recall “us-east-1 first”.
- riknos314 1y agoEach AWS service may choose different pipeline ordering based on the risks specific to their architecture. In general: You don't deploy to the largest region first because of the large blast radius. You may not want to deploy to the largest region last because then if there's an issue that only shows up at that scale you may need to roll every single region back (divergent code across regions is generally avoided as much as possible). A middle ground is to deploy to the largest region second or third.
- Liftyee 1y agoDamn. This is why Duolingo isn't working properly right now.
- redeux 1y agoIt’s a good day to be a DR software company or consultant
- sineausr931 1y agoOn a bright note, Alexa has stopped pushing me merchandise.
- neuroelectron 1y agoSounds like a circular error with monitoring is flooding their network with metrics and logs, causing DNS to fail and produce more errors, flooding the network. Likely root cause is something like DNS conflicts or hosts being recreated on the network. Generally this is a small amount of network traffic but the LBs are dealing with host address flux, causing the hosts to keep colliding host addresses as they attempt to resolve to a new host address which are being lost from dropped packets and with so many hosts in one AZ, there's a good chance they end up with a new conflicting address.
- czhu12 1y agoOur entire data stack (Databricks and Omni) are all down for us also. The nice thing is that AWS is so big and widespread that our customers are much more understanding about outages, given that its showing up on the news.
- busymom0 1y agoFor me Reddit is down and also the amazon home page isn't showing any items for me.
- aaronbrethorst 1y agoMy ISP's DNS servers were inaccessible this morning. Cloudflare and Google's DNS servers have all been working fine, though: 1.1.1.1, 1.0.0.1, and 8.8.8.8
- Readerium 1y ago99.999 percent lol
- motbus3 1y agoAlways a lovely Monday when you wake just in time to see everything going down
- suralind 1y agoI wonder how their nines are going. Guess they'll have to stay pretty stable for the next 100 years.
- itqwertz 1y agoDid they try asking Claude to fix these issues? If it turns out this problem is AI-related, I'd love to see the AAR.
- ecommerceguy 1y agoJust tried to get into Seller Central, returned a 504.
- twistedpair 1y agoI just saw services that were up since 545AM ET go down around 12:30PM ET. Seems AWS has broken Lambda again in their efforts to fix things.
- wcchandler 1y agoThis is usually something I see on Reddit first, within minutes. I’ve barely seen anything on my front page. While I understand it’s likely the subs I’m subscribed to, that was my only reason for using Reddit. I’ve noticed that for the past year - more and more tech heavy news events don’t bubble up as quickly anymore. I also didn’t see this post for a while for whatever reason. And Digg was hit and miss on availability for me, and I’m just now seeing it load with an item around this. I think I might be ready to build out a replacement through vibe coding. I don’t like being dependent on user submissions though. I feel like that’s a challenge on its own.
- kccqzy 1y agoReddit itself is having issues. I have multiple comments fail to post. And the Reddit user page leads me to a 404.
- qingcharles 1y agoYeah, Reddit has been half-working all morning. Last time this happened I had an account get permabanned because the JavaScript on the page got stuck in a no-backoff retry loop and it banned me for "spamming." Just now it put me in rate limit jail for trying to open my profile one time, so I've closed out all my tabs.
- postexitus 1y agoMost Reddit API is down as well.
- ryanisnan 1y agoAnecdotally, I think you should disregard this. I found out about this issue first via Reddit, roughly 30 minutes after the onset (we had an alarm about control plane connectivity).
- midtake 1y agoReddit is worthless now, and posting about your tech infrastructure on reddit is a security and opsec lapse. My workplace has reddit blocked at the edge. I would trust X more than reddit, and that is with X having active honeypot accounts (it is even a meme about Asian girls). In fact, heard about this outage on X before anywhere else.
- deleted 1y ago[deleted]
- wartywhoa23 1y agoSomeone vibecoded it down.
- stego-tech 1y agoNot remotely surprised. Any competent engineer knows full well the risk of deploying into us-east-1 (or any “default” region for that matter), as well as the risks of relying on global services whose management or interaction layer only exists in said zone. Unfortunately, us-east-1 is the location most outsourcing firms throw stuff, because they don’t have to support it when it goes pear-shaped (that’s the client’s problem, not theirs). My refusal to hoard every asset into AWS (let alone put anything of import in us-east-1) has saved me repeatedly in the past. Diversity is the foundation of resiliency, after all.
- mcintyre1994 1y ago> as well as the risks of relying on global services whose management or interaction layer only exists in said zone. Is this well known/documented? I don't have anything on AWS but previously worked for a company that used it fairly heavily. We had everything in EU regions and I never saw any indication/warning that we had a dependency on us-east-1. But I assume we probably did based on the blast radius of today's outage.
- captainkrtek 1y agoSome of the “global” and “edge” services depend on us-east-1. See: https://docs.aws.amazon.com/whitepapers/latest/aws-fault-isolation-boundaries/global-services.html https://docs.aws.amazon.com/whitepapers/latest/aws-fault-iso... “In the aws partition, the IAM service’s control plane is in the us-east-1 Region, with isolated data planes in each Region of the partition.“ Also, intra-region, many of the services use eachother, and not in a manner where the public can discern the dependency map.
- toephu2 1y agoHalf the internet goes down because part of AWS goes down... what happened to companies having redundant systems and not having a single point of failure?
- jppope 1y agoIronically for most companies its cheaper to just say if AWS goes down half of the internet goes down so people will understand
- rdm_blackhole 1y agoMy app deployed on Vercel and therefore indirectly deployed on us-east-1 was down for about 2 hours today then came back up and then went down again 10 minutes ago for 2 or 3 minutes. It seems like they are still intermittent issues happening.
- lawlessone 1y agoAm i imagining it or are more things like this happening in recent weeks than usual?
- Isuckatcode 1y agoMan , I just wanted to enjoy celebrating Diwali with my family but been up from 3am trying to recover our services. There goes some quality time
- melozo 1y agoEven internal Amazon tooling is impacted greatly - including the internal ticketing platform which is making collaboration impossible during the outage. Amazon is incapable of building multi-region services internally. The Amazon retail site seems available, but I’m curious if it’s even using native AWS or is still on the old internal compute platform. Makes me wonder how much juice this company has left.
- NelsonMinar 1y agoI saw a quote from a high end AWS support engineer that said something like "submitting tickets for AWS problems is not working reliably: customers are advised to keep retrying until the ticket is submitted".
- willsmith72 1y agoIt seems reasonable to me that Amazon (retail) would build better AZ redundancy into their services than say Snapchat or a bank
- melozo 1y agoSure, but it’s not reasonable that internal collaboration platforms built for ticketing engineers about outages doesn’t work during the outage. That would be something worth making multi-region at a minimum.
- willsmith72 1y agoOf course, referring to this > The Amazon retail site seems available
- nettlin 1y ago> The Amazon retail site seems available, but I’m curious if it’s even using native AWS or is still on the old internal compute platform. Some parts of amazon.com seem to be affected by the outage (e.g. product search: https://x.com/wongmjane/status/1980318933925392719 https://x.com/wongmjane/status/1980318933925392719)
- neon_me 1y ago"serverless"
- JPKab 1y agoThe length and breadth of this outage has caused me to lose so much faith in AWS. I knew from colleagues who used to work there how understaffed and inefficient the team is due to bad management, but this just really concerns me.
- llmslave 1y ago"Tech people" are long gone, most projects are death marches of technical debt
- 0x5345414e 1y agoThis is having a direct impact on my wellbeing. I was at Whole Foods in Hudson Yards NYC and I couldn’t get the prime discount on my chocolate bar because the system isn’t working. Decided not to get the chocolate bar. Now my chocolate levels are way too low.
- tonymet 1y ago"alexa turn on coffee pot" stopped working this morning, and I'm going bonkers.
- jdlyga 1y agoAlexa is super buggy now anyway. I switched my Echo Dot to Alexa+, and it fails turning on and off my Samsung TV all the time now. You usually have to do it twice.
- tonymet 1y agoi agree. the new LLM is better for dialog and Q&A, but they haven't properly tested intents and IOT integration at all.
- clbrmbr 1y agoHow can this be? I had great luck with GPT3 way back when… and I didn’t have function calling or chat… had to parse the JSON myself, extraction “action” and “response-text” fields… How has this been so hard for AMZN? Is it a matter of token cost and trying to use small models?
- tonymet 1y agothat's a reasonable theory. they've likely delayed the launch this long due to the inference cost compared to the more basic Alexa engine. I would also guess the testing is incomplete. Alexa+ is a slow roll out so they can improve precision/recall on the intents with actual customers. Alexa+ is less deterministic than the previous model was wrt intents
- dabinat 1y agoMy site was down for a long time after they claimed it was fixed. Eventually I realized the problem lay with Network Load Balancers so I bypassed them for now and got everything back up and running.
- YouAreWRONGtoo 1y agoI don't get how you can be a trillion dollar company and still suck this much.
- jjice 1y agoWe got off pretty easy (so far). Had some networking issues at 3am-ish EDT, but nothing that we couldn't retry. Having a pretty heavily asynchronous workflow really benefits here. One strange one was metrics capturing for Elasticache was dead for us (I assume Cloudwatch is the actual service responsible for this), so we were getting no data alerts in Datadog. Took a sec to hunt that down and realize everything was fine, we just don't have the metrics there. I had minor protests against us-east-1 about 2.5 years ago, but it's a bit much to deal with now... Guess I should protest a bit louder next time.
- motiejus 1y agoToo big to recover.
- jrm4 1y agoHey wait wasn't the internet supposed to route around...?
- BryanBeshore 1y agohttps://www.youtube.com/shorts/liL2VXYNyus https://www.youtube.com/shorts/liL2VXYNyus
- cpncrunch 1y ago"The root cause is an underlying internal subsystem responsible for monitoring the health of our network load balancers." https://health.aws.amazon.com/health/status?path=service-history https://health.aws.amazon.com/health/status?path=service-his...
- vmnb 1y agoAh, it's just what I thought. An underlying internal subsystem.
- Aldipower 1y agoaltavista.com is also down!
- the-chitmonger 1y agoI'm not sure if this is directly related, but I've noticed my Apple Music app has stopped working (getting connection error messages). Didn't realize the data for Music was also hosted on AWS, unless this is entirely unrelated? I've restarted my phone and rebooted the app to no avail, so I'm assuming this is the culprit.
- raw_anon_1111 1y agoFrom the great Corey Quinn Ah yes, the great AWS us-east-1 outage. Half the internet’s on fire, engineers haven’t slept in 18 hours, and every self-styled “resilience thought leader” is already posting: “This is why you need multi-cloud, powered by our patented observability synergy platform™.” Shut up, Greg. Your SaaS product doesn’t fix DNS, you're simply adding another dashboard to watch the world burn in higher definition. If your first reaction to a widespread outage is “time to drive engagement,” you're working in tragedy tourism. Bet your kids are super proud. Meanwhile, the real heroes are the SREs duct-taping Route 53 with pure caffeine and spite. https://www.linkedin.com/posts/coquinn_aws-useast1-cloudcomputing-activity-7386091910320308224-73kb?utm_source=share&utm_medium=member_ios&rcm=ACoAAAokImIBCa4JR7mlUp5o3BseZArP8KlN550 https://www.linkedin.com/posts/coquinn_aws-useast1-cloudcomp...
- Rooster61 1y agoThis wins all the internets today. Probably not a day where internets are particularly valuable, but it wins them nonetheless.
- raw_anon_1111 1y agoHe once said about an open source project that I was the third highest contributor on at AWS “This may be the worst AWS naming of 2021.” It was one of the proudest moments in my career. Yes I know it’s sad…
- t1234s 1y agoDo events like this stir conversations in small to medium size businesses to escape the cloud?
- rovr138 1y agoDepends how small. I have clients and I’ve heard “even Amazon is down, we can be down” more than once.
- FigurativeVoid 1y agoIt would have to be catastrophic for most businesses to make think about escaping the cloud. The cost of migration and maintenance are massive for small and medium businesses.
- tonymet 1y agoThis isn't a "cloud failure". All of these apps would be running now had they spent the additional 5% development costs to add failover to another region.
- decimalenough 1y agous-east1 is supposed to consist of a number of "availability zones" that are independent and stay "available" even if one goes down. That's clearly not happening, so yes, this is very much a cloud failure.
- tonymet 1y agoIt's an AWS failure, perhaps, but it's not a reason to write off "the cloud"
- TacticalCoder 1y ago[dead]
- ilikecakeandpie 1y agoProbably, but they're usually dropped after they come back up
- tonymet 1y agoI don't think blaming AWS is fair, since they typically exceed their regional and AZ SLAs AWS makes their SLAs & uptime rates very clear, along with explicit warnings about building failover / business continuity. Most of the questions on the AWS CSA exam are related to resiliency . Look, we've all gone the lazy route and done this before. As usual, the problem exists between the keyboard and the chair.
- dijit 1y agoNot sure any of their SLA’s are covered here. If they don’t obfuscate the downtime (they will, of course), this outage would put them at, what, two nines? Thats very much out of their SLA. People also keep talking about it as if its one region, but there are reports in this thread of internal dependencies inside AWS which are affecting unrelated regions with various services. (r53 updates for example)
- tonymet 1y agoSounds like your lesson is "yes we should continue shaming AWS rather than fix our app"
- rester324 1y agoIt sounds like you think the SLA is just toilet paper? When in reality it's a contract which defines AWS's obligations. So the lesson here is that they broke their contract big time. So yes. Shaming is the right approach. Also it seems you missed somehow the other 1700+ comments agreeing with shaming
- tonymet 1y agoI wouldn't go that far. The SLA is a contract, and they are clear on the remedy (up to 100% refund if they don't hit 95% uptime in a month). Just like reading medication side effects, they are letting you know that downtime is possible, albeit unlikely. All of the documentation and training programs explain the consequence of single-region deployments. The outage was a mistake. Let's hope it doesn't indicate a trend. I'm not defending AWS. I'm trying to help people translate the incident into a real lesson about how to proceed. You don't have control over the outage, but you do have control over how your app is built to respond to a similar outage in the future.
- chermi 1y agoStupid question, why isn't the stock down? Couldn't this lead to people jumping to other providers and at the very least require some pretty big fees for do dramatically breaking SLA? Is it just not a biggest fraction of revenue to matter?
- Archonical 1y agoMaybe since Amazon is reporting Q3 numbers soon and this will only show up in Q4 numbers?
- ilikecakeandpie 1y agoLots of stock brokerages consumer offerings are served through.... ....AWS!
- chasd00 1y agoRobinhood is down ;)
- Capricorn2481 1y agoNon-technical people don't really notice these things. They hear it and shrug, because usually it's fixed within a day. CNBC is supposed to inform users about this stuff, but they know less than nothing about it. That's why they were the most excited about the "Metaverse" and telling everyone to get on board (with what?) or get left behind. The market is all about perception of value. That's why Musk can tweet a meme and double a stocks price, it's not based in anything real.
- chasd00 1y agowow I think most of Mulesoft is down, that's pretty significant in my little corner of the tech world.
- 8cvor6j844qw_d6 1y agoThat's unusual. I wss under the impression that having multiple available zones guarantees high availability. It seems this is not the case.
- ykl 1y agoOne of the open secrets of AWS is that even though AWS has a lot of regions and availability zones, a lot of AWS services have control planes that are dependent on / hosted out of us-east-1 regardless of which region / AZ you're using, meaning even if you are using a different availability zone in a different region, us-east-1 going down still can mess you up.
- worik 1y agoThis outage is a reminder: Economic efficiency and technical complexity are both, separately and together, enemies of resilience
- _pvzn 1y agoCan confirm, also getting hit with this.
- EbNar 1y agoMay be because of this that trying to pay with PayPal on Lenovo's website has failed thrice for me today? Just asking... Knowing how everything is connected nowadays it wouldn't surprise me at all.
- hippo77 1y agoFinally an upside to running on Oracle Cloud!
- haunter 1y agoThe Premier League said there will be only limited VAR today w/o the automatic offside system becasue of the AWS outage. Weird timeline we live in https://www.bbc.com/news/live/c5y8k7k6v1rt?post=asset%3Ad9021236-e1c2-41d4-8c1a-283a37172945#post https://www.bbc.com/news/live/c5y8k7k6v1rt?post=asset%3Ad902...
- magarnicle 1y agoA silver lining to this cloud (outage).
- iammrpayments 1y agoWhy is VAR connected to the internet? Are they trying to gather data on offside players customers to improve recommedations?
- ashikns 1y agoI worked in a similar system. The raw data from the field first goes to a cloud hosted event queue of some sort, then a database, then back to whatever app/screen on field. The data doesn't just power on-field displays. There's a lot of online websites, etc that needs to pull data from an api.
- afavour 1y agoI wouldn't be at all surprised if people pay for API access to the data. I've worked with live sports data before, it's a very profitable industry to be in when you're the one selling the data. Of course in a sane world you'd have an internal fallback for when cloud connectivity fails but I'm sure someone looked at the cost and said "eh, what's the worst that could happen?"
- l33tnull 1y agoI can't do anything for school because Canvas by Instructure is down because of this.
- arrty88 1y agoI expect gcp and azure to gain some customers after this
- kevinsundar 1y agoAWS pros know to never use us-east-1. Just don't do it. It is easily the least reliable region
- j45 1y agoMore and more I want to be could agnostic or multi-cloud.
- rwke 1y agoWith more and more parts of our lives depending on often only one cloud infrastructure provider as a single point of failure, enabling companies to have built-in redundancy in their systems across the world could be a great business. Humans have built-in redundancy for a reason.
- IOT_Apprentice 1y agoApparently IMDb, an Amazon service is impacted. LOL, no multi region failover.
- iwontberude 1y agoworst outage since xmas time 2012
- artyom 1y agoAmazon has spent most of its HR post-pandemic efforts in: • Laying off top US engineering earners. • Aggressively mandating RTO so the senior technical personnel would be pushed to leave. • Other political ways ("Focus", "Below Expectations") to push engineering leadership (principal engineers, etc) to leave, without it counting as a layoff of course. • Terminating highly skilled engineering contractors everywhere else. • Migrating serious, complex workloads to entry-level employees in cheap office locations (India, Spain, etc). This push was slow but mostly completed by Q1 this year. Correlation doesn't imply causation? I find that hard to believe in this case. AWS had outages before, but none like this "apparently nobody knows what to do" one. Source: I was there.
- 1970-01-01 1y agoCompletely detached from reality, AMZN has been up all day and closed up 1.6%. Wild.
- Loughla 1y agoUntil this impacts their bottom line, how is that unexpected? Will we see mass exits from their service? Who knows. My money says no though. How many companies can just ride the "but it's not our fault" to buy time with customers until it's fixed?
- deleted 1y ago[deleted]
- homeonthemtn 1y ago"We should have a fail back to US-West." "It's been on the dev teams list for a while" "Welp....."
- nullorempty 1y agoIt won't be over until long after AWS resolves it - the outages produce hours of inconsistent data. It especially sucks for financial services, things of eventual consistency and other non-transactional processes. Some of the inconsistencies introduced today will linger and make trouble for years.
- president_zippy 1y agoI wonder how much better the uptime would be if they made a sincere effort to retain engineering staff. Right now on levels.fyi, the highest-paying non-managerial engineering role is offered by Oracle. They might not pay the recent grads as well as Google or Microsoft, but they definitely value the principal engineers w/ 20 years of experience.
- valdiorn 1y agoI missed a parcel delivery because a computer server in Virginia, USA went down, and now the doorbell on my house in England doesn't work. What. The. Fork. How the hell did Ring/Amazon not include a radio-frequency transmitter for the doorbell and chime? This is absurd. To top it off, I'm trying to do my quarterly VAT return, and Xero is still completely borked, nearly 20 hours after the initial outage.
- hacker_homie 1y agoIt's always DNS
- AtomicOrbital 1y agohttps://m.youtube.com/watch?v=KFvhpt8FN18 https://m.youtube.com/watch?v=KFvhpt8FN18 clear detailed explanation of the AWS outage and how properly designed systems should have shielded the issue with zero client impact
- dorongrinstein 1y agoAnyone needing multi-cloud WITH EASE, please get in touch. https://controlplane.com https://controlplane.com I am the CEO of the company and started it because I wanted to give engineering teams an unbreakable cloud. You can mix-n-match services of ANY cloud provider, and workloads failover seamlessly across clouds/on-prem environments. Feel free to get in touch!
- rsanheim 1y agoI wonder what kind of outage or incident or economic change will be required to cause a rejection of the big commercial clouds as the default deployment model. The costs, performance overhead, and complexity of a modern AWS deployment are insane and so out of line with what most companies should be taking on. But hype + microservices + sunk cost, and here we are.
- newZWhoDis 1y agoHonest answer? The outage would need to last about a week.
- babl-yc 1y agoI don't expect the majority of tech companies to want to run their own physical data centers. I do expect them to shift to more bare-metal offerings. If I'm a mid to large size company built on DynamoDB, I'd be questioning if it's really worth the risk given this 12+ hour outage. I'd rather build upon open source tooling on bare metal instances and control my own destiny, than hope that Amazon doesn't break things as they scale to serve a database to host the entire internet. For big companies, it's probably a cost savings too.
- baobabKoodaa 1y ago> For big companies, it's probably a cost savings too. For any sized company, moving away from big clouds back onto traditional VPS or bare-metal offerings will lead to cost savings.
- guerrilla 1y agoSo then we should expect it in the long-term regardless of outages anyway for the sake of growth slone.
- einsteinx2 1y agoI think that prediction severely underestimates the amount of cargo culting present at basically every company when it comes to decisions like this. Using AWS is like the modern “no one ever got fired for buying IBM”.
- hamonrye 1y ago[dead]
- a-dub 1y agoi am amused at how us-east-1 is basically in the same location as where aol kept its datacenters back in the day.
- fastball 1y agoOne of my co-workers was woken up by his Eight Sleep going haywire. He couldn't turn it off because the app wouldn't work (presumably running on AWS).
- cranberryturkey 1y agoSling still down at 11:42PM PST
- cranberryturkey 1y agoKraken can't do deposits either at 3:38am PST
- ct_list 1y ago[dead]
- ronakjain90 1y agowe[1] operate out of `us-east-1` but chose to not use any of the cloud based vendor lockin (sorry vercel, supabase, firebase, planetscale etc). Rather a few droplets in DigitalOcean(us-east-1) and Hetzner(eu). We serve 100 million requests/mo, few million user generated content(images)/mo at monthly cost of just about $1000/mo. It's not difficult, it's just that we engineers chose convenience and delegated uptime to someone else. [1] - https://usetrmnl.com https://usetrmnl.com
- deleted 1y ago[deleted]
- throw-10-13 1y agoimagine spending millions on devops and sre to still have your mission critical service go down because amazon still has baked in regional dependencies
- teunlao 1y agous-east-1 down again. We all know we should leave. None of us will.
- michaelcampbell 1y agoAnthem Health call center disconnected my wife numerous times yesterday with an ominous robo-message of "Emergency in our call center"; curious if that was this. Seems likely, but what a weird message.
- ky_vulnerable 1y agoDo we know what caused the outage yet?
- AbstractH24 1y agoAre there websites that do post-mortems for how the single points of failure impacted the entire internet? Not just AWS, but Cloudflare and others too. Would be interesting to review them clinically.
- jodrellblank 1y agoAnother time to link The Machine Stops by E.M. Forster, 1909: https://web.cs.ucdavis.edu/~rogaway/classes/188/materials/the%20machine%20stops.pdf https://web.cs.ucdavis.edu/~rogaway/classes/188/materials/th... > “The Machine,” they exclaimed, “feeds us and clothes us and houses us; through it we speak to one another, through it we see one another, in it we have our being. The Machine is the friend of ideas and the enemy of superstition: the Machine is omnipotent, eternal; blessed is the Machine.” .. > "she spoke with some petulance to the Committee of the Mending Apparatus. They replied, as before, that the defect would be set right shortly. “Shortly! At once!” she retorted" .. > "there came a day when, without the slightest warning, without any previous hint of feebleness, the entire communication-system broke down, all over the world, and the world, as they understood it, ended."
- 0xbadcafebee 1y agoWe never went down in us-east-1 during this incident. We have tons of high-traffic sites/services. Not multi-region, not multi-cloud. You're gonna hear mostly complaints in this thread, but simple, resilient, single-region architecture is still reliable as hell in AWS, even in the worst region.