14 ms·
AWS us-east-2 outage
our alerts just went crazy, and we're having issues even logging in to the AWS dashboard
possibly related: https://news.ycombinator.com/item?id=32267154 https://news.ycombinator.com/item?id=32267154
- rwky 4y agoSame. Our EC2 instances can't connect RDS and just got 500 errors on the dashboard.
- kaustubhvp 4y agoZoom is having connectivity issues. Nothing on https://status.zoom.us/ https://status.zoom.us/ page yet.
- jmartens 4y agoCan confirm based on what Metrist is seeing. Looks to be a larger issue in the US East, also seeing Cloudfare and Datadog with issues.
- kaustubhvp 4y agoyeah looks like us-east-2 has networking issues
- sjm-lbm 4y agoSame here - we were finally able to log in to the console, but we're in us-east-2 and are having a ton of issues.
- allenrb 4y agous-east-2 customer here, having a variety of "strange" issues including inability to reach an RDS database, other users in my firm having VPN reconnect trouble to that region.
- allenrb 4y agoFWIW, most of our issues just resolved ~1 minute ago. We'll see if it remains stable.
- kaustubhvp 4y agoAWS acknowledged the issue https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status
- allenrb 4y ago"[10:11 AM PDT] We are investigating network connectivity issues for some instances and increased error rates and latencies for the EC2 APIs within the US-EAST-2 Region."
- allenrb 4y ago[10:25 AM PDT] We can confirm that some instances within a single Availability Zone (USE2-AZ1) in the US-EAST-2 Region have experienced a loss of power. The loss of power is affecting part of a single data center within the affected Availability Zone. Power has been restored to the affected facility and at this stage the majority of the affected EC2 instances have recovered. We expect to recover the vast majority of EC2 instances within the next hour. For customers that need immediate recovery, we recommend failing away from the affected Availability Zone as other Availability Zones are not affected by this issue.
- rwky 4y agoFunny I failed away from the zone and RDS still doesn't work, connections fail. Edit: 2 minutes after I post this it starts working.
- corobo 4y agoAhh Always check HN before trying to diagnose weird issues that shouldn't be connected
- jmartens 4y agoToo bad that is the best we have!
- Syonyk 4y agoAnd the reason that works is because HN is mostly hosted on its own stuff, without weird dependencies on anything beyond "the servers being up" and "TCP mostly working."
- EddySchauHai 4y agoI believe it's on AWS after its two servers broke at the same time the other day.
- dang 4y agoLiving a bit more dangerously at the moment as HN is still running temporarily on AWS. (I'd link to the threads about this from a few weeks ago but am on my phone ATM.)
- corobo 4y agoI did notice it being a little slow but I'm also on 4G at the moment (it got the blame)
- annamarie 4y agoSame
- whendon 4y agoTons of issues in us-east-2 here as well
- ilikecakeandpie 4y agous-east-2 customer here also having some issues
- deleted 4y ago[deleted]
- nnf 4y agoLots of things intermittently unreachable in us-east-2 for us, across multiple AWS accounts.
- MikeMacMan 4y agoI'm seeing issues in 2a but not 2b. Anyone having issues in 2b?
- NovemberWhiskey 4y agoJust FYI; availability zone names in AWS are randomized between accounts - your "a" can be someone else's "f" Physical identifiers for availability zones are like "use2-az1", ref. https://docs.aws.amazon.com/ram/latest/userguide/working-with-az-ids.html https://docs.aws.amazon.com/ram/latest/userguide/working-wit...
- bradyd 4y agoAvailability zones are not guaranteed to have the same name across accounts (ie. us-east-2a in one account might be us-east-2d in another). You would need to use the AZ-ID to determine if they are the same. https://docs.aws.amazon.com/prescriptive-guidance/latest/patterns/use-consistent-availability-zones-in-vpcs-across-different-aws-accounts.html https://docs.aws.amazon.com/prescriptive-guidance/latest/pat...
- silverlyra 4y agoAWS availability zones are randomly shuffled for each AWS account – your us-east-2a won't (necessarily) be the same as another user's (or even another account in the same organization): https://docs.aws.amazon.com/ram/latest/userguide/working-with-az-ids.html https://docs.aws.amazon.com/ram/latest/userguide/working-wit... You'll need to see which availability zone ID (e.g., use2-az3) corresponds to each zone in your account: https://aws.amazon.com/premiumsupport/knowledge-center/vpc-map-cross-account-availability-zones/ https://aws.amazon.com/premiumsupport/knowledge-center/vpc-m... edit: AWS identified this as a power loss in a single zone, use2-az1.
- blahyawnblah 4y agoI wonder if this is done because people have a tendency or something to always create resources in 'A' (or some other AZ) and this helps spread things around. And if I would have read the page the link points to better, that's exactly the reason
- 02thoeva 4y agoThanks for sharing. We've just spent the last hour debugging our website, thinking we had issues. This explains it.
- jmartens 4y agoInterestingly, we saw a bunch of other services degrade (Zoom, Zendesk, Datadog) before AWS services themselves degrade.
- kaustubhvp 4y agoLooks like a lot of services are impacted including Cloudflare, Ping, Zoom and Datadog.
- svnpenn 4y agoLooks like Snap, Crackle and Pop are down as well.
- collinvandyck76 4y agoHah, frequency illusion strikes again? I just learned about the derivatives past Jerk yesterday.
- CoastalCoder 4y ago> Looks like Snap, Crackle and Pop are down as well. I don't work on cloud stuff, so I'm genuinely unsure if this is a joke.
- fragmede 4y agoIt's a joke but I only knew that because Snap is/was (as of S1) hosted on GCP and not AWS. Crackle happens to be the name of a video on demand company. It's a reference to the breakfast cereal of the same name.
- ebabani 4y agoPop is also a screen sharing/pairing tool, so the joke was great.
- tatersolid 4y agoPedantic clarification for the unfamiliar: the breakfast cereal is named Rice Krispies while Snap, Crackle, and Pop are the names of the cartoon mascots on the box.
- qeternity 4y agoCloudflare uses AWS? For what?
- aaur0 4y agoAWS reported the outage here : https://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status
- monocasa 4y ago> Severity > Informational lol
- Analemma_ 4y agoThree things are certain: death, taxes, and useless cloud status dashboards.
- chrismartin 4y agoLies, damned lies, and self-reported statistics that affect a company's SLA refund liability.
- mlrtime 4y agoAnd SE/SRE reviews
- hakube 4y agoI'm running Terraform and it appears to be stuck now. What do I do??
- NeckBeardPrince 4y agoWait
- danw1979 4y agoDepends what it’s stuck doing, but you might ctrl-c it and later manually unlock the state file (by carefully coordinating with colleagues and deleting the dynamo DB lock object if you’re using the s3 backend) when the outage is over.
- eurasiantiger 4y agoThanks, this comment made it very clear to me that I never want to touch a terraform system.
- deathanatos 4y agoTF makes API calls to the underlying cloud. If those hang, you'll have to wait for them to time out. Whether TF can update the state & release its locks would depend on where those were hosted. If they're in the downed AZ, then ofc. it can't do that, and manual intervention will be required afterwards. I forget if you can make those objects regional when stored in AWS or not. (You can in some other storages.) … what would you expect to happen here?
- mdaniel 4y agoFun fact, for a lot of providers, it'll hang on any error, not just cloud ones. I presume it's due to the gRPC communication mechanism and the terraform binary blocking until the provider answers "yes or no" to the request
- datatrashfire 4y agoI think any system is susceptible to problems like this if the underlying hardware becomes unavailable. Using dynambodb to obtain locks on s3 is a pretty common pattern in AWS development. This has more to do with AWS than Terraform.
- mrwnmonm 4y agoAnd I thought us-east-2 is the way to escape us-east-1's problems.
- snapcaster 4y agoHaha literally had this same thought. us-east-2 is our default region for most stuff and so far that's been good. I think this is the first AWS downtime in last couple years that hit our systems directly
- nemothekid 4y agoI've somehow dodged region outages on AWS for years, and here's my first one. So many alerts firing off in unexpected ways.
- Justin_K 4y agoCan confirm....
- eatonphil 4y agoWonder if this is why Zoom is down. Wasn't able to connect just now. The connection proxy/sites were giving 504s.
- bluedino 4y agoWebEx too
- jasonjayr 4y agoIIRC Zoom signed up with Oracle Cloud when COVID hit and they needed to scale like crazy. https://www.oracle.com/customers/zoom/ https://www.oracle.com/customers/zoom/ I'm not sure if Zoom has any Critical infra in AWS though.
- easton 4y agoI interviewed there a few months ago for DevOps, and one of the people I interviewed with said that most of Zoom was in AWS (they liked that I had AWS stuff on my resume).
- adrr 4y agoOkta is degraded as well.
- bigfatfrock 4y agoSuppose it's time to setup multi-az and pay to insure against AWS' own failures. I don't know why I previously thought their EC2 uptime claims were sufficient. Lesson learned.
- okdood64 4y agoWhat are you hosting? Multi-AZ seems like a bare minimum for basic reliability. That said it's not a panacea. There's all sorts of cascading/downstream "weirdness" that can result on AWS's own services through the loss of an AZ.
- Johnny555 4y agoAre you sure you understand their uptime claims? They offer a 99.99% SLA for regional availability, but only 99.5% for individual instances (and even then, they only owe you a 10% service credit for affected instances) https://aws.amazon.com/compute/sla/ https://aws.amazon.com/compute/sla/ 99.5% availability allows up to about 3 and a half hours of downtime a month. 99.99% means around 4 minutes a month. So if you can't handle hours of downtime, you should definitely be multi-AZ.
- chronid 4y agoMulti-AZ is a requirement on production level loads if you cannot sustain prolonged downtime. Datacenters do end up completely dying now and then, you really want to have a good strategy in that case. Or not, if that's not required.
- jmartens 4y agoAnyone else notice similar issues in US-West-2 a few hours before this issue in US-East-2?
- bkruse 4y agoLots of issues in us-east-2 for instances for us but also other regions when connecting to RDS
- jedberg 4y agoSorry all I jinxed it. Yesterday I was in a meeting and said "The only regional outages AWS has ever had were in us-east-1, so we should just move to us-east-2." Now I guess we have to move to us-west-2. :) Update: looks like it's only one zone anyway, so my statement still stands!
- acwan93 4y agoIn all seriousness, we've been deploying everything on us-west-2, and it seems to have dodged most of the outages recently. Is there something special about that data center?
- arecurrence 4y agoClassically, us-east-1 received most of the hate given its immense size (it used to be several times larger than any other) and status as the first large aws data center. It also seemed to launch new aws features first but that may have been my imagination. If true, I'm sure always running the latest builds was not great for stability. us-west-2 has had outages as well but it is less common, even rare. I've been pushing companies to make their initial deployments onto us-west-2 for over ten years now. I occasionally get kudos messages in my inbox :)
- teknopaul 4y ago93.99999
- doubled112 4y agoThere are six 9s in there. Pretty solid!
- rjh29 4y agoMaybe Amazon should make us-east-1's actual datacenter change depend on the customer, as they do with the AZs :P
- creeble 4y ago
- packetslave 4y agoUpdate from AWS: they lost power to (part of?) a single DC in the use2-az1 availability zone. 10:25 AM PDT We can confirm that some instances within a single Availability Zone (USE2-AZ1) in the US-EAST-2 Region have experienced a loss of power. The loss of power is affecting part of a single data center within the affected Availability Zone. Power has been restored to the affected facility and at this stage the majority of the affected EC2 instances have recovered. We expect to recover the vast majority of EC2 instances within the next hour. For customers that need immediate recovery, we recommend failing away from the affected Availability Zone as other Availability Zones are not affected by this issue.
- alfalfasprout 4y agoInteresting to see it's been a loss of power that caused this. Usually the better datacenters have multiple levels of power redundancy including emergency backup generators.
- throwawaymaths 4y agoInsert clip of O'Brien explaining to cardassians why there are backups for backups
- laumars 4y agoIn case anyone is unaware of the reference, that’s taken from Star Trek Deep Space 9 https://youtu.be/UaPkSU8DNfY https://youtu.be/UaPkSU8DNfY
- jmartens 4y agoYa, am I surprised by this too. Like, you have one job, keep the power on.
- packetslave 4y agoIt depends entirely on how AWS architected their power redundancy. Given that the outage affected a portion of one DC in one AZ, we can make some assumptions, but the truth is we just don't know. It could be that their shared-fate scope is an entire data hall, or a set of rows, or even an entire building given that an AZ is made up of multiple datacenters. I don't know that AWS has ever published any kind of sub-AZ guarantees around reliability. Datacenter power has all kinds of interesting failure modes. I've seen outages caused by a cat climbing into a substation, rats building a nest in a generator, fire-fighting in another part of the building causing flooding in the high-voltage switching room, etc.
- packetslave 4y agoOn the bright side, this is certainly a good test for "exactly how resilient are our AWS-based systems to the loss of a single availability zone?"
- dudeinjapan 4y agoVery resilient, provided that its not the AZ I'm using.
- jupp0r 4y agoChaos monkey is on the move again.
- jdugan 4y agoWe have RDS and ECS issues in us-east-2
- ctur 4y agoI find this spreadsheet handy for thinking about AWS region-wide outages and frequency. This seems to be the first major us-east-2 outage, indeed, vs us-east-1 and other regions. https://docs.google.com/spreadsheets/d/1Gcq_h760CgINKjuwj7WuRmLXHIdvsUdzNQCg0g4QvVs/edit https://docs.google.com/spreadsheets/d/1Gcq_h760CgINKjuwj7Wu... (from https://awsmaniac.com/aws-outages/ https://awsmaniac.com/aws-outages/)
- joshstrange 4y agoJust lost my email provider (https://status.postmarkapp.com/incidents/240161 https://status.postmarkapp.com/incidents/240161) to this and I'd bet my services are degraded/down. I know it's never a "good" time for an outage but this sure does suck for me right now. We've got an event this weekend and people can't sign up right now to buy tickets/etc.
- numpad0 4y ago$ dig news.ycombinator.com ;; ANSWER SECTION: news.ycombinator.com. 1 IN A 50.112.136.166 $ dig -x 50.112.136.166 ;; ANSWER SECTION: 166.136.112.50.in-addr.arpa. 300 IN PTR ec2-50-112-136-166.us-west-2.compute.amazonaws.com. saving couple keypresses just in case
- linsomniac 4y agoI understand that us-east is AWS's oldest and biggest facility, but Amazon seems to have more money than Croesus, why aren't they fixing/rebuilding/replacing us-east with something more modern?
- ctvo 4y agous-east is a geographic distinction within which there are multiple regions. us-east-1 and us-east-2 are not the same. This outage occurred in us-east-2. Within an AWS region there are multiple data centers. They call their data centers availability zones. The availability zone AZ1 was the one impacted, and within that availability zone, most likely only a subset of servers. us-east-1 is the region you're thinking of that has issues. Mostly due to being the largest region (I think?) and like you mentioned, the oldest.
- dragonwriter 4y ago> us-east-1 is the region you're thinking of that has issues. Mostly due to being the largest region (I think?) and like you mentioned, the oldest. Also, because shared and global AWS resource are (or at least often behave as if they are) intimately tied to us-east-1.
- philihp 4y agoMy first instinct would be to guess that something like this happened because of some intentional and well-meaning effort to upgrade some critical part of their infrastructure. Just my hunch given that it happened during the middle of the week in the middle of the day, and came back relatively quickly. The quick but not instantaneous bounce back has the hallmark of someone following a carefully laid out worst case contingency plan. I look forward to the postmortem.
- _xnmw 4y agoBecause money can't fix everything? In fact sometimes having too much money makes it worse, as YC startup wisdom says.
- deleted 4y ago
- mathgladiator 4y agoNo wonder my error logs were clean. I can get into my hosts, but my LB isn't routing. Sad face.
- smm11 4y agoHad Rackspace login issues earlier today. Hmmmmmm ...
- ChrisArchitect 4y agoWe have always been at war with us-east-2.
- poxrud 4y agoThe fact that so many popular sites/services are experiencing issues due to a single AZ failure makes me think that there is a serious shortage of good cloud architects/engineers in the industry. It would be one thing if this was a Regional failure, but a single AZ failure should not have any noticeable effect.
- dangero 4y agoFor most businesses a little down time here and there is a calculated risk versus more complex infrastructure. You can’t assume all the cloud architects are idiots — they have to report their task list and cost of infrastructure to someone who can give feedback on various options based on comparative resource requirements and risks. Zone downtime still falls under an AWS SLA so you know about how much downtime to accept and for a lot of businesses that downtime is acceptable.
- ilikecakeandpie 4y agoThis is true, but I think it would be more acceptable if the region were down vs the single AZ
- water-your-self 4y agoIt gives me a bad gut feeling when you imply that multiple instances of a service is more complex than a single instance which cannot be duplicated easily. I also disagree that it is inherently more costly to run a service in multiple locations.
- mrits 4y agoYou should get into the database business. A lot of money to be made there if things are so trivial for you.
- willmadden 4y agoThe sounds of crickets is deafening!
- jonatron 4y agoI just set up a few small sites (not live yet) on us-east-2, because us-east-1 has a poor reputation. I wanted to avoid multi-region to keep things simple, but now I'm thinking I might have to spend the additional time on it. Not ideal when there's no dedicated ops.
- dmalvarado 4y agoI dunno how else to put it. Having EVERYTHING on AWS is a national security threat. This isn't good, and someone who can do something about it needs to.
- harryh 4y agoThere is no "someone" who could do anything about this.
- gtirloni 4y agoGood thing we don't have EVERYTHING on AWS, so no threat detected.
- smt88 4y agoOur own applications are hosted on Azure, but we had an outage today anyway. It was because apparently Netlify and Auth0 use AWS and went down, which took down our static sites and our authentication. The nature of our business means it wasn't a big deal, but I could imagine lots of people were in the same boat.
- outworlder 4y agoYeah, let's place everything in large colos instead. Those never fail, right?
- Nextgrid 4y agoBut the colos aren't usually managed by a single control plane controlled by a single company, so while they can all fail, they will generally do so independently.
- scubbo 4y agoHaving _everything_ on a single AZ of AWS is, indeed, a problem. Having everything well-architected on AWS is...well, it's a problem for reasons of monopoly and cost, but it's not a problem for availability.
- weeeeelp 4y agoAnyone's ECR endpoints went out during the outage? We've had timeouts while pulling images onto our k8s cluster post-restart
- tpl 4y agoPersonally only saw intermittent failures through out. Rather minor production as far as AWS outages go.
- myroon5 4y agounfortunately one of the largest regions: https://github.com/patmyron/cloud/#ip-addresses-per-region https://github.com/patmyron/cloud/#ip-addresses-per-region
- hamzatahirdevc 4y ago