8 ms·
Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ https://status.aws.amazon.com/ > 8:22 AM PST We are investigating inc
by ipmb 5y ago
Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ https://status.aws.amazon.com/
> 8:22 AM PST We are investigating increased error rates for the AWS Management Console.
> 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US-EAST-1. Customers may be able to access region-specific consoles going to https://console.aws.amazon.com/ https://console.aws.amazon.com/. So, to access the US-WEST-2 console, try https://us-west-2.console.aws.amazon.com/ https://us-west-2.console.aws.amazon.com/
- bobviolier 5y agohttps://status.aws.amazon.com/ https://status.aws.amazon.com/ still shows all green for me
- banana_giraffe 5y agoIt's acting odd for me. Shows all green in Firefox, but shows the error in Chrome even after some refreshes. Not sure what's caching where to cause that.
- junyoon 5y agofirefox has more aggressive caching than other browsers I think
- jeremyjh 5y agoThey are still lying about it, the issues are not only affecting the console but also AWS operations such as S3 puts. S3 still shows green.
- lsaferite 5y agoIt's certainly affecting a wider range of stuff from what I've seen. I'm personally having issues with API Gateway, CloudFormation, S3, and SQS
- pbalau 5y ago> We are experiencing API and console issues in the US-EAST-1 Region
- jeremyjh 5y agoI read it as console APIs. Each service API has its own indicator, and they are all green.
- midasuni 5y agoOur corporate ForgeRock 2FA service is apparently broken. My services are behind distributed x509 certs so no problems there.
- Rantenki 5y agoYep, I am seeing failures on IAM as well: aws iam list-policies An error occurred (503) when calling the ListPolicies operation (reached max retries: 2): Service Unavailable
- silverlyra 5y agoSame here. Kubernetes pods running in EKS are (intermittently) failing to get IAM credentials via the ServiceAccount integration.
- bamboozled 5y agoI'm seeing errors for things that worked fine, like policies that had no issue now are saying "access denied". I'm wondering if the cause of the outage has to do with something changing in the way IAM is interpreted ?
- packetslave 5y agoIAM is a "global" service for AWS, where "global" means "it lives in us-east-1". STS at least has recently started supporting regional endpoints, but most things involving users, groups, roles, and authentication are completely dependent on us-east-1.
- Rapzid 5y agoI still can't create/destroy/etc CloudFront distros. They are stuck in "pending" indefinitely.
- guenthert 5y agoUh, four minutes to identify the root cause? Damn, those guys are on fire.
- czbond 5y ago:) I imagine it went like this theoretical Slack conversation: > Dev1: Pushing code for branch "master" to "AWS API". > <slackbot> Your deploy finished in 4 minutes > Dev2: I can't react the API in east-1 > Dev1: Works from my computer
- flerchin 5y agoOutage started at 731 PST from our monitoring. They are on fire, but not in a good way.
- Frost1x 5y agoIdentify or to publicly acknowledge? Chances are technical teams knew about this and noticed it fairly quickly, they've been working on the issue for some time. It probably wasn't until they identified the root cause and had a handful of strategies to mitigate with confidence that they chose to publicly acknowledge the issue to save face. I've broken things before and been aware of it, but didn't acknowledge them until I was confident I could fix them. It allows you to maintain an image of expertise to those outside who care about the broken things but aren't savvy to what or why it's broken. Meanwhile you spent hours, days, weeks addressing the issue and suddenly pull a magic solution out of your hat to look like someone impossible to replace. Sometimes you can break and fix things without anyone even knowing which is very valuable if breaking something had some real risk to you.
- sirmarksalot 5y agoThis sounds very self-blaming. Are you sure that's what's really going through your head? Personally, when I get avoidant like that, it's because of anticipation of the amount of process-related pain I'm going to have to endure as a result, and it's much easier to focus on a fix when I'm not also trying to coordinate escalation policies that I'm not familiar with.
- 5y ago
- stephenr 5y agoWhen I brought up the status page (because we're seeing failures trying to use Amazon Pay) it had EC2 and Mgmt Console with issues. I opened it again just now (maybe 10 minutes later) and it now shows DynamoDB has issues. If past incidents are anything to go by, it's going to get worse before it gets better. Rube Goldberg machines aren't known for their resilience to internal faults.
- jabiko 5y agoYeah, but I still have a different understanding what "Increased Error Rates" means. IMHO it should mean that the rate of errors is increased but the service is still able to serve a substantial amount of traffic. If the rate of errors is bigger than, let's say, 90% that's not an increased error rate, that's an outage.
- thallium205 5y agoThey say that to try and avoid SLA commitments.
- jiggawatts 5y agoSome big customers should get together and make an independent org to monitor cloud providers and force them to meet their SLA guarantees without being able to weasel out of the terms like this…
- giorgioz 5y agoI'm trying to login in the AWS Console from other regions but I'm getting HTTP 500. Anyone managed to login in other regions? Which ones? Our backend is failing, it's on us-east-1 using AWS Lambda, Api Gateway, S3
- jesboat 5y ago> This issue is affecting the global console landing page, which is also hosted in US-EAST-1 Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region? At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer https://internal.site.example/~someEngineer was expected to be multi-region, and was, because there was enough infrastructure for making things multi-region that it was usually pretty easy.
- balls187 5y agoMAANG* How long before Meta takes over for Facebook?
- all_usernames 5y ago
- dang 5y agoOk, we've changed the URL to that from https://us-east-1.console.aws.amazon.com/console/home https://us-east-1.console.aws.amazon.com/console/home since the latter is still not responding. There are also various media articles but I can't tell which ones have significant new information beyond "outage".
- JPKab 5y agoAs a user of Sagemaker in us-east-1, I deeply fucking resent AWS claiming the service is normal. I have extremely sensitive data, so Sagemaker notebooks and certain studio tools make sense for me. Or DID. After this I'm going back to my previous formula of EC2 and hosting my own GPU boxes. Sagemaker is not working, I can't get to my work (notebook instance is frozen upon launch, with zero way to stop it or restart it) and Sagemaker Studio is also broken right now. The length of this outage has blown my mind.
- markus_zhang 5y agoLooks like they removed some 9s from availability in one day. I wonder if more are considering moving away from cloud.
- wahern 5y agoYou don't use AWS because it has better uptime. If you've been around the block enough times, this story has always rung hollow. Rather, you use AWS because when it is down, it's down for everybody else as well. (Or at least they can nod their head in sympathy for the transient flakiness everybody experiences.) Then it comes back up and everybody forgets about the outage like it was just background noise. This is what's meant by "nobody ever got fired for buying (IBM|Microsoft)". The point is that when those products failed, you wouldn't get blamed for making that choice; in their time they were the one choice everybody excused even when it was an objectively poor choice. As for me, I prefer hosting all my own stuff. My e-mail uptime is better than GMail, for example. However, when it is down or mail does bounce, I can't pass the buck.
- pbreit 5y agoI like how 6 hours in: "Many services have already recovered".
- bamboozled 5y agoNot even close for us.