32 ms·
AWS Cognito is having issues and health dashboards are still green
- garymoon 6y agoWe experienced 504 errors from Cognito but seems to be that other services are affected as well
- ricms93 6y agoWe have experienced the same on multiple accounts
- simlevesque 6y agoI'm getting tired of that bullshit. Just admit it.
- that_guy_iain 6y agoAWS' conditions for what healthy is and everyone elses is completely different. Kinda makes me wonder what their internals are like.
- SteveNuts 6y agoProbably similarly to the "Downfall" Hitler parody, or the "he's delusional, take him to the infirmary" from Chernobyl, take your pick.
- ben509 6y agoI fully admit, with no reservations at all, that you're getting tired of that bullshit.
- mushufasa 6y agowhat a scam. who can hold them accountable for cheating those who paid for uptime guarantees? I guess the lawyers of those who paid for uptime guarantees...
- rexreed 6y agoThis is a relevant comment. I agree. Who will hold them accountable for violating uptime guarantees? Nobody? Then what's the point or purpose of an uptime guarantee? Marketing value?
- freehunter 6y agoIf it’s in a contract, companies can sue. And big enough customers who have lost enough money due to the outage would definitely threaten to sue to recover that money.
- redisman 6y agoWouldn't the contract specify the remedy? Or do they actually on paper really promise uptimes that on a single isolated data-center level no one can keep?
- kenhwang 6y agoWe just ask nicely. Never really had a problem getting a huge % discount on the bill because of an outage. Extra bonus for us since there's no bottom line impact since we can tolerate some downtime (just annoying for engineering).
- gt565k 6y agowah wah wah, my boss is chewing my ear out cuz I can't explain to him that it's a vendor (AWS) outage host on-premise then, and deal with power outages, and infrastructure management, and failing hardware shit happens, drink a coffee and sit back for 30 mins
- gtsop 6y agoHis point is not about aws inability to have 0 downtime. His point is some people apparently pay based on a guarantee for uptime that is not delivered. He is criticizing a company not providing the service it sells and our inability to hold them accountable
- dbenny 6y agoApparently they can't update the status page because of the outage. This happened a few years ago with the massive s3 outage.
- dubcanada 6y ago7:30 AM PST: We are currently blue on Kinesis, Cognito, IoT Core, EventBridge and CloudWatch given an increase in errors for Kinesis in the US-EAST-1 Region. It's not posted on SHD as the issue has impacted our ability to post there. We will update this banner if there continue to be issues with the SHD. Was posted 8 minutes ago.
- _0o6v 6y ago> It's not posted on SHD as the issue has impacted our ability to post there. Is that not a massive catch-22 for a service dashboard?
- snazz 6y agoThis has happened a few times before, actually. Dogfooding is good, but not for status pages! Cloudflare does it right for their status page (https://www.cloudflarestatus.com https://www.cloudflarestatus.com). They don't use Cloudflare itself for it (you can tell because /cdn-cgi/trace returns nothing), the actual backend is Atlassian Statuspage, their TLS certificate is issued by Let's Encrypt instead of Cloudflare itself, and it's on a completely separate domain for DNS purposes.
- toong 6y agoHave you checked https://www.githubstatus.com/ https://www.githubstatus.com/ ? ;-)
- lights0123 6y agoGitHub doesn't own their own datacenters.
- judge2020 6y agoThey do use their own registrar though: $ whois cloudflarestatus.com Registrar: Cloudflare, Inc.
- freehunter 6y agoReminds me of a recent outage from IBM Cloud, where the VPN was hosted on IBM Cloud so employees couldn’t log in to fix it, and the email is hosted on IBM Cloud so support teams couldn’t email customers to let them know and even access to their Twitter was behind the non-functional VPN so they couldn’t tweet during the outage either.
- kamyarg 6y ago
- reese_john 6y agoKinesis seems to be down to me. Everything is melting, it is like they have Chaos Monkey perpetually on in us-east-1
- snvzz 6y agoThis is what's called a SNAFU.
- LeoTM 6y agoExperiencing 504's from Cognito too, our users can't log in. "amazon-cognito-identity-js": "^3.2.2" "aws-amplify": "^2.2.2"
- simlevesque 6y agoyeah every amplify app should be down/super slow right now.
- tysoncadenhead 6y agoMaybe AWS should put their dashboards on GCP
- ti_ranger 6y ago> Maybe AWS should put their dashboards on GCP Then the status page would be almost entirely useless ...
- dsagal 6y agoBanner on top of https://status.aws.amazon.com/ https://status.aws.amazon.com/ just has an update from 8:36AM PST -- just removed -- even thought it's only 7:42AM PST. I guess it's really manual firefighting there.
- dtjones 6y agoWe're getting 504 for our well-known jwks file And request timeouts against cognito-idp.us-east-1.amazonaws.com And the cognito console won't load
- turdnagel 6y agoIsn't it common practice to host your status board on someone else's infrastructure? In 2017 there was an S3 issue that supposedly affected their ability to post. I believe they said that they were updating how they posted to the status board so that there would no longer be a dependency on S3. Well, I guess whatever they're dependent on now broke.
- WrtCdEvrydy 6y agoS3 East didn't affect the ability but they couldn't swap out the green checkmark for the red checkmark... which is just hilarious.
- tuwtuwtuwtuw 6y agoIt's common practice for small players but Amazon, Microsoft Azure and Google Cloud host their status pages on their own servers because they value the marketing aspect higher than a functioning status page for their customers.
- Frost1x 6y agoI find it surprising how many people forget how much underlying business motives drive pretty much every action they make and how this is quickly forgotten by many. No matter how much you value science and engineering, it ultimately doesn't matter to the business unless that aligns directly with their revenue stream. Sometimes it does, sometimes it doesn't.
- tuwtuwtuwtuw 6y agoYes. But I wonder if self-hosting their status page is really the correct decision from a marketing perspective. The people who consumes the status page on say Google Cloud probably know that Google self-hosting it is a bad decision from a technical point of view. So to the only people who care, their choice appear stupid. So I don't really understand what they gain by doing it. I think maybe I am wrong about it being a marketing concern and that the choice is more related to internal politics and incompetent management.
- arusahni 6y agoAll my CloudWatch alerts are firing "OK" transitions, and AWS ES isn't displaying any known instances
- tibbar 6y ago503s from CloudWatch for us.
- unilynx 6y agoThere's a lot more going on over there... - 7 cloudfront distributions created today are still in "InProgress", a few already for more than one hour - The support case I created about it doesn't show up in my support portal. Direct link to it does work though
- mcphilip 6y agoyeah, I’m seeing event bridge errors and am unable to load cloudwatch log groups. happy short staff day!
- tootie 6y agoI think the issue is that Kinesis is a single point of failure for a ton of systems. When it goes down, loads of other system's workflows can't operate. AWS is famous for eating their own dog food and someone just poisoned it.
- karlkatzke 6y agoMaybe they bought the dogfood from the Amazon Marketplace, but it was counterfeit.
- sk5t 6y agoEventBridge has been struggling for about the past 14 hours as well, which means Cloudwatch Events is not too happy; and, I have the impression CWE underpins a surprising diversity of other things at AWS.
- s_dev 6y agoCan anyone explain why status pages are so difficult. Theres even statups like status.io dedicated to this one thing. It really does seem that anytime there is an outage more often than not the status page is showing all green traffic lights. Making it redundant as a tool to corroborate whats happening. How did AWS status page compare with status.io/aws?
- xyzzy123 6y agoWhen your company gets sufficiently large, outages become political. Failure happens at the speed of computing but agreeing that something is failing in a way that customers need to be told about is a slower process. Even when status pages are fully automatic (rather than manually updated), there will tend to be gaming of the metrics that constitute that. Ideally you would just be monitoring your SLOs and publishing that to customers... that doesn't seem to be how it works, anywhere.
- deleted 6y ago[deleted]
- freehunter 6y agoAnd not just outages, but security incidents. I’ve worked at/with/for many companies as both an employee and a consultant where the top priority wasn’t to have fewer security incidents, but to have fewer security incidents that would require disclosure. Publicly disclosing an incident to a customer is embarrassing and potentially damaging but almost equally as damaging is telling other teams you had an incident. Now anything that goes wrong is your fault by default because “it’s probably related to that incident” and any new security policies are blamed on the other team: “we wouldn’t have to do that if Ops didn’t mess up last month”. The answer to “is this service suffering an outage” is seriously complex and hard to determine. The answer to “is this a security incident” is 10x harder and 100x more political because the industry is still just so wildly immature.
- drchopchop 6y agoAdditionally, you're penalized for doing it "right", because you're often competing against companies which rarely say that anything's wrong (ahem, Mailchimp). You look worse, because you're being transparent about service status, which creates the perception that you're generally less stable.
- zxcvbn4038 6y agoI think we are learning everything that uses AWS Kinesis internally which is cool. It’s always fascinating to learn how AWS works on the backend.
- salil999 6y agoI work at AWS. I can tell you surely enough it's not pretty or easy to work with. Design and architecture are great here but implementation of that is pretty crap...
- sheeshkebab 6y agoWhy use it then? (api is crap, uptime is crap, limits are crap... politics?)
- cheeze 6y agoMoney. Lack of alternatives. Cheaper than GCP. Still less crappy than Azure.
- that_guy_iain 6y agoBecause business authorised it's use. The final say on using AWS doesn't belong to tech but busines and AWS is very good at the sales game. I went to one of their conferences and it was mostly business people and sales pitches.
- zxcvbn4038 6y agoThats too bad, I always imagined the backend was as magical as what AWS users see. I still wish I could have a peek at how S3 works or IAM. Not enough to get a job at AWS - I know they'd fire me the first time I left early for a parent teacher conference or took a sick day, so why put myself in that position.
- orf 6y agoI would be beyond fascinated at how IAM works under the hood.
- pluc 6y agoRule #1 of status pages: never put your status page on the same infrastructure it monitors.
- skavish 6y agoMediaconvert just stopped processing our queues two hours ago, in all our accounts. Anybody else is having it? It's green on the status board.
- camhart 6y ago"This issue has also affected our ability to post updates to the Service Health Dashboard." Last sentence of the alert at the top of the page.
- s_dev 6y agoAlways seems to be the case -- this happened before where the status pages updates were stored in ... S3. It goes beyond coincidence when this happens several times in a row. I think the other explainations sound plausible. There is no technical difficulty here that AWS can't solve -- it's political. Having an outage with a status page makes you liable for your SLAs.
- deleted 6y ago[deleted]
- mikece 6y agoIs this only affecting us-east-1 or other regions as well?
- LennyWhiteJr 6y agoJust us-east-1.
- unilynx 6y agoBut some global services run through us-east-1 - eg Cloudfront is now broken too. So this is also affecting users who don't actually run anything in us-east-1 explicitly (or in the US at all)
- mcintyre1994 6y agoI'm not seeing any issues here yet with S3 images/website buckets stored in eu-west-1 and served by CloudFront. You're right that there's definitely some internal coupling though: > If you want to require HTTPS between viewers and CloudFront, you must change the AWS Region to US East (N. Virginia) in the AWS Certificate Manager console before you request or import a certificate. From https://docs.aws.amazon.com/AmazonCloudFront/latest/DeveloperGuide/cnames-and-https-requirements.html https://docs.aws.amazon.com/AmazonCloudFront/latest/Develope...
- unilynx 6y agoExisting cloudfront is indeed fine. But creating or deleting distributions fails now. (I think it's also pretty rare for an already configured cloudfront to suffer from issues on the control planes. Cloudfront configuration updates are painfully slow even under normal circumstances, and that's probably because the configuration is heavily replicated to all POPs)
- rcardo11 6y ago> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, Lambda, LEX, Managed BlockChain, S3, Sagemaker, and Workspaces. Well, this is a major outgage
- Schweigi 6y agoIndeed, we had the first AWS Kinesis issues already at 13:50 (UTC). Now it's still ongoing after two hours. The status page didn't even update in the first 45 min or so...
- jjoonathan 6y agoThat's typical. The AWS status page is a marketing gimmick whose job is to stay green, not a good faith attempt to assess and report status. If there's an outage, seeing it accurately reflected on the status page is the exception, not the rule.
- simlevesque 6y agoIsn't that fraud ? edit: not sure why my question deserved a downvote...
- WrtCdEvrydy 6y agoIf you're small, yes, if you're AWS, it's business as usual?
- tootie 6y agoAs of this moment, there are more non-green services than I've ever seen. And it's steadily getting worse. EDIT: 15 minutes later and the board is looking worse again.
- chizhik-pyzhik 6y agoUpdating the status dashboard is pretty low priority for operators trying to resolve this issue. It requires escalation up the management chain and careful wording.
- edoceo 6y agoAlso this thread https://news.ycombinator.com/item?id=25209508 https://news.ycombinator.com/item?id=25209508
- bengalister 6y agoSame issue with AWS lambdas I got a: Received malformed response from transform AWS::Serverless-2016-10-31. It is reported now in their service health dashboard.
- Erlangen 6y agoIs this the reason I have seen connection errors in duolingo, > upstream connect error or disconnect/reset before headers. reset reason: overflow
- doseofreality 6y agofriends don’t let friends use us-east-1.
- bithavoc 6y ago"I want to have an AWS region where everything breaks with high frequency..."[0] discussed here [1] [0] https://twitter.com/apgwoz/status/1292519906433306625?s=20 https://twitter.com/apgwoz/status/1292519906433306625?s=20 [1] https://news.ycombinator.com/item?id=24103746 https://news.ycombinator.com/item?id=24103746
- maletor 6y agoIsn't that just called us-east-1?
- riyadparvez 6y agoI've read this multiple times that AWS us-east-1 region is the one that has the highest number of outages. I am eager to hear others' experiences here.
- WrtCdEvrydy 6y agous-east-1 is the zone with highest load and most new services are tested there first. rumor has it, some of the older hardware is moved there and that's why prices are a little cheaper but I have not been able to confirm that.
- chizhik-pyzhik 6y agoNot so much older hardware is moved there as it's just the oldest region with the most baggage
- dodobirdlord 6y agoIt’s not that the oldest hardware is moved there, it’s just that the oldest hardware was there to begin with. There are probably still first-generation EC2 instances running in us-east-1 on their original platforms.
- glenngillen 6y agoPeople are just projecting their own cognitive biases. As Werner has said before everything fails all the time, so you need to design your system/architecture to accept that constant. US-east-1 is by far the largest of the regions, and at that scale you can probably assume that at any given point in time there is hardware in there failing that needs to be physically replaced. As a result it's the region most well equipped to tolerate that level of constant failure (it's got 6 AZs!). It's also the the most popular of the regions, is typically one of the launch regions for new services, and runs a bunch of critical Amazon infra too. If anything it holds a special place in terms of importance for AWS to keep it up because the impact of a widespread problem here is amplified. For the same reason though any problem here is much more visible across the entire internet. Which is why the handful of outages are so memorable to people.
- tootie 6y agoAny tips on how to collect on SLA credits from this?
- runlevel1 6y agoThe procedure is outlined in each service's SLA -- though I think they're all pretty much the same. Annoyingly, they expect you to do the leg work to show when the outage happened and supply logs demonstrating that you were impacted. Might want to do some napkin math first to see if the amount credit is worth your time. The couple times my org considered pursuing it, it just wasn't worth the effort. (Though, personally, I think that speaks to a larger problem with the SLA.) Credit Request Procedure in Kinesis SLA: https://aws.amazon.com/kinesis/sla/#Credit_Request_and_Payment_Procedures https://aws.amazon.com/kinesis/sla/#Credit_Request_and_Payme...
- troelsSteegin 6y agoWould this explain the washingtonpost.com outage? That site has been displaying a "Welcome to OpenResty!" page for the past 20 minutes or so. EDIT: nevermind, the Post is back, and Kinesis is still erroring.
- leothekim 6y ago> This issue has also affected our ability to post updates to the Service Health Dashboard. This is when you fall back to the Tumblr blog for status updates. <rimshot>
- Nexeo 6y agoIs anyone else getting "Capacity unavailable" when trying to add tasks in Fargate?
- vishesh92 6y ago> "This issue has also affected our ability to post updates to the Service Health Dashboard." This is why I prefer 3rd party monitoring systems to track health of my internal monitoring systems.
- totaldude87 6y agoi tried reaching out to amazon support, apparently they are also seeing issues internally and there is a high possibility that these two are related.. Their ETA, 2 hours, and then try contacting again!
- btown 6y agoAs of 2020-11-25T17:21Z this is also causing a Heroku outage preventing new spin-ups, which presumably uses these APIs to verify instance health. https://status.heroku.com/ https://status.heroku.com/
- jadbox 6y agoHaving Lambda issues too
- drfritznunkie 6y agoCognito is one of the most frustrating AWS services I have to work with, it is almost, but not quite, entirely unlike an SP. We're using it to federate customer IDPs through user pools, but this ends up with customer configs being region specific. Has anyone figured out how to set up Cognito in multiple regions without the hijinx of having the customer setup trusts for each region? Not to mention, while multiple trusts are I think possible with ADFS (not that I've tested it), I'm pretty sure that Okta doesn't support multiple trusts, so regardless of how many regions, we'd still be SOL there...
- sk5t 6y agoEh? Brokering amongst multiple trusts (and managing protocol transition) is almost the raison d'etre for lifting token issuance out of your app and into ADFS, Okta, Auth0, etc. Of course you'll have to deal with home realm discovery--really need to go in with open eyes on that one.
- drfritznunkie 6y agoYes, but cognito endpoints and pools ids are regional and globally unique, and there is no way that I know of to setup duplicate userpools in multiple regions and have requests served by either region. That means the customer IDP side would need to have two different SAML apps configured for each region...
- sk5t 6y agoAh, I see what you mean. It does seem like you'd want a more complex arrangement of trusts to keep things simple on the leaves; or else avoid using a product that requires generating a hundred scattered security authorities.
- myleshenderson 6y agoThis was shared with me today: https://medium.com/@nealrp/aws-cross-region-cognito-replication-c764da1f29c0 https://medium.com/@nealrp/aws-cross-region-cognito-replicat...
- 0xmohit 6y agoAlmost 9 years have passed by and nothing has changed. The dashboards continue to remain green. https://news.ycombinator.com/item?id=3707590 https://news.ycombinator.com/item?id=3707590
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]
- oliverfriedmann 6y agoEC2 is also having issues as status cannot be reported properly which makes auto-scale go nuts
- swasheck 6y agoAh yes. It's the annual AWS Thanksgiving Holiday major us-east outage.
- Steve886 6y agoMany applications – Including Anchor, Adobe Spark, Flickr, SiriusXM and Roku reported disruption caused by this outage. https://news.alphastreet.com/huge-aws-outage-affects-a-wide-range-of-applications/ https://news.alphastreet.com/huge-aws-outage-affects-a-wide-...
- sushikokk 6y agoMy iRobot (roomba vacuum robot) app not working for 4 hours...
- shripadk 6y agoPaddle checkout is down as well (connected to this outage): https://twitter.com/PaddleHQ/status/1331659286649466881 https://twitter.com/PaddleHQ/status/1331659286649466881
- piewzko 6y agoNow is probably a good time to plug some of the open source alternatives to vendor locked in identity solutions: - https://github.com/ory https://github.com/ory - https://github.com/dexidp/dex https://github.com/dexidp/dex - https://github.com/authelia/authelia https://github.com/authelia/authelia - https://github.com/keycloak/keycloak https://github.com/keycloak/keycloak - https://www.gluu.org/ https://www.gluu.org/ - https://github.com/accounts-js/accounts https://github.com/accounts-js/accounts
- agustif 6y agoadd AccountsJS, a small nice modular typescript/js lib for building account systems easily
- piewzko 6y agoDid not know about that one, I added it to the list!
- lukevp 6y agoFusionauth is pretty cool. I’ve worked with the team a bit on the .net core support.
- simlevesque 6y ago- https://www.etebase.com/ https://www.etebase.com/
- kevindong 6y agoI'd expect Amazon to be better able to maintain uptime than a self-hosted option at most (but not all) companies.
- rhizome 6y agoAmazon can't diversify their providers, though. Regular Joes like us can use AWS, GCE, on premises, some non-reseller colocation provider, etc., and create failover duplicates, alternative deploy targets, or simply not ever have a complete outage due to the unlikelihood of all of these things failing at once.
- deleted 6y ago[deleted]
- agustif 6y agoWe've been having lots of issues with Vercel today, since it uses AWS under the hood I'm guessing that's related...
- rmujica 6y agoIt is now affecting ECS and EKS. Having problems scaling own nodes.
- revicon 6y agoI’m trying to find a doc on running cognito using multiple zones and I can’t find much. Anyone have a multi-az cognito deployment running right now?
- x86_64Ubuntu 6y agoZones or regions? I don't think multi-az will help you when the whole region kicks the bucket.
- LeoTM 6y agoI'm in the UK and this now may have cascaded onto VISA https://downdetector.co.uk/status/visa/map/ https://downdetector.co.uk/status/visa/map/ I am unable to order my Papa Johns pizza https://imgur.com/u5QSszv https://imgur.com/u5QSszv
- jm547ster 6y agoTopped with?
- zedpm 6y agoEvery time I check the Personal Health Dashboard, the number of issues increases; it's currently showing 13 open issues for my account. Cloudwatch logs for the last few hours are unavailable; it appears that the log agent is getting errors when it attempts to upload log events. Metrics are spotty or missing.
- symlinkk 6y agoTitle should be changed, this is a widespread AWS issue, it’s not specific to Cognito.
- outworlder 6y agoThey should rename that region to us-chaostesting-1 . Problem solved.
- PragmaticPulp 6y agoWe hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some digging, he told us that admitting your services were down was considered a death sentence for your job at his previous team at Amazon. He was so scarred from the experience that he refused to ever take responsibility for outages. Ultimately, we had to put someone else in charge of updating the status page because he just couldn't be trusted. FWIW, I have other friends who work on different teams at Amazon who have not had such bad experiences.
- Ensorceled 6y ago"Shooting the messenger" is so common that we, well, have a phrase for it.
- geertj 6y agoThere is a phrase for it, but it does not match my experience at AWS at all. (Source: been working there for 3.5 years now). Things break, we do COEs and we learn from them. If an issue was caused by operator error, the COE would look at what missing or broken processes caused the operator to be able to make this error in the first place.
- Ensorceled 6y agoAnd yet we continue to have “all greens” on the status boards during outages.
- Aperocky 6y agoThat's opposite of my experience at AWS. It's likely that the culture at AWS has changed over the past few years, it's also likely that there's a difference in culture between teams.
- kristianpaul 6y agoCloudWatch is definitely one of those "AWS Primitives" services that side effects others when having problems, something similar happened with DynamoDB some years ago.
- dodobirdlord 6y agoIn early 2019 CloudWatch had a major outage that was particularly nasty, where instead of just outright failing to report metrics it reported a small percentage of metrics. As a result a lot of autoscaling groups and DynamoDB tables that were theoretically supposed to avoid scaling in during a metric outage still scaled in, because they saw it as a 90+% traffic reduction rather than a metric outage.
- carusooneliner 6y agoWe are seeing an elevated rate of failures on our service, which depends on AWS Cognito. Tweeted an update on it: https://twitter.com/outklip/status/1331705524396625924 https://twitter.com/outklip/status/1331705524396625924
- driverdan 6y agoFive hours later and nothing has changed. For a company like Amazon this should be unacceptable. Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.
- Bombthecat 6y agoI think the deeper problem is the interconnectivity between services and their apis. It's too complicated to maintain...
- ben509 6y agoYeah, the cascading failure of all the other services is a deep architectural issue. Having lots of services that do one thing and one thing well makes a lot of sense. Breaking them out into separate components brings a level of visibility into the system. And it's AWS's whole business model. But it does mean that, fundamentally, service X is available when and only when (WAOW?) services A, B, C, etc. are all available. Its uptime is no greater than min(uptime(A), uptime(B), etc) I'm trying to rework the authentication for our application and integrate it with our parent company's systems. As we talk to other teams, I see all these architecture diagrams where the solution to every problem is Yet Another Service, to where you're running a real rube goldberg machine.
- adewinter 6y agoInterested to know what the alternative might be and why it would mean better uptime
- aledalgrande 6y agoThe alternative for AWS might be to be able to failover to e.g. Kinesis in another region.
- one2know 6y ago
- throwaway343432 6y agoLarge-scale events (LSEs) are becoming more and more common. It'll keep getting worse. AWS has to take a hard look at how they build their software. Their bad engineering practices will eventually catch up to them. You can't treat AWS the same as Alexa. Sometimes it's smarter to take your time to ship stuff instead of putting it out there. Burning out your oncall engineers is not a feasible long-term plan. AWS will be in deep trouble when/if GCE fixes their customer support.
- ipsocannibal 6y ago"Large-scale events (LSEs) are becoming more and more common." Stats on this? You seem to have insight on AWS's engineering practices. From your point of view what should be changed?
- xyst 6y agojust looking at this dashboard, I never realized how many services aws has to offer. I’d hate to be the “aws” guy
- mcintyre1994 6y ago> 2:43 PM PST Between 5:15 AM and 2:28 PM PST customers experienced increased API failure rates for Cognito User Pools and Identity Pools in the US-EAST-1 Region. This was due to an issue with Kinesis Data Streams. We have implemented a mitigation to this issue. Cognito is now operating normally. Seems like they fixed Cognito while Kinesis and many other services are still broken - presumably somehow removing the dependency on Kinesis? It’ll be really interesting if their post mortem explains this mitigation.
- TedShiller 6y agoAWS Status website is down for me. Is there a status website for AWS Status?
- astatine 6y agoThe way ahead should be an independent entity who audits systems and has responsibility to certify that the dashboard represents a true and accurate view of the actual status. Like is done so effectively with company financials. Oh, wait! EY, PWC, and who can forget Arthur Andersen! But, naturally, technology people can solve this better than anyone else, right?
- ssss11 6y agoThere is no problem here. jedi mind trick hand wave
- MR4D 6y agoAnd that confirms it for me: Amazon is officially a Day 2 Company. Happened faster than I thought, but based on reading the comments about people who work(ed) there, this seems cut and dried to me.