9 ms·
AWS us-west-2 issues
AWS appears to be having lambda and API Gateway issues again in us-west-2. Symptoms on our end look similar to the August 24th partial outage.
- hazrmard 4y agoMy job application at amazon.jobs today got interrupted by this, I think. I hope they don't get a malformed application :/
- neuralspark 4y agoMaybe it's trying to save you
- shahbaby 4y agoThis is divine intervention.
- hbn 4y agoThis is part of your test. You're failing!
- mrweasel 4y agoMaybe you broke it?
- dylan604 4y agoThat resume that was submit contained a known PDF delivered attack. Guess their AV subscription ran out on that server?!
- thallium205 4y agoHere here!
- deleted 4y ago[deleted]
- restlake 4y agoYep same symptoms for us as the August outage and support confirmed an outage for us as well
- jvolkman 4y agoWould be nice if they could update their status dashboard.
- thallium205 4y agoIt is. It's just currently set at "Informational" severity.
- sarcasticadmin 4y agoThere was a decent delay in the alerting from AWS. I was just investigating this degradation in service for some of my systems for about 1 hour before seeing this alert raised.
- atomon 4y agoWe're having issues as well
- adamkittelson 4y agoCoincidentally that's two months in a row of a us-west-2 outage on the 4th Wednesday of the month. See you all back here on Oct 26?
- kwindla 4y ago:-) August 24th was the first time we saw exactly this issue in 7 years of heavy, multi-region, AWS use. So we put in place the ability to semi-automatically route around this more quickly, but we didn't fixate on it. Two data points is a line, however. (But maybe not yet a trend?)
- qbasic_forever 4y agoMonthly certificate updates gone wrong perhaps? Oops
- zonkd1234 4y agoLambda is returning 502 errors in US-West for me as well.
- camhart 4y agoSame here
- kwindla 4y agohttps://health.aws.amazon.com/health/status https://health.aws.amazon.com/health/status is now showing the issue. Our first canaries fired for this at 9:43 PDT local time. The status page updated at 10:13.
- deleted 4y ago[deleted]
- matthoiland 4y agoOur first canaries at 9:21 PDT
- kwindla 4y agoThat's really interesting. We didn't get any user reports before our canaries fired, either. Now I have to think about what might explain the difference between your systems and ours. We're monitoring API Gateway health (more or less), because that's what we care about in this part of our infrastructure.
- dylan604 4y agohttps://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/aws-amazon-web-services/ My Roomba isn't working!!! hahahaha!
- chasd00 4y agoi wonder if any smart locks are down, people being locked in or out of their AirBNBs would be pretty funny.
- dylan604 4y agoof all things to connect to a smart home, the locks to my house would be the absolute bottom of the list. i hate coming home when the power is out and cannot open the garage door. i couldn't imagine not being able to get in at all. as far as not getting out, any lock unable to be unlocked from inside seems like something should not be allowed to be made. ever.
- readams 4y agoI think consumer smart locks still let you use the regular lock if the power is out.
- hatware 4y agoHow long do you think batteries last?
- dylan604 4y agoyou ask that like you are challenging my idea of not ever using smart locks on my home. instead, you're bringing up another reason to support that decision. so, which direction were you attempting to move the needle?
- hatware 4y agoBatteries going out have never locked me out of my home. Seems like you're putting way too much thought into something that probably won't happen, being 2022 with notifications and all. I don't justify my choices with 0.1% chances.
- ggregoire 4y agoBunch of Database error: [Amazon](500310) Invalid operation: No response body. Database error: [Amazon](500310) Invalid operation: curlCode: 28, Timeout was reached in our logs during the last hour
- mchusma 4y agoI'm seeing multiple services as down: ECS, Medialive, etc.
- ksimukka 4y agoOh no!
- deleted 4y ago[deleted]
- willio58 4y agoIt's crazy to me how long it takes Amazon to notify users that there is even an issue. I think it took 15 minutes for them to acknowledge an issue, thats a lot of time for our services.
- socialismisok 4y agoThat's because it's a political issue inside AWS. They have the technology to report it automatically, but there's a strong pressure not to post "green-i" or yellow or red, because those things impact SLA payments. So if there's any way they can spin it as not an outage they will try not to post it.
- dilyevsky 4y agoI participated in multiple sla penalty payments requests on the customer side and never once we even discussed status page color
- socialismisok 4y agoI participated in many status page color discussions and we always discussed sla. Shrug
- dilyevsky 4y agoYou mean on the AWS/service provider side? I've also participated in internal incident response on the service provider side and we again made SLA refund decisions based on actual customer impact not our status page. But then again we diligently update our status pages so there's that.
- socialismisok 4y agoYeah, on the AWS side.
- hatware 4y ago
- AaronM 4y ago[10:33 AM PDT] [10:33 AM PDT] We are investigating increased error rates for invokes in the US-WEST-2 Region. We do not yet have a root cause, but are investigating multiple potential root causes in parallel. In addition, we are implementing filters on inbound traffic from a set of sources with recent significant traffic shifts, which may help mitigate the impact. We do not yet have a solid ETA, but will continue to provide updates as we progress. [10:13 AM PDT] [10:10 AM PDT] We are investigating increased error rates for invokes in the US-WEST-2 Region. AWS AppSync AWS Batch AWS Certificate Manager AWS Cloud9 AWS CloudShell AWS Device Farm AWS Global Accelerator AWS Greengrass AWS IoT 1-Click AWS IoT Device Management AWS Lambda AWS Proton AWS Resource Access Manager AWS RoboMaker AWS Service Catalog Amazon AppStream 2.0 Amazon CloudWatch Amazon Connect Amazon Elastic Compute Cloud Amazon Elastic Container Registry Amazon Elastic MapReduce Amazon EventBridge Amazon FinSpace Amazon Kendra Amazon Lightsail Amazon Location Service Amazon Managed Workflows for Apache Airflow Amazon Nimble Studio Amazon Pinpoint Amazon SageMaker Amazon WorkSpaces EC2 Image Builder
- biermic 4y agoRedshift COPY from s3 is also affected. ----------------------------------------------- error: No response body. code: 30000 context: query: 1154965 location: xen_aws_credentials_mgr.cpp:403 process: padbmaster [pid=15893] -----------------------------------------------
- humbleharbinger 4y agoS3-megalodon on call checking in
- austinpena 4y agoSeeing API gateway issues here: https://www.taloflow.ai/is-aws-down/us-west-2 https://www.taloflow.ai/is-aws-down/us-west-2
- jmartens 4y agoMust be a specific AZ, not seeing any issues here EDIT: I see now that it's an API Gateway issue, so that's why my team isn't impacted.
- sdfdsfsd 4y ago
- CacheRules 4y agoSuccess!
- whalesalad 4y agomy least favorite region tbh
- Victerius 4y agous-east-1's revenge.
- enapupe 4y agoPeople, I've just started setting up API GW and deploying my code to another region. Rest assured the outage will normalize itself before I finish migrating. ETA: 15min-ish
- enapupe 4y agoDid I say 15 minutes? Actually, it depends.
- dylan604 4y agoThe first rule of publishing ETAs for updates is to triple the estimate. Be the hero that finishes before the published time vs finishing after! The second rule of ETAs is don't specifiy a time. I'm still amazed at the number of cults that place a specific date on the return of the savior, and even more by the people that go along with the rescheduling. The fact that I'm still waiting in Dallas for JFK's return is irritating. /s
- ramesh31 4y agoThe moment I get a whiff of AWS issues, that's basically a wrap on the day. I'm not going to spend my time wondering why something isn't working today because somewhere down the chain it is undoubtedly a silently failing AWS service driving me mad. The whole house of cards comes tumbling down and things stop working that Amazon will swear up down left and right are 100% healthy. Sadly this has become a monthly occurrence at this point. Monoculture, it turns out, is a really really bad idea.
- 0xbadcafebee 4y agoThis is a good point: we should come up with a list of what services use what cloud providers, so not all our eggs are in one basket.
- balls187 4y agoCloud9 is also down. Joy. Edit to add: I can't seem to access my NHL Season Tickets via Ticketmaster either...
- makestuff 4y agoOne would think all of those ridiculous ticketmaster fees would make them be able to afford some sort of multi region/multi cloud setup...
- fasteddie31003 4y agoI'm working on an app that tracks downtimes. I put some of my latest data up here: https://app.awareops.com/whatisdown/orgs/a7f95108-ead0-4718-8940-d61e0d80510a https://app.awareops.com/whatisdown/orgs/a7f95108-ead0-4718-...
- jvolkman 4y agoFrom the latest update: > While we have seen improvements in error rates since 10:40 AM PDT, recovery has stalled and we do not have a clear ETA on full recovery. For customers that have dependencies on API Gateway and are experiencing error rates, we do not have any mitigations to recommend to address the issue on the customer side.
- systemvoltage 4y agoHonest question and frankly scared to ask because it sounds stupid: if you have like 30 mins of downtime on AWS every year and spend 3x cost on managing their infrastructure risk by deploying it on multi-AZ and multi-region (and thereby AWS pushing reliability management back to the customer); is the value proposition of cloud just some dude to install a disk if it goes out on a rack in your office? May be there is a reverse incentive for AWS to keep their AZ's slightly unreliable so that customers spend 3x or 9x or what have you to make sure nothing ever goes down. Like what's wrong with on-prem? Lack of diesel generators? We could just have that without AWS. Bare metal datacenter. Counter to most opinions, I think managing a server isn't that difficult. I am sort of a semi-professional and prosumer that has no trouble managing servers for years on end with less downtime than a whole fricking datacenter. There is more serious discussion and new revelations around this [1], [2]. Sometimes it is hard to ask questions about layers of abstractions that have built up and no one dares to think about getting rid of them. [1] https://www.economist.com/business/2021/07/03/do-the-costs-of-the-cloud-outweigh-the-benefits https://www.economist.com/business/2021/07/03/do-the-costs-o... [2] https://oxide.computer/ https://oxide.computer/
- rrdharan 4y agoSecurity, Compliance, and Geographic footprint. If you're a large multinational, you basically face the same threats as Google/AWS/MSFT but there's no way you can hire, train and keep as good a production security team as them (well maybe better than Azure, but I digress). You can't afford the upfront contractual / capital costs to maintain datacenters in every region. And finally you can't afford the armies of lawyers and compliance engineering teams to try and reason about your data residency and things like GDPR and CCPA. In other words, you're mostly paying for production security / privacy incident response, compliance (lawyers) and datacenters.
- mannyv 4y agoHaving managed a small data center in the past and having seen what it takes to manage multiple enterprise-scale data centers, the answer is "no, AWS does it better." The company that I'm in right now has two engineers (including myself) who are building and maintaining a product that serves millions of streams a week. There's no fucking way we could have done this ourselves. One F5 would cost more than our entire total AWS bill for two years - and we'd have to have at least 4 F5s if we wanted to try to match AWS. Plus the media encoders would cost a fortune. For some things it's fine to head over to lowendbox.com and pick up a cheap VPS hosting package. We could theoretically build our stack on top of a bunch of VPSs, sync everything with rsync, etc. But then we'd be spending time building infrastructure (which is pretty much valueless) instead of our product.
- pneumatic1 4y agowhat happened to stop.lying.cloud?
- enapupe 4y agoCould any of you clarify to me why do we have "Edge" domain rather than a specific region?
- davidjfelix 4y agoYeah, you can select "edge" for some resources (API Gateway and Lambda are two that come to mind) which just means it's located in all regions plus some additional "edge" infrastructure that isn't available as a region. AWS puts some restrictions on edge resources since there isn't enough capacity or full-region functionality. Usually you pick this for stuff that is CDN oriented and front a regional service with that.
- enapupe 4y agoExactly, why did it failed yesterday then? If supposedly just us-west-2 was down :(
- davidjfelix 4y agoMy guess is if you saw an outage with edge, it was because you were being routed by us-west-2
- kolanos 4y ago3h 52m and counting
- Aaronstotle 4y agoNow I know why cloud shell wasn't working around noon PDT
- IceWreck 4y agoHey atleast it isnt us-east-2 this time