16 ms·
AWS US East is experiencing high error rates on several services
The services CloudWatch, SES, SNS, SQS, SWS, AutoScale, Cloud Formation, Directory Service, Key Mgmt and Lambda experience very high error rates for about 3 hours now.
Dynamo DB is throttling API access and seems to be having issues with the management of meta data.
- brianpetro_ 11y agoCould be why my address is now "incorrect" https://twitter.com/search?f=tweets&vertical=default&q=incorrect%20address%20Amazon&src=typd https://twitter.com/search?f=tweets&vertical=default&q=incor...
- colinbartlett 11y agoWell at least it's nice to know Amazon uses Amazon. Could still be unrelated, but awfully coincidental.
- airza 11y agowell, time to find out who has failure tolerance built in to the architecture 8^)
- drendorx39 11y agofailure tolerance is an alien technology for amazon...
- divideby0 11y agothat's completely untrue. there are many ways to do fault-tolerance in AWS. it's expensive, but it's possible. netflix even goes as far as simulating the failure of entire aws regions in their simian army testing suite: http://techblog.netflix.com/2011/07/netflix-simian-army.html http://techblog.netflix.com/2011/07/netflix-simian-army.html That's why Netflix stays up when us-east or us-west are down.
- beagledude 11y agoNetflix is down.
- veverkap 11y agoWorks fine for me
- kordless 11y agoIt is true. Amazon offloads a decent amount of fault tolerance to the application provider, as you point out here. I will also mention that Netflix does not solely rely on Amazon for running their services. They run their own decentralized caching layer: https://openconnect.netflix.com/deliveryOptions/ https://openconnect.netflix.com/deliveryOptions/
- crb 11y ago"4:52 AM PDT We want to give you more information about what is happening. The root cause began with a portion of our metadata service within DynamoDB. This is an internal sub-service which manages table and partition information. Our recovery efforts are now focused on restoring metadata operations. We will be throttling APIs as we work on recovery." (http://status.aws.amazon.com/ http://status.aws.amazon.com/)
- deleted 11y ago[deleted]
- taf2 11y agoThe main issue appears to be DynamoDB Here's a copy from the status page. 3:00 AM PDT We are investigating increased error rates for API requests in the US-EAST-1 Region. 3:26 AM PDT We are continuing to see increased error rates for all API calls in DynamoDB in US-East-1. We are actively working on resolving the issue. 4:05 AM PDT We have identified the source of the issue. We are working on the recovery. 4:41 AM PDT We continue to work towards recovery of the issue causing increased error rates for the DynamoDB APIs in the US-EAST-1 Region. 4:52 AM PDT We want to give you more information about what is happening. The root cause began with a portion of our metadata service within DynamoDB. This is an internal sub-service which manages table and partition information. Our recovery efforts are now focused on restoring metadata operations. We will be throttling APIs as we work on recovery. 5:22 AM PDT We can confirm that we have now throttled APIs as we continue to work on recovery. 5:42 AM PDT We are seeing increasing stability in the metadata service and continue to work towards a point where we can begin removing throttles.
- virtuallynathan 11y ago6:19 AM PDT The metadata service is now stable and we are actively working on removing throttles. 7:12 AM PDT We continue to work on removing throttles and restoring API availability but are proceeding cautiously. 7:22 AM PDT We are continuing to remove throttles and enable traffic progressively. 7:40 AM PDT We continue to remove throttles and are starting to see recovery. 7:50 AM PDT We continue to see recovery of read and write operations and continue to work on restoring all other operations. 8:16 AM PDT We are seeing significant recovery of read and write operations and continue to work on restoring all other operations.
- deleted 11y ago[deleted]
- deleted 11y ago[deleted]
- justinholmes 11y ago9:12 AM PDT Between 2:13 AM and 8:15 AM PDT we experienced increased error rates for API requests in the US-EAST-1 Region. The issue has been resolved and the service is operating normally.
- colinbartlett 11y agoThis is manifesting itself as downtime for a lot of companies, including Heroku: https://status.heroku.com https://status.heroku.com If you want alerts on this sort of thing, my side project StatusGator https://statusgator.io https://statusgator.io will alert you when services post downtime on their status pages. My dashboard blew up this morning with a ton of red and yellow as soon as Amazon started flaking. Edit: I suppose it's time to invest in a multi-region setup. Since StatusGator is hosted on Heroku in the US-East region, it is in theory affected by this problem though so far is still up.
- fidz 11y agoFrom Heroku Status Page: > Our service provider is still working towards resolution of this issue. We will update when we have news, or in 1 hour. I wonder why they don't tell that AWS is their service provider. Is it wrong to make the information less obscure?
- ukd1 11y agoYes, as it's shifting the blame away from their choice - which was to use AWS.
- jonaf 11y agoI expect the reason is because Heroku could switch to a new provider in the future and it would be a pain to always update every reference to their provider. Parent seems to imply there's something wrong with choosing AWS. There is not. (forgive me if I mistook the tone)
- ukd1 11y agoYou did; there is nothing wrong - but it's a choice. I.e. your site being down is your fault ultimately, not AWS's - you choose AWS (instead of many other choices OR making something with multiple, etc). Blaming it upstream is hiding passing the buck on your decision.
- nasalgoat 11y agoThis seems like another reason to not rely on Amazon-specific services, other than the obvious vendor lock-in. At least in the event of an instance outage you could conceivably migrate off Amazon to another VPS provider. No one using DynamoDB has an alternative.
- jroid 11y agoYou still depend on some provider. Or, are you talking about multi cloud installations ?
- alexbilbie 11y agoDynamoDB now support cross-region replication [0] so you can build more resilient applications with it [0] http://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.CrossRegionRepl.html http://docs.aws.amazon.com/amazondynamodb/latest/developergu...
- andrewchilds 11y agoI don't think cross-region replication would've helped in this case: "The replica tables are intended to serve as read-only copies of the data; however, it is possible to write data to a replica table. If you write data to a replica, those changes will not be propagated to the master, or to any other replicas."
- abalone 11y agoYou could at least fall back to a read-only mode. That could be very helpful, compared to going down completely.
- dialtone 11y agoYou can relatively trivially build multi-master cross-region replication in DynamoDB by using kinesis and writing to kinesis instead of DynamoDB directly. On the consuming end of Kinesis you then fan out to all the DynamoDB (or whatever other database you want to use) regions[1]. Admittedly this only works with some relatively relaxed constraints on the latency you can see, intra region latency can go up to 1 second although rarely, while cross-region is around 3 seconds. An important role is also played by the structure of your objects and how accepting they are of concurrent updates coming from different regions (which is the main reason why the default replication in DynamoDB is not multi-master). [1]: http://tech.adroll.com/blog/data/2015/06/26/kinesis.html http://tech.adroll.com/blog/data/2015/06/26/kinesis.html
- drendorx39 11y agoDynamoDB is literally a garbage. That's why Amazon does not provide any SLA for the service... Even cheap Azure Storage provides cross-region failover.
- alexbilbie 11y agoDynamoDB now support cross-region replication [0] so you can build more resilient applications with it [0] http://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.CrossRegionRepl.html http://docs.aws.amazon.com/amazondynamodb/latest/developergu...
- Ixiaus 11y agoGarbage you say? Dynamodb blazed a trail for many open source, eventually consistent kv databases. Certainly not garbage.
- drendorx39 11y agoThe only thing DynamoDB can do good is simplicity. Except for this, even MongoDB has tons more features than DynamoDB and the new version resolved performance problems existed in the previous versions.
- icefall 11y agoApples and oranges. People who call tools like this garbage fail to evaluate trade-offs at their required complexity space. Distributed systems have extreme trade-offs. More features => more bugs.
- istvan__ 11y agoMongoDB really? You are comparing a software product to a multi-region global data service? This is just not a great comparison. If you build a distributed global data service on MongoDB you could compare that to DynamoDB.
- halayli 11y agoI have a feeling you have no clue what you are talking about, nor read dynamo db paper or tried to use it in a production environment.
- mdnormy 11y agoI hate the fact that most people(including me apparently) still assume AWS is not in their "downtime" equation. Just spend the last 30min troubleshooting SMTP auth problem. Not funny when its Sunday.
- antaviana 11y agoI also noticed the random auth problem in SMTP and my knee jerk reaction was to google "status aws"...
- interesting_att 11y agoYou're not alone buddy. Been wasting hours of my life looking at this stuff too :)
- bsdpython 11y agoBetween AWS, Google, Apple, Facebook, Twitter and like I doubt I am alone in spending a ton of my time working around their various issues to run my tech stack.
- janson0 11y agoSQS is the specific service giving me a ton of trouble right now. Hope they resolve this quickly. Had rayguns about sqs all night heh. So are they saying they are throttling SQS because of the DynamoDB issue?
- oliverfriedmann 11y agoI'm not sure. I think many of the other services mentioned probably rely internally on SQS, so resolving the SQS issues might resolve most of the other issues as well. Not completely sure though whether DynamoDB would benefit from relying internally on SQS.
- janson0 11y agoYeah good point. It's something I forget sometimes that AWS uses AWS... and that even if I don't rely on a particular service specifically, a service I rely on may, in fact, rely on that service. Hopefully there is a relatively fast recovery on this. Can anyone even log into their aws console right now?
- grhmc 11y agoI can log in. I wonder if SQS uses DynamoDB, not the other way around.
- janson0 11y agoHere's the SQS Error Log right now: 3:14 AM PDT We are investigating increased error rates in the US-EAST-1 Region. 4:06 AM PDT We can confirm increased error rates for CreateQueue, SendMessage and ReceiveMessage API calls in the US-EAST-1 Region and continue to work towards resolution. 5:07 AM PDT We can confirm increased error rates for CreateQueue, SendMessage and ReceiveMessage API calls in the US-EAST-1 Region. As we work towards recovery, error rates may temporarily increase. 6:06 AM PDT We can confirm significantly increased error rates for CreateQueue, SendMessage and ReceiveMessage API calls in the US-EAST-1 Region. As we work towards recovery, error rates may temporarily increase in error rates.
- 11y ago
- drendorx39 11y agoAmazon CTO: We designed DynamoDB to operate with at least 99.999% availability :D
- ryanfitz 11y agoI've been using DynamoDB since it was released. In over 3 and 1/2 years of use, this is the first time I've experienced DynamoDB being down.
- awscat 11y agoIf down for longer than 18 minutes then they missed "5 9s" availability (.3 hours / 3.5 years). Not that it is supposed to work that way.
- yeukhon 11y agoSoftware bug caused downtime vs infrastructure / hardware availability uptime to me are a different guarantee. I am pretty sure someone did something recently to DynamoDB.
- toomuchtodo 11y agoInfrastructure guy here doing this for 14 years. Downtime is downtime. You get a pass if its "scheduled maintenance" you've notified your customers about to allow them to be prepared, but if you silently perform maintenance and it goes to shit, you've just counted against your metrics.
- yeukhon 11y agoNope. I still disagree. No service can guarantee 99.999999% unless you discount software upgrade. You just cannot. If you think those nines include software upgrades, you are probably over optimistic.
- 11y ago
- dankohn1 11y agoI noticed this because I was unable to checkout on Amazon Prime Now just now.
- pgrote 11y agoI noticed it when I couldn't stream something. The player begged it off as Silverlight issue even when the flash option is chosen. lol
- JOnAgain 11y agoLe sigh. This is impacting AirBnB and I need to check in somewhere in LA later today. Good thing all the details are in the AirBnB messaging history with the host. Time for them to just go back to email.
- neals 11y agoI'm somewhat in the same situation. I want to check-in a movie I just watched and give it appropriate rating, but IMDB is down. I guess we'll just have wait, right?
- clebio 11y agoAny knowledge or evidence that IMDB runs on AWS, and that the two are thus correlated?
- colechristensen 11y agoFor starters, Amazon owns IMDb.
- deleted 11y ago[deleted]
- nhumrich 11y agoWell, Amazon owns IMDB, so it's probably a reasonable assumption.
- jspaetzel 11y agoQuick lookup says yes, # host imdb.com imdb.com has address 207.171.166.22 imdb.com has address 72.21.210.29 imdb.com has address 72.21.206.80 # host 207.171.166.22 22.166.171.207.in-addr.arpa domain name pointer 166-22.amazon.com.
- drendorx39 11y agoYep. AirBnb is not working
- 11y ago
- Hughlon 11y agoVideos and Alexa ia also down
- eatonphil 11y agoThis appears to have completely blown away all Alexa data. Even searching for google.com returns nothing.
- vreauobere 11y agoother down sites: medium.com, getpocket.com, idonethis.com
- vreauobere 11y agoother down sites: medium.com, getpocket.com, idonethis.com
- geertj 11y agoAddress verification on amazon.com doesn't work for me at the moment, blocking me from making any orders. Not sure if this is related.
- SoulMan 11y agoNothing should be effected in non- USEast regions as per the status page
- crablar 11y ago?
- kureikain 11y agoWhoever uses autoscaling and, especially lifecycle notification with SQS will be in trouble now(I'am). The morning is going to be started. Traffic will be ramped up, and not sure if new sevrers will be launched because CloudWatch is failed. Polling SQS to find lifecycle notification message fail too.
- tcas 11y agoWe were in the middle of a large infrastructure change starting at 4:30am this morning, including taking our application offline. I'm very thankful that we did dry runs along with timing how long certain operations like RDS restores should take and planned for abort steps in case something goes wrong. We noticed that RDS and ElastiCache backup and restores were taking much longer than expected, and once the first set of errors about Dynamo DB came in we decided to abort and try it again at a different date. An hour later we got notifications that RDS was having issues as well. I'm disappointed that it takes so long to update the AWS status page when things aren't working properly.
- confiq 11y agosimilar story here... They status.aws page has serious delays
- tnolet 11y agodocker, wercker, travis.ci are also affected. Can't login or stuff is really sluggish.
- crypt1d 11y agoAudible seems affected by this as well. I've made some purchases with my credits but the books are still not showing up in my library...and the checking out process is very slow.
- edanm 11y agoFYI happened to me as well, but is resolved now.
- sauere 11y agoTinder is down due to this, now my life is pointless.
- divideby0 11y agoSign-ins to AWS console also appear to be timing out: https://www.evernote.com/l/ABkKLgp3RjRDe5uV4pMlyVg1uzkW41DG4SEB/image.png https://www.evernote.com/l/ABkKLgp3RjRDe5uV4pMlyVg1uzkW41DG4...
- lxfontes 11y agoIf you can't get in the console, use awscli. It is responding fine!
- neoecos 11y agoThe AWS KMS is not working. Critital payment applicaction down =S.
- rational-future 11y agoWhy would you run a "critical payment application" in US East? This datacenter has 10x the downtime of West or Ireland.
- rsynnott 11y agoNot often you see a red status symbol on Amazon's status page (yellow is normally considered more than enough to indicate that the product is totally broken). Don't think I've ever seen _ten_ of them before.
- frequent 11y agonothing like having a short-movie done in 48hrs using only web services and then WeVideo goes stale just before I download... 2hrs before submission deadline :(
- driverdan 11y agoWhy is us-east-1 so terrible? All of the downtime this year has been Virginia.
- divideby0 11y agous-east-1 is where they typically deploy new features/hardware first (with the exception of efs which went to us-west first for some reason). it's also by far the largest region, with the most tenants and the heaviest traffic, so it's approaching the limit on what's physically possible to do in a public data center.
- toomuchtodo 11y agoIt's the primary AWS region. You spin up your resources by default there unless you explicitly select another region in the console.
- raverbashing 11y agoIt's probably a good idea to pick other regions, especially the ones closest to yourself However, us-east is usually the cheapest one as well
- scott_karana 11y agoIs it cheaper than the downtime?
- xenoclast 11y agoAs an AWS customer you need to be aware that the service health of all AWS services and not just the ones you use directly are important. You say you don't use SQS or SNS? When they go down, you might not be able to get Logs or even login to the web Console. Same goes for things like AutoScaling, OpsWorks, etc.
- gfosco 11y agoThat's the beauty of micro-service architectures. You don't have a single monolithic point of failure, you have dozens of smaller ones.
- dbarlett 11y agoS3 and VPC themselves appear to be fine, as noted on the dashboard, but the S3 VPC endpoints in EC2 are not ("we are also experiencing increased error rates accessing VPC endpoints for S3"). I was able to restore my sites by removing the endpoints from the routing tables.
- leesalminen 11y agoNot related to the AWS outage, but Rackspace CDN customers are in for a world of hurt today as well. https://status.rackspace.com/index/viewincidents?group=28 https://status.rackspace.com/index/viewincidents?group=28
- heapcity 11y agoWhat is happening to stock price? oh; its sunday; forgot we can't trade.
- bdcravens 11y agoI don't recall any outages having a material effect on Amazon's stock price.
- varelse 11y agoAmazon can apparently go up or down 50% depending on whether a butterfly sneezes in Australia (it's wildly danced between 284 and 580 in the past 12 months alone). In the background of such high volatility, it would be hard to pinpoint such a material effect from such a small disruption (in the big picture of course, I'm betting there are some pretty angry customers today due to the loss of a few sigma of reliability from this outage alone). Now if a study were published indicating customers were switching providers over incidents this, then I think you'd have some material evidence. But is anyone else better yet? Azure was out for 12 hours last year apparently... http://www.datacenterknowledge.com/archives/2015/01/23/cloud-reliability-aws-had-fewer-errors-than-azure-google-cloud-in-2014/ http://www.datacenterknowledge.com/archives/2015/01/23/cloud...
- archimedespi 11y agoReddit is down right now with a 503 - they're on AWS.
- aaawow 11y agoAmazon Echo doesn't work from 4am PST
- Tinyyy 11y agoIts pretty interesting to see how much our internet relies on cloud services like AWS, and all that are brought down with issues like this.
- williamcotton 11y agoFree Rugby World Cup! http://universalsports.com/ http://universalsports.com/ "RWC2015ppv.com has been affected by an internet outage. Watch here. Not all mobile devices are compatible"
- onetom 11y agoFor ADHD I would recommend Concerta/Ritalin/Adderall; that would enable you to read for not just a few seconds more but even minutes more before you judge. (I'm not joking here, I'm using 36mg Concerta for about about 4yrs now. Before that I also often came across as an asshole.) There is also a less evasive way of improving your online communication quality: http://www.paulgraham.com/disagree.html http://www.paulgraham.com/disagree.html
- jsprogrammer 11y agoThe feedback is valid. If they can't figure it out in 10 minutes, then they can't figure it out. It doesn't help that the inner workings are obfuscated through multiple license agreements and hidden/closed source. We are constrained by time. We can't invest our time into investigating every claim that we come across. Drugs can help some, but they are not necessarily the answer in all cases (including this case).
- onetom 11y agonadam said 10 seconds, not minutes. we are using the starter version combined with dynamodb in production and we found the payment structure very clear; no obfuscation whatsoever, unlike a microsoft or adobe pricing matrix ;) (it's made by the same guy who made the very open source clojure programming language, btw, and he is very much against obfuscation) anyway, it's a competitive advantage for those guys' who are building a bank on top of it in brazil (https://www.youtube.com/watch?v=7lm3K8zVOdY https://www.youtube.com/watch?v=7lm3K8zVOdY). we feel we can avoid writing a lot of authorization and audit log code by using datomic, maybe u can save such work too.
- nadams 11y agoThat was from my phone which is why it was sparse...but... > The Datomic software consists of the peer library and the transactor > These components connect to one of several storage service options I'm pretty well versed in computer science and buzzwords - but that means nothing to me. That was from the overview which I expected a dead simple "this is what it does and why you need it". There are several important questions that don't seem to be addressed: - I'm guessing it's a database proxy that's intelligent? - Is it better than HAProxy? - How is it different/better? And most importantly: - Do I need to modify my current code base to interact with this thing? In case you think I'm being overly dramatic - here are 2 examples: https://www.statuspage.io/ https://www.statuspage.io/ On the front page I know exactly what they offer. https://slack.com/is https://slack.com/is On the product page I can see exactly what slack is and offers.
- rocky1138 11y agoBetter to have this happen on a Sunday than a Monday.
- deleted 11y ago[deleted]
- sidcool 11y agoWow, very interesting to see how much of the infrastructure directly or indirectly depends on AWS.
- kiallmacinnes 11y agoGreat! Almost every takeaway in Dublin has moved to Zuppler for their online ordering... Zuppler is hosted out of AWS.
- zkhalique 11y agoThis is why 99.99999999% uptime is a fallacy It is not really measuring the time you're going to be up. That interpretation is based on faulty assumptions. It's like the statement "the sun will burn out before one bit is flipped" is wrong. It is quite likely that by that time, all the bits will be gone. https://signalvnoise.com/posts/3067-lets-get-honest-about-uptime https://signalvnoise.com/posts/3067-lets-get-honest-about-up...
- istvan__ 11y agoYes it is without a date range. 99.999% / year means something while 99.999% uptime does not mean anything by itself.
- rdtsc 11y agoYeah it is bullshit. They just had a failure so now they can claim, oh it is still 9 9s it is just that it is over 400 billion years averaged, not like you assumed, 10 billions. So legally still cool though...
- DinkyG 11y agoDynamoDb never promised 10 nines of availability, so it's a bit silly to hold them to that.
- chetanahuja 11y agoThe aws services stack is deep and deeply intertwined. I've always viewed depending on such stacks in production with skepticism and I'd recommend everybody else does that too. This might come across as tooting our horn a bit. But it's more about sounding a warning to other startups providing SaaS service built on public cloud. My own misgivings about relying on a cloud provider specific stack (both for the reasons of visibility/debuggability as well as for vendor lock-in) meant that PacketZoom services were not affected by this failure at all because we only use them as one of the many providers of raw machines. We use our own techniques to load-balance/failover among multiple cloud providers too (so even if the raw compute/network went away, our service would take a perf hit but not be completely down).
- not_kurt_godel 11y agoOr you could just run in multiple regions. Using multiple cloud providers limits your ability to take advantage of provider-specific features - why waste time writing your own load balancer when you could use ELB + multiple regions?
- chetanahuja 11y ago"Or you could just run in multiple regions." Not when the original goal of the very service is to have presence in all geographical regions. If aws us-east is hit, I want the users to transparently failover to a server on east coast (perhaps one hosted by google or softlayer) rather than be directed all the way to us-west or eu. And as for ELB, one doesn't use ELB for a custom protocol that load-balances/fails-over itself from the client :-)
- fapjacks 11y agoYou know this is interesting. There were no symptoms for us at all that something was wrong with Amazon itself, and their status page was not updated in a timely fashion. I spent a few hours (in the middle of the night working with my laptop in bed next to my wife) trying to figure out what in the hell was wrong, only to find out through the grapevine that it was Amazon. This is extremely frustrating when providers are having problems and actively working on a solution yet their status page still has glowing recommendations of their service.
- qaqy 11y agoclouds are so great
- samstave 11y agoThe last time there were API outages in AWS, our autoscaling logic could not determine the number of running instances, so it felt it had too few. It kept launching instances, and due to the API outage we couldn't manually kill the instances either... So we wound up with over 1,000 of these machines running which then due to our fan out of their DB they needed to load into memory from other machines, our whole environment crashed until we could kill off the erroneously launched instances. This meant an effective full reboot of our entire platform... It's was not a fun weekend.