27 ms·
Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
- JCM9 5y agoObviously one hopes these things don’t happen, but that’s an impressive and transparent write up that came out quickly.
- markranallo 5y agoIts not transparent at all. A massive amount of services were hard down for hours like SNS and were never acknowledged on the status page or in this write-up. This honestly reads like they don't truly understand the scope of things effected.
- nijave 5y agoIt sounded like the entire management plane was down and potentially part of the "data" plane too (management being config and data being get/put/poll to stateful resources) I saw in the Reddit thread someone mentioned all services that auth to other services on the backend were effected (not sure how truthful it is but that certainly made sense)
- WC3w6pXxgGd 5y ago> Our networking clients have well tested request back-off behaviors that are designed to allow our systems to recover from these sorts of congestion events, but, a latent issue prevented these clients from adequately backing off during this event. Sentences like this are confusing. If they are well-tested, wouldn't this issue have been covered?
- Fordec 5y agoBetween this and Log4j, I'm just glad it's Friday.
- eigen-vector 5y agoUnfortunately, patching vulns can't be put off for Monday.
- deleted 5y ago[deleted]
- _jal 5y agoYou are clearly not involved in patching.
- Fordec 5y agoSimply already patched. Company sizes and number of attack surfaces vary. 22 hours is plenty of time for an input string filter on a centrally controlled endpoint and a dependency increment with the right CI pipeline.
- zeko1195 5y agolol no
- deleted 5y ago[deleted]
- shagie 5y agoConsider the possible ways for a string to be injected into any of the following: Apache Solr Apache Druid Apache Flink ElasticSearch Flume Apache Dubbo Logstash Kafka If you've got any of them, they're likely exploitable too. That list comes from: https://unit42.paloaltonetworks.com/apache-log4j-vulnerability-cve-2021-44228/ https://unit42.paloaltonetworks.com/apache-log4j-vulnerabili... The attack surface is quite a bit larger than many realize. I recently had a conversation with a person who wasn't at a Java shop so wasn't worried... until he said "oh, wait, ElasticSearch is vulnerable too?" You'll even see it in things like the connector between CouchBase and ElasticSearch ( https://forums.couchbase.com/t/ann-elasticsearch-connector-4-3-3-4-2-13-fixes-log4j-vulnerability/32402 https://forums.couchbase.com/t/ann-elasticsearch-connector-4... ).
- eigen-vector 5y agoExceeded character limit on the title so I couldn't include this detail there, but this is the post-mortem of the event on December 7 2021.
- jtchang 5y agoThe complexity that AWS has to deal with is astounding. Sure having your main production network and a management network is common. But making sure all of it scales and doesn't bring down the other is what I think they are dealing with here. It must have been crazy hard to troubleshoot when you are flying blind because all your monitoring is unresponsive. Clearly more isolation with clearly delineated information exchange points are needed.
- dijit 5y ago“But AWS has more operations staff than I would ever hope to hire” — a common mantra when talking about using the cloud overall. I’m not saying I fully disagree. But consolidation of the worlds hosting necessitates a very complicated platform and these things will happen, either due to that complexity, failures that can’t be foreseen or good old fashioned Sod’s law. I know AWS marketing wants you to believe it’s all magic and rainbows, but it’s still computers.
- beoberha 5y agoI work for one of the Big 3 cloud providers and it’s always interesting when giving RCAs to customers. The vast majority of our incidents are due to bugs in the “magic” components that allow us to operate at such a massive scale.
- mperham 5y agoI wish it contained actual detail and wasn’t couched in generalities.
- nijave 5y agoThat was my take. Seems like boilerplate you could report for almost any incident. Last year's Kinesis outage and the S3 outage some years ago had some decent detail
- propter_hoc 5y agoDoes anyone know how often an AZ experiences an issue as compared to an entire region? AWS sells the redundancy of AZs pretty heavily, but it seems like a lot of the issues that happen end up being region-wide. I'm struggling to understand whether I should be replicating our service across regions or whether the AZ redundancy within a region is sufficient.
- nemothekid 5y agoI've been naively setting up our distributed databases in separate AZs for a couple years now, paying, sometimes, thousands of dollars per month in data replication bandwidth egress fees. As far as I can remember I've never never seen an AZ go down, and the only region that has gone down has been us-east-1.
- cyounkins 5y agoIs that separate AZs within the same region, or AZs across regions? I didn't think there were any bandwidth fees between AZs in the same region.
- TheP1000 5y agoThat is incorrect. Cross az fees are steep.
- electroly 5y agoIt's $0.01/GB for cross-AZ transfer within a region.
- nemothekid 5y agoIn reality it's more like $0.02/GB. You pay $0.01 on sending and $0.01 on receiving. I have no idea why ingress isn't free.
- rfraile 5y ago
- pinche_gazpacho 5y agoYeah, cloudwatch APIs went to the drain. Good for them for publishing this at least.
- iwallace 5y agoMy company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by their "green board" is meaningless. During the event they were late and reluctant to report it and its significance. My point is that they are wrongly incentivized to keep the board green at all costs.
- tootie 5y agoSecond hand info but supposedly when an outage hits they go all hands on resolving it and no one who knows what's going on has time to update the status board which is why it's always behind.
- voidfunc 5y agoNot AWS, but Azure: highly doubt. At least at Azure the moment you declare an outage there is a incident manager to handle customer communication. Bullshit someone at Amazon doesn’t have time to update the status.
- amalter 5y agoI mean, not to defend them too strongly, but literally half of this post mortem is addressing the failure of the Service Dashboard. You can take it on bad faith, but they own up to the dashboard being completely useless during the incident.
- dijit 5y agoOnce is a mistake. Twice is a coincidence. Three times is a pattern. But this… This is every time.
- doctor_eval 5y agoFour times is a policy.
- hourislate 5y agoBroadcast storm. Never easy to isolate, as a matter of fact it's nightmarish...
- foobarbecue 5y ago"... the networking congestion impaired our Service Health Dashboard tooling from appropriately failing over to our standby region. By 8:22 AM PST, we were successfully updating the Service Health Dashboard." Sounds like they lost the ability to update the dashboard. HN comments at the time were theorizing it wasn't being updated due to bad policies (need CEO approval) etc. Didn't even occur to me that it might be stuck in green mode.
- dgivney 5y agoIn the February 2017 S3 outage, AWS was unable to move status icons to the red icon because those images happened to be stored on the servers that went down. https://twitter.com/awscloud/status/836656664635846656 https://twitter.com/awscloud/status/836656664635846656
- deleted 5y ago[deleted]
- wongarsu 5y agoHasn't this exact thing (something in US-east-1 goes down, AWS loses ability to update dashboard) happened before? I vaguely remember it was one of the S3 outages, but I might be wrong. In any case, AWS not updating their dashboard is almost a meme by now. Even for global service outages the best you will get is a yellow.
- foobarbecue 5y agoYeah, probably. I haven't watched it this closely before during an outage. I have no idea if this happens in good faith, bad faith, or (probably) a mix.
- bpodgursky 5y agoDNS? Of course it was DNS. It is always* DNS.
- tezza 5y agoIt is often BGP, regularly DNS, frequently expired keys, sometimes a bad release and occasionally a fire
- shepherdjerred 5y agoThis wasn’t caused by DNS. DNS was just a symptom.
- simlevesque 5y ago> Customers accessing Amazon S3 and DynamoDB were not impacted by this event. However, access to Amazon S3 buckets and DynamoDB tables via VPC Endpoints was impaired during this event. What does this even mean ? I bet most people use DynamoDB via a VPC, in a Lambda or in EC2
- jtoberon 5y agoVPC Endpoint is a feature of VPC: https://docs.aws.amazon.com/vpc/latest/privatelink/endpoint-services-overview.html https://docs.aws.amazon.com/vpc/latest/privatelink/endpoint-...
- discodave 5y agoYour application can call DynamoDB via the public endpoint (dynamodb.us-east-1.amazonaws.com). But if you're in a VPC (i.e. practically all AWS workloads in 2021), you have to route to the internet (you need public subnet(s) I think) to make that call. VPC Endpoints create a DynamoDB endpoint in your VPC, from the documentation: "When you create a VPC endpoint for DynamoDB, any requests to a DynamoDB endpoint within the Region (for example, dynamodb.us-west-2.amazonaws.com) are routed to a private DynamoDB endpoint within the Amazon network. You don't need to modify your applications running on EC2 instances in your VPC. The endpoint name remains the same, but the route to DynamoDB stays entirely within the Amazon network, and does not access the public internet."
- aeyes 5y agoI call my DynamoDB tables via the public endpoint and it was severely impaired - high error rate and very high (second) latency.
- all2well 5y agoFrom within a VPC, you can either access DynamoDB via its public internet endpoints (eg, dynamodb.us-east-1.amazonaws.com, which routes through an Internet Gateway attachment in your VPC), or via a VPC endpoint for dynamodb that's directly attached to your VPC. The latter is useful in cases where you want a VPC to not be connected to the internet at all, for example.
- llaolleh 5y agoI wonder if they could've designed better circuit breakers for situations like this. They're very common in electrical engineering, but I don't think they're as common in software design. Something we should try to design and put in, actually for situations like this.
- EsotericAlgo 5y agoThey’re a fairly common design pattern https://en.m.wikipedia.org/wiki/Circuit_breaker_design_pattern https://en.m.wikipedia.org/wiki/Circuit_breaker_design_patte.... However, they certainly aren’t implemented with the frequency they should be at service level boundaries resulting in these sorts of cascading failures.
- riknos314 5y agoOne of the big issues mentioned was that one of the circuit breakers they did have (client back off), didn't function properly. So they did have a circuit breaker in the design, but it was broken.
- kevin_nisbet 5y agoNetflix was talking alot about circuit breaks a few years ago, and had the Hystrix project. Looks like Hystrix is discontinued, so I'm not sure if there are good library solutions that are easy to adopt. Overall I don't see it getting talked about that frequently... beyond just exponential backoff inside a retry loop. - https://github.com/Netflix/Hystrix https://github.com/Netflix/Hystrix - https://www.youtube.com/watch?v=CZ3wIuvmHeM https://www.youtube.com/watch?v=CZ3wIuvmHeM I think talks about Hystrix a bit, but I'm not sure if it's the presentation I'm thinking of from years ago or not.
- isbvhodnvemrwvn 5y agoIn the JVM land resilience4j is the de facto successor of Hystrix: https://github.com/resilience4j/resilience4j https://github.com/resilience4j/resilience4j
- bamboozled 5y agoStill doesn’t explain the cause of all the IAM permission denied requests we saw against policies which are again working fine without any intervention. Obviously networking issues can cause any number of symptoms but it seems like an unusual detail to leave out to me. Unless it was another ongoing outage happening at the same time.
- notimetorelax 5y agoIt’s so hard to know what was the state of the system when the monitoring was out. Wouldn’t be surprised if they don’t have the data to investigate it now.
- a45a33s 5y agohow are auth requests supposed to reach the auth server if the networking is broken?
- bamboozled 5y agoI’d accept this as an answer if I received a timeout or a message to say that. Permission denied is something altogether because it implies the request reached an authorisation system, was evaluated and denied.
- sbierwagen 5y agoFail-secure + no separate error for timeouts maybe? If the server can't be reached then it just denies the request.
- nayuki 5y ago> Operators instead relied on logs to understand what was happening and initially identified elevated internal DNS errors. Because internal DNS is foundational for all services and this traffic was believed to be contributing to the congestion, the teams focused on moving the internal DNS traffic away from the congested network paths. At 9:28 AM PST, the team completed this work and DNS resolution errors fully recovered. Having DNS problems sounds a lot like the Facebook outage of 2021-10-04. https://en.wikipedia.org/wiki/2021_Facebook_outage https://en.wikipedia.org/wiki/2021_Facebook_outage
- human 5y agoThe rule is that it’s always DNS.
- jessaustin 5y agoDNS seemed to be involved with both the Spectrum business internet and Charter internet outages overnight. So much for diversifying!
- shepherdjerred 5y agoIt’s quite a bit different… Facebook took themselves offline completely because of a bad BGP update, whereas AWS had network congestion due to a scaling event. DNS relies on the network, so of course it’ll be impacting if networking is also impacted.
- bdd 5y agono. it wasn't a "bad bgp update". bgp withdrawal of anycast addresses was a desired outcome of a region (serving location) getting disconnected from the backbone. if you'd like to trivialize it, you can say it was configuration change to the software defined backbone.
- cyounkins 5y agoMy favorite sentence: "Our networking clients have well tested request back-off behaviors that are designed to allow our systems to recover from these sorts of congestion events, but, a latent issue prevented these clients from adequately backing off during this event."
- dilap 5y agoOh you know, an editing error -- they accidentally dropped the word "not".
- discodave 5y agoI saw pleeeeeenty of untested code at Amazon/AWS. Looking back it was almost like the most important services/code had the least amount of testing. While internal boondoggle projects (I worked on a couple) had complicated test plans and debates about coverage metrics.
- dastbe 5y agomy take is that the overwhelming majority of services insufficiently invest in making testing easy. the services that need to grow fast due to customer demand skip the tests while the services that aren't going much of anywhere spend way too much time on testing.
- DrBenCarson 5y agoThis is almost always the case. The most important services get the most attention from leaders who apply the most pressure, especially in the first ~2y of a fast-growing or high-potential product. So people skip tests.
- foobiekr 5y agoreality most of the real world successful projects are mostly untested because that's not actually a high ROI endeavor. it kills me to realize that mediocre code you can hack all over to do unnatural things is generally higher value in phase I than the same code done well in twice the time.
- DenisM 5y ago> Customers accessing Amazon S3 and DynamoDB were not impacted by this event. We've seen plenty of S3 errors during that period. Kind of undermines credibility of this report.
- marcinzm 5y agoDAX, part of DynamoDB from how AWS groups things, was throwing internal server errors for us and eventually we had to reboot nodes manually. That's separate from the STS issues we had in terms of our EKS services connecting to DAX.
- banana_giraffe 5y agoYeah > For example, while running EC2 instances were unaffected by this event That ignores the 3-5% drop in traffic I saw in us-east-1 on EC2 instances that only talk to peers on the Internet with TCP/IP during this event.
- manquer 5y agoI guess you have to read this kind of items hidden in careful language, running the instances had no problem, it is different matter they had limited connectivity!from AWS point of view they don't seem to see user impact but services from their point of view. Perhaps that distinction has value if your workloads did not depend on network connectivity externally for example say S3 access without vpc and only compute some DS/ ML jobs perhaps.
- grogers 5y agoHow are you measuring this? Remember that cloudwatch was apparently also losing metrics, so aggregating CW metrics might show that kind of drop.
- banana_giraffe 5y agoYeah, I know. This was based off instance store logging on these instances. For better or worse, they're very simple ports of pre-AWS on-prem servers, they don't speak AWS once they're up and running.
- JCM9 5y agoQueue the armchair infrastructure engineers. The reality is that there’s a handful of people in the world that can operate systems at this sheer scale and complexity and I have mad respect for those in that camp.
- tristor 5y agoSome of us are in that camp and are looking at this outage and also pointing out that they continuously fail to accurately update their status dashboard in this and prior outages. Yes, doing what AWS does is hard, and yes outages /will/ happen, it is no knock on them that this outage occurred, what is a knock is that they haven't communicated honestly while the outage was ongoing.
- JCM9 5y agoThey address that in the post, and between Twitter, HN and other places there wasn’t anyone legit questioning if something was actually broken. Contacts at AWS also all were very clear that yes something was going on and being investigated. This narrative that AWS was pretending nothing was wrong just wasn’t true based on what we saw.
- DrBenCarson 5y agoI'm going to leave it at this: the dashboards at AWS aren't automated. Say what you will, but I can automate a status dashboard in a couple days--yes, even at AWS scale. No reason the dashboard should be green for hours while their engineers and support are aware things aren't working.
- notinty 5y agoApparently VP approval is required to update it, i.e. they're a farce.
- bradknowles 5y agoUh, no. You can’t. If you could, then you would already have been hired and you would have already solved this problem. What you can do at what you think is AWS scale has no bearing on what you could actually do at real AWS scale.
- StreamBright 5y ago"At 7:30 AM PST, an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network. " Very detailed.
- almostdeadguy 5y ago> The AWS container services, including Fargate, ECS and EKS, experienced increased API error rates and latencies during the event. While existing container instances (tasks or pods) continued to operate normally during the event, if a container instance was terminated or experienced a failure, it could not be restarted because of the impact to the EC2 control plane APIs described above. This seems pretty obviously false to me. My company has several EKS clusters in us-east-1 with most of our workloads running on Fargate. All of our Fargate pods were killed and were unable to be restarted during this event.
- ClifReeder 5y agoStrong agree. We were using Fargate nodes in our us-east-1 EKS cluster and not all of our nodes dropped, but every coredns pod did. When they came back up their age was hours older than expected, so maybe a problem between Fargate and the scheduler rendered them “up” but unable to be reached? Either way, was surprising to us that already provisioned compute was impacted.
- silverlyra 5y agoSaw the same. The only cluster services I was running in Fargate were CoreDNS and cluster-autoscaler; thought it would help the clusters recover from anything happening to the node group where other core services run. Whoops. Couldn't just delete the Fargate profile without a working EKS control plane. I lucked out in that the label selector the kube-dns Service used was disjoint from the one I'd set in the Fargate profile, so I just made a new "coredns-emergency" deployment and cluster networking came back. (cluster-autoscaler was moot since we couldn't launch instances anyway.) I was hoping to see something about that in this announcement, since the loss of live pods is nasty. Not inclined to rely on Fargate going forward. It is curious that you saw those pod ages; maybe Fargate kubelets communicate with EKS over the AWS internal network?
- soheil 5y agoHaving an internal network like this that everything on the main AWS network so heavily depends on is just bad design. One does not create a stable high tech spacecraft and then fuels it with coal.
- discodave 5y agoIt's 2006, you work for an 'online book store' that's experimenting with this cloud thing. Are you going to build a whole new network involving multi-million dollar networking appliances?
- deleted 5y ago[deleted]
- yegle 5y agoWas this outage only impact us-east-1 region? I think I saw other regions affected in some HN comments but this summary did not mention anything to suggest it has more than 1 region impacted.
- teej 5y agoThere are some AWS services, notably STS, that are hosted in us-east-1. I don’t have anything in us-east-1 but I was completely unable to log into the console to check on the health of my services.
- grouphugs 5y agostop using aws, i can't wait till amazon is hit so hard everyday they can't maintain customers
- azundo 5y ago> This resulted in a large surge of connection activity that overwhelmed the networking devices between the internal network and the main AWS network, resulting in delays for communication between these networks. These delays increased latency and errors for services communicating between these networks, resulting in even more connection attempts and retries. This led to persistent congestion and performance issues on the devices connecting the two networks. I remember my first experience realizing the client retry logic we had implemented was making our lives way worse. Not sure if it's heartening or disheartening that this was part of the issue here. Our mistake was resetting the exponential backoff delay whenever a client successfully connected and received a response. At the time a percentage but not all responses were degraded and extremely slow, and the request that checked the connection was not. So a client would time out, retry for a while, backing off exponentially, eventually successfully reconnect and then after a subsequent failure start aggressively trying again. System dynamics are hard.
- colechristensen 5y ago> System dynamics are hard. And have to be actually tested. Most of them are designs based on nothing but uninformed intuition. There is an art to back pressure and keeping pipelines optimally utilized. Queueing doesn’t work like you think until you really know.
- vmception 5y ago> Most of them are designs based on nothing but uninformed intuition. Or because they read it on a Google|AWS Engineering blog
- vmception 5y agoOr they regurgitated a bullshit answer from a system design prep course while pretending to think of it on the spot just to get hired
- gfodor 5y agoWhy is this hard, and can’t just be written down somewhere as part of the engineering discipline? This aspect of systems in 2021 really shouldn’t be an “art.”
- markus_zhang 5y ago>At 7:30 AM PST, an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network. Just curious, is this scaling an AWS job or a client job? Looks like an AWS one from the context. I'm wondering if they are deploying additional data centers or something else?
- whatever1 5y agoNoob question, but why does network infrastructure need dns? Why the full ipv6 address of the various components do not suffice to do business?
- nijave 5y agoIt's basically used for service discovery. At a certain point, you have too many different devices which are potentially changing to identify them by IP. You want some abstraction layer to separate physical devices from services and DNS lets you so things like advertise different IPs at different times in different network zones
- AUX4R6829DR8 5y agoThat "internal network" hosts an awful lot of stuff- it's not just network hardware, but services that mostly use DNS to find each other. Besides that, it's just plain useful for network devices to have names. (Source: Work at AWS.)
- betaby 5y agoProblem is that I have to defend our own infrastructure real availability numbers vs cloud's fictional "five nines". It's a loosing game.
- Spivak 5y agoAll I’m hearing is that you can make up your own availability numbers and get away with it. When you define what it means to be up or down then reality is whatever you say it is. #gatekeep your real availability metrics #gaslight your customers with increased error rates #girlboss
- WoahNoun 5y agoWhat are you trying to imply with that last hashtag?
- Spivak 5y agoIt’s a meme; search it on Twitter. It’s a play on “live, laugh, love” that started as a way for young women to mock pandering displays of female empowerment but has grown in scope so that it can be used to mock anyone. #gatekeep #gaslight #girlboss or the male equivalent #mansplain #manipulate #malewife
- mpyne 5y agoSome orgs really do have lousy availability figures (such as my own, the Navy). We have an environment we have access to for hosting webpages for one of the highest leaders in the whole Dept of Navy. This environment was DOWN (not "degrade availability" or "high latencies"), literally off of the Internet entirely, for CONSECUTIVE WEEKS earlier this year. Completely incommunicado as well. It just happened to start working again one day. We collectively shrugged our shoulders and resumed updating our part of it. This is an outlier example but even our normal sites I would classify as 1 "nine" of availability at best.
- garbagecoder 5y ago
- tyingq 5y ago"Amazon Secure Token Service (STS) experienced elevated latencies" I was getting 503 "service unavailable" from STS during the outage most of the time I tried calling it. I guess by "elevated latency", they mean from anyone with retry logic that would keep trying after many consecutive attempts?
- jrockway 5y agoI suppose all outages are just elevated latency. Has anyone ever had an outage and said "fuck it, we're going out of business" and never came back up? That's the only true outage ;)
- hericium 5y ago5xx errors are servers or proxies giving up on requests. Increased timeouts resulting in successful requests may have been considered "elevated latency" (but rarely this would be a proper way to solve similar issue). They treat 5xx errors as non-errors but this is not the case with rest of the world. "Increased timeouts" is Amazon's untruthful term for "not working at all".
- WaxProlix 5y agoSTS is the worst with this. Even for other internal teams, they seem to treat dropped requests (ie, timeouts which represent 5xxs on the client side) as 'non faults', and so don't treat those data points in their graphs and alarms. It's really obnoxious. AWS in general is trying hard to do the right thing for customers, and obviously has a long ways to go. But man, a few specific orgs have some frustrating holdover policies.
- hericium 5y ago> AWS in general is trying hard to do the right thing for customers You are responding to a comment that suggests they're misrepresenting the truth (which wouldn't be the first time even in last few days) in communication to their customers. As always, they are doing the right thing for themselves only. EDIT: I think that you should mention being an Engineer at Amazon AWS in your comment.
- sneak 5y ago"impact" occurs 27 times on this page. What was wrong with "affect"?
- pohl 5y agoThe easiest way to avoid confusing affect with effect is to use other words.
- wjossey 5y agoI’ve been running platform teams on aws now for 10 years, and working in aws for 13. For anyone looking for guidance on how to avoid this, here’s the advice I give startups I advise. First, if you can, avoid us-east-1. Yes, you’ll miss new features, but it’s also the least stable region. Second, go multi AZ for production workloads. Safety of your customer’s data is your ethical responsibility. Protect it, back it up, keep it as generally available as is reasonable. Third, you’re gonna go down when the cloud goes down. Not much use getting overly bent out of shape. You can reduce your exposure by just using their core systems (EC2, S3, SQS, LBs, Cloudfrount, RDS, Elasticache). The more systems you use, the less reliable things will be. However, running your own key value store, api gateway, event bud, etc., can also be way less reliable than using their’s. So, realize it’s an operational trade off. Degradation of your app / platform is more likely to come from you than AWS. You’re gonna roll out bad code, break your own infra, overload your own system, way more often than Amazon is gonna go down. If reliability matters to you, start by examining your own practices first before thinking things like multi region or super durable highly replicated systems. This stuff is hard. It’s hard for Amazon engineers. Hard for platform folks at small and mega companies. It’s just, hard. When your app goes down, and so does Disney plus, take some solace that Disney in all their buckets of cash also couldn’t avoid the issue. And, finally, hold cloud providers accountable. If they’re unstable and not providing service you expect, leave. We’ve got tons of great options these days, especially if you don’t care about proprietary solutions. Good luck y’all!
- daguava 5y agoYou've written up my thoughts better than I can express them myself - I think what people get really stuck on when something like this happens is the 'can I solve this myself?' aspect. A wait for X provider to fix it for you situation is infinitely more stressful than an 'I have played myself, I will now take action' situation. Situations out of your (immediate) resolution control feel infinitely worse, even if the customer impact in practice of your fault vs cloud fault is the same.
- electroly 5y agoI couldn't possibly disagree more strongly with this. I used to drive frantically to the office to work on servers in emergency situations, and if our small team couldn't solve it, there was nobody else to help us. The weight of the outage was entirely on our shoulders. Now I relax and refresh a status page.
- wly_cdgr 5y agoTheir service board is always as green as you have to be to trust it
- AtlasBarfed 5y ago"At 7:30 AM PST, an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network. This resulted in a large surge of connection activity that overwhelmed the networking devices between the internal network and the main AWS network, resulting in delays for communication between these networks. These delays increased latency and errors for services communicating between these networks, resulting in even more connection attempts and retries." So was this in service to something like DynamoDB or some other service? As in, did some of those extra services that AWS offers for lockin (and that undermines open source projects with embrace and extend) bomb the mainline EC2 service? Because this kind of smacks of "Microsoft Hidden APIs" that office got to use against other competitors. Does AWS use "special hardware capabilites" to compete against other companies offering roughtly the same service?
- nijave 5y agoYes and other cloud providers (Google, Microsoft) probably have similar. Besides special network equipment, they use PCIe accelerator/coprocessors on their hypervisors to offload all non-VM activity (Nitro instances) They also recently announced Graviton ARM CPUs
- User23 5y agoA packet storm outage? Now that brings back memories. Last time I saw that it was rendezvous misbehaving.
- rodmena 5y agoI was alive, however because I could not breath, I died. Bob was fine himself, but someone shot him, so he is dead, (but remember bob was fine) --- What a joke
- rodmena 5y agoumm... But just one thing, S3 was not available at least for 20 minutes.
- divbzero 5y ago> This congestion immediately impacted the availability of real-time monitoring data for our internal operations teams, which impaired their ability to find the source of congestion and resolve it. Disruption of the standard incident response mechanism seems to be a common element of longer lasting incidents.
- femiagbabiaka 5y agoIt is. And to add, all automation that we rely on in peace time can often complicate cross cutting wartime incidents by raising the ambient complexity of an environment. Bainbridge for more: https://blog.acolyer.org/2020/01/08/ironies-of-automation/ https://blog.acolyer.org/2020/01/08/ironies-of-automation/
- daenney 5y agoYup. There was a GCP outage a couple of years ago like this. I don’t remember the exact details, but it was something along the lines of a config change went out that caused systems to incorrectly assume there were huge bandwidth constraints. Load shedding kicked in to drop lower priority traffic which ironically included monitoring data rendering GCP responders blind and causing StackDriver to go blank for customers.
- sponaugle 5y agoIndeed - Even the recent facebook outage outlined how slow recovery can be if the primary investigation and recovery methods are directly impacted as well. Back in the old days some environments would have POTS dial-in connections to the consoles as backup for network problems. That of course doesn't scale, but it was an attempt to have an alternate path of getting to things. Regrettably if a backhoe takes out all of the telecom at once that plan doesn't work so well.
- qwertyuiop_ 5y agoHouse of cards
- onion2k 5y agoThis congestion immediately impacted the availability of real-time monitoring data for our internal operations teams I guess this is why it took ages for the status page to update. They didn't know which things to turn red.
- jetru 5y agoComplex systems are really really hard. I'm not a big fan of seeing all these folks bash AWS for this, and not really understanding the complexity or nastiness of situations like this. Running the kind of services they do for the kind of customers, this is a VERY hard problem. We ran into a very similar issue, but at the database layer in our company literally 2 weeks ago, where connections to our MySQL exploded and completely took down our data tier and caused a multi-hour outage, compounded by retries and thundering herds. Understanding this problem under the stressful scenario is extremely difficult and a harrowing experience. Anticipating this kind of issue is very very tricky. Naive responses to this include "better testing", "we should be able to do this", "why is there no observability" etc. The problem isn't testing. Complex systems behave in complex ways, and its difficult to model and predict, especially when the inputs to the system aren't entirely under your control. Individual components are easy to understand, but when integrating, things get out of whack. I can't stress how difficult it is to model or even think about these systems, they're very very hard. Combined with this knowledge being distributed among many people, you're dealing with not only distributed systems, but also distributed people, which adds more difficulty in wrapping this around your head. Outrage is the easy response. Empathy and learning is the valuable one. Hugs to the AWS team, and good learnings for everyone.
- xamde 5y agoThere is this really nice website which explains how complex systems fail: https://how.complexsystems.fail/ https://how.complexsystems.fail/
- fastball 5y agoI think most of the outrage is not because "it happened" but because AWS is saying things like "S3 was unaffected" when the anecdotal experience of many in this thread suggests the opposite. That and the apparent policy that a VP must sign off on changing status pages, which is... backwards to say the least.
- jetru 5y agoThere's definitely miscommunication around this. I know I've miscommunicated impact, or my communication was misinterpreted across the 2 or 3 people it had to jump before hitting the status page. For example, The meaning of "S3 was affected" is subject to a lot of interpretation. STS was down, which is a blocker for accessing S3. So, the end result is S3 is effectively down, but technically it is not. How does one convey this in a large org? You run S3, but not STS, it's not technically an S3 fault, but an integration fault across multiple services. If you say S3 is down, you're implying that the storage layer is down. But it's actually not. What's the best answer to make everyone happy here? I cant think of one.
- revskill 5y agoMost of rate limiter system often drop invalid requests, it's not optimal as i see. The better way is, we should have two queues, one for valid messages and one for invalid messages.
- xyst 5y agoI am not a fan of AWS due to their substantial market share on cloud computing. But as a software engineer I do appreciate their ability to provide fast turnarounds on root cause analyses and make them public.
- waz0wski 5y agoThis isn't a good example of an RCA - as other commenters have noted, it's outrightly lying about some issues during the incident, and using creative language to dance around other problems many people encountered. If you want to dive into postmortems, there are some repos linking other examples https://github.com/danluu/post-mortems https://github.com/danluu/post-mortems https://codeberg.org/hjacobs/kubernetes-failure-stories https://codeberg.org/hjacobs/kubernetes-failure-stories
- amznbyebyebye 5y agoI’m glad they published something, that too so quick. Ultimately these guys are running a business. There are other market alternatives, multibillion dollar contracts at play, SLAs, etc. it’s not as simple as people think.
- herodoturtle 5y agoI am grateful to AWS for this report. Not sure if any AWS support staff are monitoring this thread, but the article said: > Customers also experienced login failures to the AWS Console in the impacted region during the event. All our AWS instances / resources are in EU/UK availability zones, and yet we couldn't access our console either. Thankfully none of our instances were affected by the outage, but our inability to access the console was quite worrying. Any idea why this was this case? Any suggestions to mitigate this risk in the event of a future outage would be appreciated.
- deleted 5y ago[deleted]
- plasma 5y agoThey posted on the status page to try using the alternate region endpoints like us-west.console.Amazon.com (I think) at the time, but not sure if it was a true fix.
- raffraffraff 5y agoSomething they didn't mention is AWS Billing alarms. These rely on metrics systems which were affected by this (and are missing some data). Crucially, billing alarms only exist in the us-east-1 region, so if you're using them, your impacted no matter where you're infrastructure is deployed. (That's just my reading of it)
- londons_explore 5y agoIdea:. Network devices should be configured to automatically prioritize the same packet flows for the same clients as they served yesterday. So many overload issues seem to be caused by a single client, in a case where the right prioritization or rate limit rule could have contained any outage, but such a rule either wasn't in place or wasn't the right one due to the difficulty of knowing how to prioritize hundreds of clients. Using more bandwidth or requests than yesterday should then be handled as capacity allows, possibly with a manual configured priority list, cap, or ratio. But "what I used yesterday" should always be served first. That way, any outage is contained to clients acting differently to yesterday, even if the config isn't perfect.
- stevefan1999 5y agoIn a nutshell: thundering herd.
- Ensorceled 5y agoThere are a lot of comments in here that boil down to "could you do infrastructure better?" No, absolutely not. That's why I'm on AWS. But what we are all ACTUALLY complaining about is ongoing lack of transparent and honest communications during outages and, clearly, in their postmortems. Honest communications? Yeah, I'm pretty sure I could do that much better than AWS.
- moogly 5y agoHm. This post does not seem to acknowledge what I saw. Multiple hours of rate-limiting kicking in when trying to talk to S3 (eu-west-1). After the incident everything works fine without any remediations done on our end.
- sponaugle 5y ago"Our networking clients have well tested request back-off behaviors that are designed to allow our systems to recover from these sorts of congestion events, but, a latent issue prevented these clients from adequately backing off during this event. " That is an interesting way to phrase that. A 'well-tested' method, but 'latent issues'. That would imply the 'well-tested' part was not as well-tested as it needed to be. I guess 'latent issue' is the new 'bug'.
- paulryanrogers 5y agoHas anyone been credited by AWS for violations of their SLAs?
- atoav 5y agoA "service event"?!