15 ms·
Summary of the Amazon Kinesis Event in the Northern Virginia (US-East-1) Region
- codesparkle 6y agoFrom the postmortem: At 9:39 AM PST, we were able to confirm a root cause [...] the new capacity had caused all of the servers in the fleet to exceed the maximum number of threads allowed by an operating system configuration.
- lytigas 6y ago> During the early part of this event, we were unable to update the Service Health Dashboard because the tool we use to post these updates itself uses Cognito, which was impacted by this event. Poetry. Then, to be fair: > We have a back-up means of updating the Service Health Dashboard that has minimal service dependencies. While this worked as expected, we encountered several delays during the earlier part of the event in posting to the Service Health Dashboard with this tool, as it is a more manual and less familiar tool for our support operators. To ensure customers were getting timely updates, the support team used the Personal Health Dashboard to notify impacted customers if they were impacted by the service issues. I'm curious if anyone here actually got one of these.
- 0x11 6y agoI can't say for sure that the company I work for didn't, but it certainly didn't make it's way to me and there are only 8 of us.
- vishnugupta 6y agoThis won't be a first. The status page was hosted in S3. It is hilarious in the hindsight, but understandable.
- capableweb 6y ago> but understandable Is it really? I get the value of eating your own dogfood, it improves things a lot. But your status page? Such a high importance, low difficulty thing to build that dogfeeding it gives you small amount of benefits (dogfeed something bigger/more complex instead) in the good case, and high amount of drawback when things go wrong (like when your infrastructure goes down, so does your status page). So what's the point?
- KingOfCoders 6y agoArrogance.
- tpetry 6y agoI can really imafgine what happened: Engineer wants to host dashboard at different provider for resilience. Manager argues that they cant do this, it would be embarassing if anybody found out. And why choose another provider? Aws has multiple AZs and cant be down everywhere at the same moment. Engineer then says „fu it“ and just builds it on a single solution.
- ufmace 6y agoMy employer is a pretty big spender with AWS. I didn't hear anything about anybody getting status updates from a "Personal Health Dashboard" or anywhere else. I can't be 100% sure such an update would have made its way to me, but given the amount of buzzing, it's hard to believe that somebody had info like that and didn't share it.
- newscom59 6y agoThe PHD is always updated first, long before the global status page is updated. Every single one of my clients that use AWS got updates on the PHD literally hours before the status page was even showing any issues, which is typical. It’s the entire point of the PHD. Through reading Reddit and HN during this event I learned that most people apparently aren’t even aware of the existence of the PHD and rely solely on the global status page, despite the fact that there is a giant “View my PHD” button at the very top of the global status page, and additionally there is a notification icon on the header of every AWS console page that lights up and links you directly to the PHD whenever there is an issue. The PHD is always where you should look first. It is, by design, updated long before the global status page is.
- robbiemitchell 6y ago> despite the fact that there is a giant “View my PHD” button at the very top of the global status page If you don’t know what the PHD is, a big button pointing to it won’t do anything. People ignore big boxes of irrelevant stuff all the time. AWS user of ~8 years and I’ve never heard of the PHD nor this sequencing of updating it first.
- mwarkentin 6y agoYes, we had some messages coming through in our PHD.
- loriverkutya 6y agoI can confirm we got the Personal Health Dashboard notifications.
- freeone3000 6y agoThe failure to update the Service Health Dashboard was due to reliance on internal services to update. This also happened in March 2017[0]. Perhaps a general, instead of piecemeal, approach to removing dependencies on running services from the dashboard would be valuable here? 0: https://aws.amazon.com/message/41926/ https://aws.amazon.com/message/41926/
- roman_sf 6y agoBut this time it was a _different_ dependency. They just want to make sure all dependencies are ruled out before migrating to GCP :)
- tnolet 6y agoThis is a pretty damn decent post mortem so soon after the outage. Also gives an architectural analysis of how Kinesis works which is something they had not have to do at all.
- ris 6y agoThe one thing I want to know in cases like this is: why did it affect multiple Availability Zones? Making a resource multi-AZ is a significant additional cost (and often involves additional complexity) and we really need to be confident that typical observed outages would actually have been mitigated in return.
- EdwardDiego 6y agoIndeed. We're paying (and designing our systems to work on multiple AZs) to reduce the risk of outages, but then their back-end services are reliant on services in a sole region?
- qz2 6y agoCorrect. I, as many people have, discovered this when something broke in one of the golden regions. In my case cloudfront and ACM. Realistically you can’t trust one provider at all if you have high availability requirements. The justification is apparently that the cloud is taking all this responsibility away from people but from personal experience running two cages of kit at two datacenters the TCO was lower and the reliability and availability higher. Possibly the largest cost is navigating Harry-Potter-esque pricing and automation laws. The only gain is scaling past those two cages. Edit: I should point out however that an advantage of the cloud is actually being able to click a couple of buttons and get rid of two cages worth of DC equipment instantly if your product or idea doesn't work out!
- lttlrck 6y agoHarry-Potter-esque pricing? Is that a reference to the difficulty of calculating the cost of visiting all the rides at universal? That's my best guess...
- qz2 6y agoIt's more a stab at the inconsistency of rules around magic. "Well this pricing rule only works on a Tuesday lunch time if you're standing on one leg with a sausage under each arm and a traffic cone on your head" And there are a million of those to navigate.
- pps43 6y ago> the new capacity had caused all of the servers in the fleet to exceed the maximum number of threads allowed by an operating system configuration. [...] We didn’t want to increase the operating system limit without further testing Is it because operating system configuration is managed by a different team within the organization?
- sitharus 6y agoMore likely they need to understand what effect changing the thread limit would have - for example it could increase kernel memory usage or increase scheduler latency. It’s not something you want to mess with in an outage.
- sudhirj 6y agoI’ve heard AWS follows a you build it, you run it policy, so that seems unlikely. Just seems prudent to not mess with OS settings in a hurry.
- mcqueenjordan 6y agoNope. It's just a case of "stop the bleeding before starting the surgery."
- Androider 6y agoIf you start haphazardly changing things while firefighting without testing, you might make things even worse. And there's worse things than downtime, for instance if the system appears to work but you're actually silently corrupting customer data.
- metaedge 6y agoI would have started the response with: First of all, we want to apologize for the impact this event caused for our customers. While we are proud of our long track record of availability with Amazon Kinesis, we know how critical this service is to our customers, their applications and end users, and their businesses. We will do everything we can to learn from this event and use it to improve our availability even further. Then move on to explain...
- sigzero 6y agoWhat they did was fine.
- temp0826 6y agous-east-1 is AWS’s dirty secret. If ddb had gone down there, there would likely be a worldwide and multi-service interruption.
- ignoramous 6y agoroot-cause tldr: ...[adding] new capacity [to the front-end fleet] had caused all of the servers in the [front-end] fleet to exceed the maximum number of threads allowed by an operating system configuration [number of threads spawned is directly proportional to number of servers in the fleet]. As this limit was being exceeded, cache construction was failing to complete and front-end servers were ending up with useless shard-maps that left them unable to route requests to back-end clusters. fixes: ...moving to larger CPU and memory servers [and thus fewer front-end servers]. Having fewer servers means that each server maintains fewer threads. ...making a number of changes to radically improve the cold-start time for the front-end fleet. ...moving the front-end server [shard-map] cache [that takes a long time to build, up to an hour sometimes?] to a dedicated fleet. ...move a few large AWS services, like CloudWatch, to a separate, partitioned front-end fleet. ...accelerate the cellularization [0] of the front-end fleet to match what we’ve done with the back-end. [0] https://www.youtube.com/watch?v=swQbA4zub20 https://www.youtube.com/watch?v=swQbA4zub20 and https://assets.amazon.science/c4/11/de2606884b63bf4d95190a3c2390/millions-of-tiny-databases.pdf https://assets.amazon.science/c4/11/de2606884b63bf4d95190a3c...
- frankietaylr 6y agoI wonder how many of them are already logged engineering tasks which never got prioritized because of the aggressive push to add features.
- deleted 6y ago[deleted]
- joneholland 6y agoRunning out of file handles and other IO limits is embarrassing and happens at every company, but I’m surprised that AWS was not monitoring this. I’m also surprised at the general architecture of Kinesis. What appears to be their own hand rolled gossip protocol (that is clearly terrible compared to raft or paxos, a thread per cluster member? Everyone talking to everyone? An hour to reach consensus?) and the front end servers being stateful period breaks a lot of good design choices. The problem with growing as fast as Amazon has is that their talent bar couldn’t keep up. I can’t imagine this design being okay 10 years ago when I was there.
- marcinzm 6y agoI don't think it's about growing fast so much as, from those I talked to, Amazon now has a fairly bad reputation in the tech community. You only go to work there if you don't have a better option (Google, Facebook, etc) or have some specialty skill they're willing to pay for. Pay is below other FAANG companies and the work culture isn't great (toxic even some would say). edit: They also had the most disorganized and de-centralized interview approach from all the FAANG companies I talked with. Which isn't growing pains this far in, it's just bad management and process.
- imajoredinecon 6y agoInteresting re interview experience I interviewed as a new grad SWE and the process was totally straightforward, and way lower friction (albeit much less human interaction, which made it feel even more impersonal) than almost everywhere else I applied: initial online screen, online programming task, and then a video call with an engineer where you explained your answer to the programming task.
- akhilcacharya 6y agoThe process for new grads and interns is different from industry hires and is decided by team.
- 6y ago
- fafner 6y agoFrom the summary I don't understand why front end servers need to talk to each other ("continuous processing of messages from other Kinesis front-end servers"). It sounds like this is part of building the shard map or the cache. Well in the end an unfortunate design decision. #hugops for the team handling this. Cascading failures are the worst.
- karmakaze 6y agoSeems to me that the root problem could also be fixed by not using presumably blocking application threads talking to each of the other servers. Any async or poll mechanism wouldn't require N^2 threads across the pool.
- why-el 6y agoI wonder if the new wonders coming out from linux (io_uring...) would have made this a better design, but that work in the kernel is still in active development.
- tmk1108 6y agoHow does the architecture of Kinesis compare to Kafka? If you scale up the number of Kafka brokers can you hit similar problem? Or does Kafka not rely on creating threads to connect to each other broker
- aloknnikhil 6y agoKafka uses a thread pool for request processing. Both the brokers and the consumer clients use the same request processing loop. This goes a bit more in-depth: https://jaceklaskowski.gitbooks.io/apache-kafka/content/kafka-server-KafkaServer.html https://jaceklaskowski.gitbooks.io/apache-kafka/content/kafk...
- hintymad 6y agoA tangential question, why would AWS even use the term "microservice"? A service is a service, right? I'm not sure what the term "microservice" signifies here.
- arduinomancer 6y agoIt’s because service can be confused with “AWS Service” which is not the same as a microservice (a component of a full service)
- londons_explore 6y agoOne requirement on my "production ready" checklist is that any catastrophic system failure can be resolved by starting a completely new instance of the service, and it be ready to serve traffic inside 10 minutes. That should be tested at least quarterly (but preferably automatically with every build). If Amazon did that, this outage would have been reduced to 10 mins, rather than the 12+ hours that some super slow rolling restarts took...
- why-el 6y agoThe same OS limits would apply to new instances, unless they knew the root cause and forced new instances to be configured with larger descriptor limits, which is....well, hindsight is 20/20, no?
- WookieRushing 6y agoThis only works for stateless services. If you’ve got frontends that take longer than 10 mins to serve traffic then you have a problem. But if you’re running a DB or a storage system, 10 mins is a blink of an eye. Storage systems in particular can run a few hundred TB per node and moving that data to another node can take over an hour. In this case, the frontends have a shard map which is definitely not stateless. This is typically okay if you have a fast load operation which blocks other traffic until shard map is fully loaded
- londons_explore 6y agoIt's possible (albeit much harder) for stateful services too. It basically boils down to "We must be able to restore the minimum necessary parts of a full backup in under 10 minutes". Take wikipedia as an example. I'd expect them to be able to restore a backup of the latest version of all pages in 10 minutes. It's 20GB of data, and I assume it's sharded at least 10 ways. That means each instance will have to grab 2GB from the backups. Very do-able. As a service gets bigger, you typically scale horizontally, so the problem doesn't get harder. Restoring all the old page versions and re enabling editing might take longer, but that's less critical functionality.
- WJW 6y ago
- steelframe 6y ago> Cellularization is an approach we use to isolate the effects of failure within a service, and to keep the components of the service (in this case, the shard-map cache) operating within a previously tested and operated range. This had been under way for the front-end fleet in Kinesis, but unfortunately the work is significant and had not yet been completed. Translation: The eng team knew that they had accumulated tech debt by cutting a corner here in order to meet one of Amazon's typical and insane "just get the feature out the door" timelines. Eng warned management about it, and management decided to take the risk and lean on on-call to pull heroics to just fix any issues as they come up. Most of the time yanking a team out of bed in the middle of the night works, so that's the modus operandi at Amazon. This time, the actual problem was more fundamental and wasn't effectively addressable with middle-of-the-night heroics. Management rolled the "just page everyone and hope they can fix it" dice yet again, as they usually do, and this time they got snake eyes. I guarantee you that the "cellularization" of the front-end fleet wasn't actually under way, but the teams were instead completely consumed with whatever the next typical and insane "just get the feature out the door" thing was at AWS. The eng team was never going to get around to cellularizing the front-end fleet because they were given no time or incentive to do so by management. During/after this incident, I wouldn't be surprised if management didn't yell at the eng team, "Wait, you KNEW this was a problem, and you're not done yet?!?" Without recognizing that THEY are the ones actually culpable for failing to prioritize payments on tech debt vs. "new shiny" feature work, which is typical of Amazon product development culture. I've worked with enough former AWS engineers to know what goes on there, and there's a really good reason why anybody who CAN move on from AWS will happily walk away from their 3rd- and 4th-year stock vest schedules (when the majority of your promised amount of your sign-on RSUs actually starts to vest) to flee to a company that fosters a healthy product development and engineering culture. (Not to mention that, this time, a whole bunch of peoples' Thanksgiving plans were preempted with the demand to get a full investation and post-mortem written up, including the public post, ASAP. Was that really necessary? Couldn't it have waited until next Wednesday or something?)
- roman_sf 6y agoHahaha, well this time Jessy got paged so yeah... the summary got priority over turkey.
- lend000 6y agoEven today I had a few minutes of intermittent networking outages around 9:30am EST (which started on the day of the incident), and compared to other regions, I frequently get timeouts when calling S3 from us-east-1 (although that has been happening since forever).
- ipsocannibal 6y agoSo the cause of outage boils down to not having a metric on total file descriptors with an alarm if usage gets within 10% of the Max and a faulty scaling plan that should of said "for every N backend hosts we add we must add X frontend hosts". One metric and a couple of lines in a wiki could have saved Amazon what is probably millions in outage related costs. One wonders if Amazon retail will start hedging its bets and go multicloud to prevent impacts on the retail customers from AWS LSE's.
- bithavoc 6y agoThey’re calling it an “Event”, title should say “Summary of the Amazon Kinesis Outage...”
- zxcvbn4038 6y agoThey didn’t really discuss their remediation plans but maybe having one fleet of servers for everything isn’t the best setup. I’d love to know which OS setting they ran into. In their defense this is exactly the sort of change that never shows up in testing because the dev and qa environments are always smaller then production. I’m wondering how many people Amazon fired over this incident - that seems to be their goto answer to everything.
- rswail 6y agoMinor detail, but is anyone else irritated by the use of the word "learnings" instead of "lessons"? "To learn" is a verb. Nouning verbs seems to be an unnecessary operationalization.
- jaikant77 6y ago"and it turned out this wasn’t driven by memory pressure. Rather, the new capacity had caused all of the servers in the fleet to exceed the maximum number of threads allowed by an operating system configuration." An auto scaling irony for AWS! We seem to be back to the late 1990s :)
- terom 6y agoUnsurprising to see such outages also tickling bugs/issues in the fallback behavior of dependent services that were intended to tolerate outages. There must be some classic law of cascading failures caused by error handling code :) > Amazon Cognito uses Kinesis Data Streams [...] this information streaming is designed to be best effort. Data is buffered locally, allowing the service to cope with latency or short periods of unavailability of the Kinesis Data Stream service. Unfortunately, the prolonged issue with Kinesis Data Streams triggered a latent bug in this buffering code that caused the Cognito webservers to begin to block on the backlogged Kinesis Data Stream buffers. > And second, Lambda saw impact. Lambda function invocations currently require publishing metric data to CloudWatch as part of invocation. Lambda metric agents are designed to buffer metric data locally for a period of time if CloudWatch is unavailable. Starting at 6:15 AM PST, this buffering of metric data grew to the point that it caused memory contention on the underlying service hosts used for Lambda function invocations, resulting in increased error rates.