13 ms·
Building resilient services at Prime Video with chaos engineering
- BiteCode_dev 6y agoIs this a case of "netflix did those articles, so we'll do the same so geeks like us too"?
- scarmig 6y agoIf nothing else, the article showcases how to use the recently released AWSSSMChaosRunner[0], which I hadn't heard about. Beyond that, it's good when multiple different high quality blog posts show how to do what's fundamentally the same thing. [0] https://github.com/amzn/awsssmchaosrunner https://github.com/amzn/awsssmchaosrunner
- danellis 6y agoIt does seem like a fad. Kudos to those who can create a new field and profit from it. On the one hand, "chaos engineering" seem a bit like "we don't understand our architecture well enough to know what its failure modes are, so let's just poke it and see what happens," but on the other hand, it seems at least a little bit analogous to fuzzing, which is certainly a technique that yields useful results that would have otherwise been overlooked until it was too late.
- kokooo 6y agoThere are a million different ways a computer can fail. I think we're asking too much of people to be able to know all the pitfalls of every system they create. But also this 'new field' just seems like something we've already been doing just with a different name. You're kind of expected to make sure your system can work if the computer suddenly shuts off, or a dependency is lost, or the network is slow. Have we not been doing this??
- danellis 6y agoIt does seem antithetical to engineering, though. Are there other engineering fields that take this approach?
- eru 6y ago> On the one hand, "chaos engineering" seem a bit like "we don't understand our architecture well enough to know what its failure modes are, [...] That approach seems like a good idea, even if you think you know what the failure modes are.
- mohave529 6y agoMy first instinct was to agree with this, but from my experience it's extremely difficult to properly communicate failure modes 100% of the time across different teams in very large organizations. Dependencies that are fuzzy arise for example when a service A proxies data for client service B from some other service C. It doesn't help that the organization of teams in a company often severs lines of communication between teams who explicitly don't have dependencies but implicitly do. As a result, information gets lost in the process. Having a last line of defense in the form of a "chaos engineering" team may actually be the natural response of large organizations to counter the inherent messiness that is produced as a result of bureaucracy.
- lhoff 6y agoThat and additionally it has implications for the development team as well. Using "chaos engineering" shifts the mindset of the developers. As a developer you now expect things to fail. You know that the "we make it work first and make it resilient later" approach will bite you sooner then later so you think resilience from the first line of code.
- twblalock 6y agoYou don't know what all your failure modes are. You probably think you do, but you don't. That's the point of chaos testing.
- cheeze 6y agoMy guess? Someones manager was like "Hey lets publish this wiki as a blog post." Sure, part of it is trying to seem cool and whatnot, but that's fine with me. The output is the same, more knowledge shared.
- marmshallow 6y agoDefinitely seems like a case of someone that wanted to publish a blog post for some internal kudos, but I'm not complaining. Always good to share knowledge.
- smabie 6y agoIt sounds like complaining to me
- jwilber 6y agoNever thought I’d see anti-article sentiment on HN, and more generally, not sure what sort of response you expect to generate. Why comment at all? To avoid more of the same, I’ll say that the article itself is pretty detailed, but whereas Netflix blog-posts are more general, this one feels hyper-specific. That’s great if you’re tied to aws, but I think the former has longer-lasting + further-reaching utility. Hope they do more.
- BiteCode_dev 6y agoNot an anti-article sentiment, but I dislike the trend of big companies pretending they care about the nerds just because they realize they need them for one particular objective. It's an anti-marketing sentiment.
- jwilber 6y agoThat makes more sense. I think it’s worth framing differently, though. Big companies aren’t writing the doc, per se. Engineers interested in the material are. And often, they want the opportunity to do so - it’s something of a stamp for “expertise” in their favor. Assuming the post is actual quality, it’s a marketing win for the company, a professional/developmental win for the employee, and an informational win for everyone else.
- chromedev 6y agoJust remember anything you purchase on Amazon Prime, you don't technically own it.
- nkristoffersen 6y agoFrom my understanding you don’t technically own most media (music, movies, software) regardless of the format. It’s a license. Even if you purchase a CD or DVD, etc.
- kortilla 6y agoThat misses the point. I have DVDs that I purchased that will work as long as I posses them and they were made before Netflix was even a company. There may technically be a license attached to them, but there is no practical way for any company to revoke my usage of them and the failure mode if any of those companies cease to exist is for them to continue working.
- eru 6y agoOriginal commenter should have written: > Just remember anything you purchase on Amazon Prime, you don't practically own it.
- paxys 6y agoIsn't it exactly the same for a song or movie you purchased from Amazon and downloaded to your hard drive?
- chrisco255 6y agoCan you download prime movies to disk? I thought it was streamed only.
- lmm 6y agoNo, that's nonsense. You do own the medium. More relevantly, the first sale doctrine means copyright does not restrict most of the ownership rights you'd expect to have.
- swayamraina 6y agoThis mentions it cannot be used against AWS lambda. Not sure why?
- vasco 6y agoLambda is fully managed, you can't use SSM to SSH into the underlying hardware because it's not exposed to you.
- fatninja 6y agoFrom the blog, most of the "chaos" is done by Amazon SSM agent running in ec2 instances. Lambda might not have this agent.
- paxys 6y agoI think it's simply that they didn't need that use case, so didn't build for it.
- neo01124 6y agoHi! Author of the article here. The AWSSMChaosRunner approach can't be used for Lambda because of what @vasco said. You can take a look at a different approach here for failure injection in Lambda - https://medium.com/@adhorn/failure-injection-gain-confidence-in-your-serverless-application-ce6c0060f586 https://medium.com/@adhorn/failure-injection-gain-confidence...
- dschuetz 6y agoI'm sure now that Amazon Prime Video won't ever get a decent UI, instead the current will remain horrible, but resilient at least.
- tyingq 6y agoI'm willing to accept a mediocre UI that doesn't start auto playing every video I hover over for more than a second.
- cgriswald 6y agoIndeed. It’s also better than another UI which requires an esoteric feature of a remote control that I don’t even use and offers no way of switching that stupid choice off, preventing me from ever getting more information about the shows it is presenting beyond the title and banner. It’s also better than one that always says I’m logged out when I start it, and then magically logs me in a minute or two later... while I’m in the process of trying to log in. (Or any of the several other ones that require me to log in nearly every time.) Of course calling any of these a “User Interface” is almost laughable. SLOM is a better term. That’s “Small Library Obfuscatory Method” for the uninitiated. Or maybe they’re just trying to save bandwidth by making viewers spend viewing time searching pointlessly through their offerings that are ordered seemingly according to the pattern of pigeon droppings outside their office.
- what_ever 6y agoNetflix now provides an option to disable that.
- MaxBarraclough 6y agoMuting your sound makes that dark pattern somewhat less annoying.
- sabana 6y agoYou must love Netflix then (you can turn autoplay off, but I'm sure you know that, right?)
- fatninja 6y agoWe tried a similar chaos tool in our company built in-house. Simulated most of the scenarios mentioned here using SSM/other scripts. At first everyone was interested and after some time the interest faded. Our problem was lack of visualization across the app ecosystem i.e how will it impact the app ecosystem when a batch of ec2 instances are suddenly spiking on CPU and what will be the impact to end user. Turns out people care only if there is an end user impact and doesn't really care about random anomalies. And to build the capabilities required for measuring the impact + automating the workflow of the actual chaos tests is a lot of work
- neo01124 6y agoStress testing a whole app ecosystem end-end and preventing/mitigating end user impact is generally a part of "gamedays" - https://wa.aws.amazon.com/wat.concept.gameday.en.html https://wa.aws.amazon.com/wat.concept.gameday.en.html. A library like AWSSSMChaosRunner would be a core component of building gameday like capability. But building a full gameday framework is out of the scope of this discussion.
- fxtentacle 6y agoSounds like cloud is getting ever closer to dedicated servers. One would hope that for the 10x price increase over bare metal, someone would abstract away things like the underlying hardware or OS failing. But apparently, no, you have to pay the premium and do all the work. The only use-cases that I could imagine for such a tool on EC2 would be if you either don't use containers, or if you oversubscribe your virtual servers by having higher container limits than what the instance can endure. In the first case, the proper fix is to use containers. Docker can do CPU limiting for you so that one service spiking won't affect its neighbors on the same instance. In the second case, I'd go bare metal and then hardware is so cheap that there's very little temptation to oversubscribe on RAM or CPU.
- neo01124 6y agoHi! Author of the article here. The core concern is not about the capabilities of the compute abstraction being used (bare metal, containers or functions) or testing OS capabilities. The aim is to validate mitigations which are in place to counter turbulent scenarios (For example: massive spike in traffic, network outage, dependency is down, etc). These scenarios generally originate outside the given system. These kind of questions should be asked and systematically validated (quoting the article): * Have you tested how the system behaves when the underlying instances have a sustained CPU spike? * Is the system behavior understood under different stress? * Is there sufficient monitoring? * Have the alarms been validated? * Are there any countermeasures implemented? For example, is auto-scaling set up, and does it behave as expected? Are timeouts and retries appropriate?
- fxtentacle 6y agoI believe we just have a rather different approach here. "Have you tested how the system behaves when the underlying instances have a sustained CPU spike?" Since dedicated boxes are cheap, I'd just buy 5x the CPU resources that I reasonably need and call it a day. If there ever is a more than 5x traffic spike, then docker will prevent it from being a noisy neighbor, so the affected services will just become slower than usual. But even a 10x traffic multiplier would just produce a 2x slowdown, which should be tolerable for most users. I agree that on clouds you want to save costs by only booking what you need. But bare metal, you can usually afford to keep spare capacity around all the time. As such, I wouldn't plan for the system to behave well under stress. I'd try to always have enough resources around so that stress never happens. At the end of the day, this seems like a developer time vs. resource costs trade-off and for most companies, developers are sparse and resources are plentiful, so they'll have a very different trade-off from big FAANG companies. "For example, is auto-scaling set up, and does it behave as expected?" If your system is usually 90% idle, I wonder if you'll ever need that auto-scaling. Also, I'd say my customers can endure it if page load time goes up from 100ms to 200ms. So in my opinion, there is little need for auto-scaling for most companies.
- what_ever 6y agoI have been getting failed to load error almost every week on Prime Video when I try to load a series at dinner time on weekdays (pacific timezone). Never have this issue with Netflix.
- agarzenm 6y agoIt is nice to see these posted. I wonder if the other engineering related sources at AWS get as much attention. They have (in my opinion) an extremely good library of articles in their so called builders library. For example these two below articles. https://aws.amazon.com/builders-library/avoiding-fallback-in-distributed-systems/ https://aws.amazon.com/builders-library/avoiding-fallback-in... https://aws.amazon.com/builders-library/leader-election-in-distributed-systems/ https://aws.amazon.com/builders-library/leader-election-in-d... These topics are extremely hard to solve from scratch yet they are distilled pretty well in the above articles and they include a further reading section. I would implore others to have a gander. I wish the same could be said for their documentation.
- steve_gh 6y agoYou can do Chaos Engineering quite easily with Ruby, because you can raise an exception in a thread from another thread. Many years ago I built a simple tool which allowed you to specify a series of exceptions, and their frequency. The library would run up an extra thread in your process, and simply drop bombs (i.e. raise exceptions according to the required distribution) across the other threads at random. It worked a treat for ensuring high availability in IoT systems
- raverbashing 6y agoOne thing I wonder (and find difficult to simulate) are failures in external services. Sure you can unit test your function with a 500 for example, but you never know in which ways the function/library can fail (Not to mention the cases where it says everything worked but it didn't)
- neo01124 6y agoThe article does talks about how to inject latency or packet-loss into calls to particular external services. This should help you test many service->dependency failure scenarios around retries, timeouts and circuit-breakers. Injecting specific error codes or exceptions is a bit more complicated but it is possible with other approaches, for example: Chaos toolkit.
- stuaxo 6y agoUsing the Prime video app on my Android TV it seems like some basics aren't really attended to. The aspect ratio on the thumbnail for film I watched was stretched the other day, which is not the lack of attention to detail I'd expect from such a rich company. Apart from this, it just isn't a smooth experience to navigate or search in. As usual in development nobody is paying any attention to responsiveness. There is no reason this app shouldn't work as well as the youtube app in these respects.
- benbristow 6y agoAmazon products are never polished or fully fleshed out. You can spend a few minutes on their flagship site and find weird design choices or basic CSS mistakes. A shame considering the amount of money they have. You'd expect a bit more, but then again Amazon are in the business of scale, quantity over quality. At least it keeps them from having a monopoly on literally everything.
- simonebrunozzi 6y agoUptime of essentially 100%. Hard to criticize when they accomplished something so rare in tech.
- benbristow 6y agoJust because something is up doesn’t mean it’s good. There’s an innuendo in there somewhere...
- 1DRACOSEA8 6y agoI see what you did there, Karen.
- SEJeff 6y agolaughs in us-east-1
- 6y ago
- neo01124 6y agoAuthor of the article here. Please take a look at the underlying library here (AWSSSMChaosRunner) - https://github.com/amzn/awsssmchaosrunner https://github.com/amzn/awsssmchaosrunner
- naringas 6y agowhat about disabling the short advertisement preview played before what I actually want to watch? sure I can skip it, but why should I have to?
- Havoc 6y ago>The key to chaos engineering is injecting failure in a controlled manner. Doesn’t that sorta defeat the point of “chaos” a bit?
- neo01124 6y agoThat is more about the "engineering" bit.
- rootedbox 6y agoIf I use any prime video apps.. they think I'm in Canada(which I'm not).. and only gives me Canadian selections(which are horrible).. if I use prime video in a browser.. it works fine.. I have no idea why it does this and support has no clue.
- _xerces_ 6y agoAlways start out streaming in really, really low quality on the Firestick and gets stuck there despite us having gigabit fiber. Usually have to stop playing, exit and then go back in again then it works and plays in HD. Never happens on Netflix or Hulu or any other app, just Prime Video.