4 ms·
This is called chaos engineering and many companies built tooling to do exactly this. Netflix pioneered/proselytized it years ago. Since you likely don't just r
by rdoherty 6y ago
This is called chaos engineering and many companies built tooling to do exactly this. Netflix pioneered/proselytized it years ago. Since you likely don't just rely upon AWS services if your app is in AWS, you want something either on your servers themselves or built into whatever low level HTTP wrapper you use. Use that library to do fault injection like high latency, errors, timeouts, etc.
- deleted 6y ago[deleted]
- jedberg 6y agoThis type of service would be a compliment to those techniques, not replace them. Ideally we could have both.
- fizwhiz 6y agoCame here to say this exact thing. There have been a variety of techniques to achieve this - some intrusive to your binaries (i.e. they require embedding specific libraries) and others that are more "external" (ex: tc/iptables). The "real" challenge is not creating chaos but managing it and verifying that your apps are resilient to said chaos.
- tempsolution 6y agoSorry but this is just strangely naive. Yes, you should have integration test setups that allow you to swap out any network dependency with in-memory stubs, with fault-injected proxies, etc. But all of that can never replace the behavior and chaos of a real outage. On top of that, a lot of the more complicated scenarios are immensely difficult to setup and you first also need to know HOW a region can fail. I agree with you on one thing though. AWS, and actually every cloud provider should ship a Chaos Engineering toolbox that comes with ready-to-use, realistic failure simulations that you can use in tests. I.e. drop-in replacements for their SDK clients that then just start running through predefined failure scenarios.
- infogulch 6y agoIt's harder to do chaos engineering if you're not engineering it. What this is really asking for is the service provider to sell chaos engineering as a service (CEAAS?), on the services they provide. I've wanted this kind of thing for testing cloud infrastructure before: you read about various failure states and scenarios you might want to handle from docs but there's no way to trigger them so you just have to hope that they work as described and your code is correct. At the least, let users simulate the effect of the failures that are part of your API. This would be great for testing the pieces of the stack that the provider is responsible for, but you may still want to inject chaos into the part of your stack that you do control.
- segmondy 6y agoNetflix runs on AWS, they are doing chaos engineering quite alright on it.