4 ms·
[Disclaimer: I work as a software engineer at Amazon (opinions my own, obvs)] The chaos aspect of this would certainly increase the evolutionary pressure on yo
by davidrupp 6y ago
[Disclaimer: I work as a software engineer at Amazon (opinions my own, obvs)]
The chaos aspect of this would certainly increase the evolutionary pressure on your systems to get better. You would need really good visibility into what exactly was going on at the time your stuff fell over, so you could know what combination(s) to guard against next time. But there is definitely a class of problems this would help you discover and solve.
The problem with the testing aspect, though, is that test failures are most helpful when they're deterministic. If you could dictate the type, number, and sequence of specific failures, then write tests (and corresponding code) that help make your system resilient to that combination, that would definitely be useful. It seems like "us-fail-1" would be more helpful for organic discovery of failure conditions, less so for the testing of specific conditions.
- cogman10 6y ago> The problem with the testing aspect, though, is that test failures are most helpful when they're deterministic. Let's not let `perfect` get in the way of `good`. Certainly having a 100% traceable system would be ideal, most systems are not that. There is still a TON of low hanging and easy to find issues that would automatically fall out of a system of random fails. Even if engineers have to spend some time figuring out what the hell is going on, it would overall improve their system because it would shine a bright shiny flashlight on the system to let them know "Hey, something is rotten here". From there, more deterministic tests and better tracing can be added.
- spaetzleesser 6y ago"The chaos aspect of this would certainly increase the evolutionary pressure on your systems to get better. You would need really good visibility into what exactly was going on at the time your stuff fell over, so you could know what combination(s) to guard against next time. But there is definitely a class of problems this would help you discover and solve. " Error conditions you already know about are easy to test and code against. I would guess most system failures come from conditions nobody expected. When you have a randomly failing system you can discover them. Fixing problems and testing for them then will be easy in comparisons. For example: Years ago I worked on a video streaming solution. Some day we got a device that would garble our traffic randomly, slow it down and so on. things started crashing left and right. Within a month we squashed hundreds of bugs and had a rock solid system that basically impossible to crash. I always wondered about AWS for other cloud systems how you could prepare for problems. It's hard to predict what can fail in what ways and you can't really force error conditions. I really like this idea of a cloud where everything breaks.