4 ms·
Has anyone else deployed a Chaos Monkey in production? I can imagine it would be a tough sell to the CEO. :)
by jsingleton 10y ago
Has anyone else deployed a Chaos Monkey in production?
I can imagine it would be a tough sell to the CEO. :)
- klapinat0r 10y agoHow so? The benefits are worth it, and I doubt any CEO will be argue against having fault tolerant code :) You catch bugs, and no one says you can't run Chaos Monkey in staging or a similar environment if it really is a tough sell.
- takeda 10y agoAgreed, it should be hard to explain benefits even to non technical people. It's like doing a fire drill, if you do it frequently when the actual fire happens you will know what to do. Similarly with infrastructure, it might not be good handling rare events, but once these events are not rare you will learn to handle them. The biggest issue IMO is explaining need to make things more resilient. Actually the technical people (mainly developers) might be the biggest obstacle, because it adds more work for them (with no visible benefit to them, because when application fails it's ops who get woken up).
- birdman3131 10y agoThe drawbacks of potentially causing downtime and therefore having the potential to drive away customers as well as obtain an image of unreliability can be much more damaging than not using it in the first place. Customer image means quite a bit.
- solatic 10y agoChaos Monkey is sort of like Advanced Continuous Deployment. Most shops are still struggling with the basics. You cant even think of trying to sell this running to the C-level until you've proven that you can at least walk (automated deployment and rollback). I remember reading years and years ago about bandit algorithms... this kind of ops work is at a level that's found only in a few different companies.
- aaronblohowiak 10y agoYes. The "sell" can be tricky for some people, until your first production issue. Never let a crisis go to waste. Machines go away all the time in the cloud. This tool increases the frequency so you can ensure your system handles it gracefully. Some people believe their system can tolerate this class of failures, but without continuous validation, that is more of a hope than a certainty.
- mlafeldt 10y agoYes, we have. See https://medium.com/production-ready/chaos-monkey-for-fun-and-profit-87e2f343db31 https://medium.com/production-ready/chaos-monkey-for-fun-and... As for chaos being a hard sell: https://medium.com/production-ready/chaos-engineering-a-shift-in-mindset-d8fbfc8c5dc2 https://medium.com/production-ready/chaos-engineering-a-shif...
- relix42 10y agoIn my opinion, it doesn't count unless its in production. Why? Your customers use your production environment, not your test environment. Something will cause loss of an instance for you: * Mistaken termination * AWS retirement (and you missed the email) * Cable trip in the data center * <Something else we can come up with> * <This list goes on> So, vaccinate against the loss of an instance cratering your service. Give your prod environment a booster shot (with Chaos Monkey or something like it) every hour of every day. Then, when anything from the above list happens you're infrastructure handles it gracefully and without intervention. Continued booster shots ensure that this stability continues through config changes, software version changes, OS changes, tooling changes, etc. I think the better question is "Why wouldn't you do this?"