6 ms·
Introducing Chaos Engineering
- ellysetaylor21 12y agoLol at monkey picture, full rambo style :p
- dirtyaura 12y agoI assume that other big internet companies also practice chaos engineering under a different name, but having such a name for the job is awesome. It highlights the difference to traditional stress testing. Names have surprising power. Growth Hacker was a bit annoying but very effective title trend and it helped to communicate the different approach to traditional marketing efforts. I think Chaos Engineer has the same potential.
- numlocked 12y agoYou're probably right re:other companies doing similar things. Certainly Google does something similar via their "disaster days": http://queue.acm.org/detail.cfm?id=2371516 http://queue.acm.org/detail.cfm?id=2371516
- herge 12y agoMan, I really have the impression that my video-on-demand service has a better understanding of ensuring availability and risk-management than either my bank or any government service. Could you imagine say a utility company creating a title called 'Chaos Engineer'.
- hortonew 12y agoIt's so true. The entertainment/consumer shopping industries seems to have surpassed banks/governments in availability. It's also pretty ridiculous knowing that some of these companies are tasked with keeping our data/money safe, and yet they still don't allow extremely complex passwords.
- tracker1 12y agoScaling financial transactions requires real sharding, but you can't eliminate ACID from the mix... certain functionality just doesn't work like that. In the end, one system needs to be responsible for a given account... best case you could have is only some accounts/users are affected. It's a different kind of problem. Facebook, for example uses Cassandra with a very wide distribution of nodes, with an immediate ack on receive... you can't do that with a bank... transactional payment systems won't tolerate it. You could do a few things to alleviate this issue... read-only replicas, etc... even then it's a matter of failing soft (disabling only those systems/services that are unavailable instead of the whole thing).
- deleted 12y ago[deleted]
- Kronopath 12y agoCan you imagine how difficult that would be to justify in a bank or government service? It would be very hard to convince someone in a position of power at an institution like that that deliberately introducing random failures and problems into your production services would be a good idea. Even if it is a good idea, it would be hard to make that argument. If Netflix's system fails, the worst thing that could happen is you don't get to watch Orange is the New Black. If your bank's system fails, you may lose payments, fail to pay your bills on time, or worse.
- opendais 12y agoYou run a staging environment to mirror production [with fewer nodes per PoP] and throw things at it. The only real difference is the load on staging would be generated artificially [e.g. ACH, credit card processing] using faked staging-only accounts. You even have a fake banking website for security audits that is a clone of production with the fake accounts. No risk to production and probably 80% of the benefits. Of course its probably 125% of the effort.
- diminoten 12y agoShow me a staging environment that's as robust as a production environment...
- deeviant 12y agoMy last company I worked at, they had like 30% of the staging environment running production stuff. They also had only one deploy target, which is drum roll production, which meant all of staging AND some of Dev/QA was live and could interact with the production environment. Also fun to see the production load at 300% of expected, burning, only to find out some event, data roll up, query, whatever was being run in duplicate on 10 different machines because somebody forget to to manually edit the configs after they rolled out a new version to staging/QA. Although I think this would qualify as "chaos engineering", I don't think it fits with what netflix is going for. Yeah, ok that had nothing to do with the OP, sorry, I just had to vent.
- deleted 12y ago[deleted]
- angersock 12y ago4chan has better availability and higher traffic than healthcare.gov. Think about that.
- elblanco 12y agoI feel like there's lots of space for b2b disruption in banking. Banks really shouldn't be in the business of building secure and robust software. They don't manufacture their own vaults either.
- kperry 12y agoNetflix puts out some great articles about architecture in the cloud. Auto-scaling, chaos monkey, and how they handle 'steal-time.' Does anyone know of any other company that publishes so much about cloud architecture? This is great stuff!
- frozenport 12y agoMost companies their size would use their own servers instead of the cloud.
- kooshball 12y agoUsing AWS does not magically give you a HA infrastructure when you have a complicated service oriented architecture like Netflix. All the stuff mentioned here are still relevant even if they're running their own DC.
- mkoller 12y agoCheck out the CloudFlare blog http://blog.cloudflare.com/ http://blog.cloudflare.com/ Here is a good read ( http://blog.cloudflare.com/technical-details-behind-a-400gbps-ntp-amplification-ddos-attack http://blog.cloudflare.com/technical-details-behind-a-400gbp... )
- chton 12y agoIt's great to see Netflix taking disaster recovery and chaos mitigation seriously. Learning to work with constant failure is one the biggest challenges to anyone working with distributed systems and scale, and concepts like the Chaos Monkey help enormously. I hope other companies follow suit, and soon.
- tomwilde 12y agoSomeone calling themselves 'the chaos commander' who uses multi-regional active-active jargon is looking for chaos engineers. Nope.
- diminoten 12y agoI wonder if Netflix has ever come out with some kind of, "So you want to get into chaos engineering, eh?" kind of article that explains the basics and some pitfalls/things to look out for.
- q_ 12y agoI gave a talk at pagerduty about how this has been done for the last few years: https://blog.pagerduty.com/2014/03/injecting-failure-at-netflix-staying-reliable-for-40-million-customers/ https://blog.pagerduty.com/2014/03/injecting-failure-at-netf...
- kperry 12y agoNice! Thanks for sharing man!
- kylekampy 12y agoHow does one go about becoming a chaos engineer? I imagine up to this point it is a field one falls into accidentally and gains experience over time. I can imagine it becoming a topic taught at a college level in the near future.
- herge 12y agoNot really, you probably start with an entry-level sysadmin/devops position, and work your way from fire-fighting incident to firefighting incident. You just need an appreciation that any down-time in a system is a symptom of larger problems, and the will to identify (and reproduce as chaos!) those problems.