13 ms·
A State of Xen – Chaos Monkey and Cassandra
- chuckcode 12y agoAre any of these anti-chaos tools open source or shared with the community? Would love to see more companies that I'm dependent on have this sort of testing and robustness...
- frankchn 12y agoChaos Monkey (and related software) is open source: https://github.com/Netflix/SimianArmy/wiki https://github.com/Netflix/SimianArmy/wiki
- Oculus 12y agoBy trading away C (Consistency), we’ve made a conscious decision to design our applications with eventual consistency in mind How does one go about writing a user-facing application with eventual consistency?
- thethimble 12y agoIf you really need consistency, you can use an external locking mechanism (Zookeeper comes to mind).
- Scaevolus 12y agoIt depends on the application. For Netflix, transient glitches like having outdated play history or catalog entries that take a while to propagate aren't hugely detrimental -- the core value of serving the videos that users request doesn't depend at all on having a consistent view of quickly-changing user data. For more complex applications (think Facebook), there are useful consistency models other than strong consistency, with causal consistency being one of the most promising: http://queue.acm.org/detail.cfm?id=2610533 http://queue.acm.org/detail.cfm?id=2610533
- cordite 12y agoIf they are often watching full episodes or movies, I wouldn't call it very quick. Though progress state would be changing near constantly.
- _delirium 12y agoDo you know whether the linked article is what Facebook is currently using, or a proposal? Facebook is an example of a webapp that fairly often has user-visible weird behavior, like "read" notifications becoming unread again. But I'm not sure if those are glitches, or a result of their consistency model.
- MoOmer 12y agoFacebook wrote Cassandra for such scenarios. https://m.facebook.com/note.php?note_id=24413138919 https://m.facebook.com/note.php?note_id=24413138919
- AYBABTME 12y agoI assume: serving content that doesn't change often (videos) and mostly needs to be available. Recording user scores and aggregating their effect later (last write win means you could update the score a second time and the first time only would be recorded - which is annoying but not end-of-the-world). Likely if they have things that require stronger consistency they use another DB.
- joshrivers 12y ago> rails new todolist ...seriously, though. All of our views are out of date by the time we see them, and all of the user inputs have to be dealt with for races and double-submits. Eventual consistency is just the patterns you already know, but bigger.
- AlisdairO 12y agoEventual consistency is much more complicated than that. Patterns for dealing with nontrivial logical operations (i.e. anything that is bigger than a single 'object') are much easier on consistent DBs than eventually consistent ones. On typical RDBMSs you maintain a per-object version number and foreign keys (possibly also a per-object checksum if you're using more relaxed isolation modes). That's generally enough to prevent races, double submits, and other concurrency artifacts over multi-object operations, and it's pretty easy to implement. The same is not true for eventually consistent systems - checking against a version number can't save you unless all your updates are trivial, single-object ones: some of your updates may succeed while others fail - and there's generally no way to perform an all or nothing operation. In general, there's no one good mechanism for maintaining eventual consistency on nontrivial operations - you have to think (hard) about it on a case by case basis.
- phamilton 12y agoEmbrace idempotence. Being able to perform am operation multiple times with the same result adds a lot of flexibility. For example, using the rails TODO list, if you add a task to "buy milk" and then on another client view the list and don't see "buy milk" there (because consistency hasn't been reached), you might be inclined to think you forgot to save it or something. Then you might enter it again. If this is not idempotent (as is the case on the classic tutorial) you will end up with 2 "buy milk" tasks instead of just one.
- jshen 12y agoand how would you make that idempotent? ;)
- phamilton 12y agoThe key is to consistently derive a primary key. With the same key, last write wins. The second create turns into an update. Making the primary key "buy-milk" would do the job. Getting a consistent key requires a bit of UX but isn't actually that difficult.
- jshen 12y agoWhat if I want to enter buy-milk twice ;)
- ghshephard 12y agoIdempotence can handle that possibility by embedding the ordinal of the purchase you wish to make: buy-first-milk buy-second-milk
- vidarh 12y agoBut now the user is burdened with having to be explicit. This is exactly the type of case where I'd much rather have the app be explicit about it's actions: allow both to be added, and at most show a notice after synchronisation saying "I've noticed you've added two buy-milk; merge them or leave them separate?". I've had nothing but bad experiences with apps that thinks they know best and try to merge records behind my back.
- nemothekid 12y agoNetflix has a couple talks about this, but it mostly comes down to for their specific use case, their users don't care about consistency. For example, for your viewing history you might not entirely care that it is 100% up to date and the last video you watched isn't available. And for the times you do care, most users will simply usually refresh the page (and since, in almost all cases eventual consistency means resolution in seconds rather than instantly, most users will have the correct info after a page refresh). I can't find them right now, but Christos Kalantzis (@chriskalan) has a number of talks on using Cassandra at Netflix w.r.t eventual consistency.
- jbellis 12y agoHere are the slides from Netflix architect Christos Kalantzis 's talk, "Eventual consistency != hopeful consistency:" http://www.slideshare.net/planetcassandra/c-summit-2013-eventual-consistency-hopeful-consistency-by-christos-kalantzis http://www.slideshare.net/planetcassandra/c-summit-2013-even... Video: https://www.youtube.com/watch?v=lwIA8tsDXXE https://www.youtube.com/watch?v=lwIA8tsDXXE But, it's important to note that Cassandra does recognize that sometimes you do need linearizable consistency. Cassandra provides lightweight transactions for this scenario: http://www.datastax.com/dev/blog/lightweight-transactions-in-cassandra-2-0 http://www.datastax.com/dev/blog/lightweight-transactions-in...
- jshen 12y agoI wonder if their chaos system causes network partitions as well as node failures.
- darkr 12y agoNot AFAIK - last time I looked it requires one or more AWS autoscaling groups, and works by semi-randomly terminating instances that match configured criteria.
- sargun 12y agoThe total toolkit they have is called the Simian Army -- https://github.com/Netflix/SimianArmy https://github.com/Netflix/SimianArmy They wrote a blog post about it here: http://techblog.netflix.com/2011/07/netflix-simian-army.html http://techblog.netflix.com/2011/07/netflix-simian-army.html I don't think they have one that introduces network partitions, but inside a datacenter, network partitions are rare.
- jon-wood 12y agoInside a single data center they're rare, but when running on AWS in multiple availability zones and regions they're a fact of life which you should be designing your systems to deal with. Given that it would be great to have automated tooling to simulate them and keep everyone on their toes.
- derek 12y ago> ...but inside a datacenter, network partitions are rare It's ... complicated. http://aphyr.com/posts/288-the-network-is-reliable http://aphyr.com/posts/288-the-network-is-reliable
- sargun 12y agoYeah, there was an excellent ACM article with him, and Bailis. I think if the network starts to partition, or fail in a datacenter, that's time to evacuate the datacenter / AZ. If a handful of machines fail, they should disengage. If more than say, 5% of the machines in the DC are having reachability issues at any given point (in a modern DC that's like ~2000 machines), it's time to shut it down.
- lsc 12y agohm. from: http://xenbits.xen.org/xsa/advisory-108.html http://xenbits.xen.org/xsa/advisory-108.html MITIGATION ========== Running only PV guests will avoid this vulnerability. Did amazon reboot all of it's VMs? or just the HVM VMs? why was neflix running on HVM VMs?
- RyanGWU82 12y agoAmazon rebooted lots of PV guests. Presumably they collocate HVM and PV guests on the same box. If there were any HVM guests on the box, then there could be the possibility of an attack. (I guess they could forcibly kick off the HVM guests, but that wouldn't be very nice.) Why shouldn't Netflix be running on HVM?
- lsc 12y ago>Why shouldn't Netflix be running on HVM? HVM, at least in the past, had a bunch more code that the guest DomU interacts with vs. fully pv guests. This has security implications. Now, my knowledge of HVM is a few years... or more like half a decade out of date, for example, I don't even know how to force a HVM guest to only use PV drivers (which would solve 90% of the problem.) and i know that more and more of this has moved into hardware, so it's possible that what was true five years ago is not true now, but... yeah, I don't let untrusted users on HVM guests for the same reason I don't let untrusted users use pygrub or load untrusted kernels directly.
- timoth 12y agoFor network and storage at least, Amazon says that their HVM guests use PV drivers according to this page: http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance-types.html http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance-... but on this page, they don't say that specifically, but mention the possibility in the PV on HVM section: http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/virtualization_types.html http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/virtualiz... Perhaps/presumably it also depends on the AMI you're using.
- grosskur 12y ago
- Aissen 12y agoIs there a blog where they post about what happened during actual downtime ? Like the one on last September 21th ?
- roncohen 12y agoI'd love to see Netflix let the chaos monkey loose on their PostgreSQL servers.
- ErikRogneby 12y ago2700+ Cassandra nodes! Anyone know how big Facebook's Cassandra is?
- nemothekid 12y agoAFAIK, Facebook no longer uses Cassandra, but apparently Apple has 75,000+ Cassandra nodes.