3 ms·
I don't agree with all his conclusions. In particular, it seems as though he is saying 'eventual consistency is no practical good' and then beating Dynamo for a
by HenryR 17y ago
I don't agree with all his conclusions. In particular, it seems as though he is saying 'eventual consistency is no practical good' and then beating Dynamo for a few paragraphs with the same stick.
The most often quoted example of Dynamo's use is the shopping cart application on every Amazon page. In the worst case, your shopping cart will mysteriously empty itself. This is a huge pain, and a potential loss for Amazon, but it's not catastrophic in the way that is implied here. Indeed, assuming the liveness of a quorum, the application will read back all conflicting entries for the shopping cart (those that aren't ordered under their vector clock timestamps) and the onus is on it to merge the conflicts. Of course, the shopping cart will take the union of all updates to ensure that nothing is dropped (and therefore some delete operations may be lost).
The key point is that some applications can do without observing a linearisable history, and the interest of this paper is that it explores the design space if you drop that requirement.
I don't understand the post's points about CAP; all three requirements are in tension. Dynamo is unusual in that it is live in the case of a network partition while still maintaining its consistency guarantees.
Similarly - those systems that use chain-replication asynchronously like he describes can still suffer from the same read-old-value-after-it-was-written consistency issue, if the reader jumps between two replicas for consecutive reads. Avoiding that can require synchronous coordination of updates (a la Paxos, e.g.) which is, I think, what the paper is driving at. Otherwise, there are still failure modes which, in order to patch up, require stronger guarantees about liveness of quorums than Dynamo needs.
I understand that Dynamo is no longer used internally at Amazon at scale, so maybe some of the practical points this post makes about the realities of central coordination held water for real deployments. Still, I don't buy the reaction that prioritising availability uber alles and designing a system that does not behave exactly like a strongly-consistent key-value store immediately invalidates it for workloads that have high availability requirements and lower consistency needs.
- gruseom 17y agoDynamo is no longer used internally at Amazon at scale That's certainly relevant. I'd love to know more. Do you know why? What are they doing instead?
- HenryR 17y agoNo - all I know is based on rumour.
- moonpolysoft 17y agoWerner Vogels, their CTO seems to think differently: https://twitter.com/Werner/statuses/5345892061 https://twitter.com/Werner/statuses/5345892061
- jsensarma1 17y agothe argument i have tried to make here is that eventual consistency does not need to be forced in tightly coupled environments that exist in a single data center. i understand that some forms of conflicts are inevitable if one wishes to update concurrently across a WAN. regarding CAP - CA is achievable in a single data center (where as i argue - one cannot tolerate partitions anyway). the problem with Dynamo is that the overheads incurred in tolerating partitions are imposed on environments that do not suffer from partitions. you are right that point in time consistency is no good for partition tolerance. i didn't mean to say that it was. all i wanted to point out was that very few, if any, commercial database deployments do concurrent writes across continents (and do synchronous replication across the same). cross data center replication is used for disaster recovery - and at best analytics - where point in time is just fine.