15 ms·
> But the real advantage of the CRDT approach is that, if you can limit your entire operation set to CRDTs, you can forego consensus algorithms such as Paxos, M
by maxmcd 6y ago
> But the real advantage of the CRDT approach is that, if you can limit your entire operation set to CRDTs, you can forego consensus algorithms such as Paxos, Multi-Paxos, Fast Paxos, Raft, Extended Virtual Synchrony, and anything else along those lines.
Things like this are said a lot, but I don't believe CRDTs provide the same consistency guarantees as Paxos/Raft.
Anyone have thoughts on why CRDTs might be favored over the internals of something like a well written eventually consistent datastore like Cassandra that doesn't use Paxos/Raft?
Is he just saying that your consensus algorithm then just becomes that math of CRDTs and not the careful complexity of a distributed consensus protocol?
(the interviewee then goes on to describe that CRDTs fit their use case well, so maybe that is all he is saying...)
- tracnar 6y agoFrom the article: > TS It's certainly the case that a lot of CRDTs are really complicated, especially in those instances where they represent some sort of complex interrelated state. A classic example would be the CRDTs used for document editing, string interjection, and that sort of thing. But the vast majority of machine-generated data is of the write-once, delete-never, update-never, append-only variety. That's the type of data yielded by the idempotent transactions that occur when a device measures what something looked like at one particular point in time. It's this element of idempotency in machine-generated data that really lends itself to the use of simplistic CRDTs. So for this particular use case, the CRDT is simple and they want to favor availability over consistency. I'm not sure if they provide further guarantees, e.g. from the description in the article it seems that temporary holes in the time series would be allowed if updates would arrive out-of-order. The article also mentions Cassandra, it seems they went for a different design so progress is more quickly visible in the DB (among other reasons I suppose). > What we wanted was a topology that looked similar to consistent hashing databases like DynamoDB or Riak or Cassandra, but we also wanted to make some minor adjustments, and we wanted all of the data types to be CRDTs [conflict-free replicated data types]. We ended up building a CRDT-exclusive database. That radically changes what is possible, specifically around how you make progress writing to the database.
- jandrewrogers 6y agoCRDTs work well for the specific case of machine-generated time-series because your records are samples. There are no complex relationships between records even for the same source. Gaps in the time series are like missing pixels in a picture, there is no specific pixel that is critical as long as you have enough of them to figure out what you are looking at in the moment. This also implies that there are no semantically meaningful update operations on specific records beyond converging values. Platforms like Cassandra can make some of the design choices they do because they don't support high write throughput (relatively). With sensor data models throughput and efficient use of bandwidth is everything, so minimizing the chattiness of ensuring consistency is critical.