7 ms·
Your statement about data propagation in that linked comment is at least misleading. A write at quorum will always be visible instantly to a read at quorum.
by _benedict 4y ago
Your statement about data propagation in that linked comment is at least misleading. A write at quorum will always be visible instantly to a read at quorum.
- Luker88 4y agoI only wrote that quorum is not a transaction (no transactions in cassandra), and that is not a consensus. While quorum looks a lot like consensus, it is not since what is returned to the client is the latest timestamp. Different nodes could have and return different data. So Quorum write+local_quorum read might fail even if you are in the datacenter that accepted the write. Quorum is also on the total copies of the data, not a quorum of DC, so in certain (weird) multi-DC setups you could have a quorum in a single DC. In general though I think that the data consistency options of cassandra (quorum/local/N) are a good idea but underdeveloped. Anyway my point was: cassandra has too many pitfalls and eliminating those restricts the use case by a lot more than people realize. Plus the naming of all features look designed to trick you into thinking soewthing else
- _benedict 4y agoNo, you wrote that you have to write and wait several minutes for data to replicate. This is straight-up false. Yes, if you mix consistency levels incorrectly you will do it wrong and maybe get stale data, but that is a different criticism. I agree that it is easy for unsophisticated users to incorrectly use consistency levels in complex topologies, and I hope we will introduce mechanisms to prevent users making such mistakes in future. But that was not your claim, and in my experience users do understand consistency levels just fine. There are lots of valid criticisms to point at various use cases with Cassandra, but this was just incorrect.
- skyde 4y agoit’s worse than waiting actually. Because client A might write value « 1 » and after client B write value « 2 ». Then client C can read Value 2 for 5 minute and then cassandra internal read-repair eventually put value 1 on all replica and value 2 is lost forever.
- _benedict 4y agoNo, that is literally impossible. If value 2 is newer than value 1, value 1 will never overwrite value 2. If a client reads value 2 at QUORUM then it will always be seen by all future queries.
- skyde 4y agoClient À write value 1 on 2 of 3 node with timestamp=3 then later client B write value 2 on 2 of the 3 node with timestamp=1 then read repair happen and (value=1 timestamp=3) is written to all 3 nodes. this is only one of many scenarios where this stupid design fail.
- _benedict 4y agoClient B’s write will only be successfully read from 1 of the nodes it wrote to, and read-repair only runs on QUORUM reads. So, no, it will not be possible to read value 2 for 5m - it will never be visible to operations at QUORUM. This would have been a valid criticism of LWW (and there are other more contrived examples), but I think (or hope) this is an explicit trade off made by anyone using Cassandra in eventual consistency mode. There are strategies to prevent this being a problem for workloads where it matters, some discussed elsewhere in the thread.
- skyde 4y agoSince "client B write" was successfully written on 2 of the 3 nodes. Any read request that read only from 2 of the 3 nodes instead of using quorum read, will be able to see what Client B wrote until it magically disappears. Quorum Read-repair is only one reason for why the value would randomly disappear. Another one is periodic anti Entropy repair!
- _benedict 4y agoNo, it won’t. Successfully written does not mean what you think it means. If there is “newer” (by timestamp) data on disk or in memtable then it will not be returned to a client, regardless of which order that data arrived. It is unlikely even to be written to disk (except the commit log). Since at least one of those nodes has the “newer” value, only one node can serve this “older” value
- pkolaczk 4y ago> Different nodes could have and return different data. Only if they miss some writes, and eventually they will converge. But if you do quorum writes and quorum reads (or local quorum W + local quorum R), this guarantees you'll read from at last one node that received all the writes issued before the read, so you get the converged value immediately, regardless of which node you ask. All nodes will eventually agree on the value, because timestamps are assigned by the coordinator or by the application, not at the replica. A single write will get the same timestamp across all replicas. Incorrect timestamps can cause a different problem - a write that happened at physical time T2 > T1 might be considered to be older than T1 by the cluster, if it was accepted by the coordinator whose clock was set in the past. Such write might simply not take any effect, as old updates would be considered newer. However, again, the resolution would be consistent on all replicas once they get all updates.
- skyde 4y agoeventually they will converge on a « random » value! that’s why we say it’s write quorum is not consensus it’s offer very weak guarantee
- pkolaczk 4y agoNo, they will converge on a value with higher timestamp. And timestamps are controllable by your app, so you can guarantee they are monotonic.
- skyde 4y agotimestamp is coming from the operating system where the client library is running. so unless all your write request are issued by the same machine then you are guaranteed to face the problem where the client machine clocks are not in sync. just comparing timestamp is obviously a design from someone that didn’t review the academic literature on distributed transaction and consensus
- wpaladin 4y ago
- GauntletWizard 4y agoAnd nobody does reads at quorum, they're slow. People do reads at closest, even when they really need quorum.
- mping 4y agoUnless they want "read your writes", otherwise bugs start appearing "why isn't the data i just put showing"?
- deleted 4y ago[deleted]
- YZF 4y agoI don't think that's true. Can't speak for everybody but for the stuff I worked on. Quorum reads at RF=3 are only twice as slow as reading a single node and it's pretty practical, if you care about it, to write and read using quorum writes/reads. It's true there's many applications where you're ok with some time not reading your latest data, and for those you can get somewhere better performance/latency for a given scale by going CL=ONE...
- staticassertion 4y agoWhat happens if you send the 'write' to 3 nodes but only 2 are up?
- alanfranz 4y agoIt depends on the consistency level you set for writing. If you have 3 nodes, set quorum cl, the write will succeed, because quorum of 3 is at least 2. If you explicitly require cl of three, and only two nodes answer, write will fail.
- Luker88 4y ago"write will fail"...but the data will still be there (you can SELECT it normally) and will be replicated to the node that was down once it is up again. Unless the node was down too much and could not fully catch up before a set time (DB TTL if I remember correctly), in which case the data might be propagated or not, and old deleted data might come up again and be repropagated by a cluster repair. So much fun to maintain
- ChrisMarshallNY 4y ago> and old deleted data might come up again and be repropagated That pretty much describes iCloud. Argh! Zombies! iCloud has terrible syncing. Here’s an example. Synced bookmarks in Safari: I use three devices, regularly; my laptop (really a desktop, most of the time), my iPad (I’m on it, now), and my iPhone. On any one of these devices, I may choose to “favorite” a page, and add it to a fairly extensive hierarchy of bookmarks, that I prefer to keep in a Bookmarks Bar. Each folder can have a lot of bookmarks. Most are ones that I hardly ever need to use, and I generally access them via a search. I like to keep the “active” ones at the top of the folder. This is especially important for my iPhone, which has an extremely limited screen (it’s an iPhone 13 Mini -Alas, poor Mini. I knew him well). The damn bookmarks keep changing order. If I drag one to the top of the menu, I do that, because it’s the most important one, and I don’t want to scroll the screen. The issue is that the order of bookmarks changes, between devices. In fact, I just noticed that a bookmark that I dragged to the top of one of my folders, yesterday, is now back down, several notches. Don’t get me started on deleting contacts, or syncing Messages. I assume that this is a symptom of DB dysfunction.