11 ms·
Cassandra is not row level consistent
- deleted 10y ago[deleted]
- supergirl 10y agoCouldn't explain the problem in a less cringy way? Sounds like a bug to me. So file a bug report?
- detaro 10y agoI think not strictly a bug, just surprising behavior. They linked to the documentation that shows that it is "expected" to work like that, and to two requests for features to help mitigate this issue.
- supergirl 10y agobug/feature whatever. CASSANDRA-6123 is exactly addressing row level consistency. but here maybe they found another case.
- refresh-creds 10y agoIf it's not obtained synchronously what does the lock achieve?
- helper 10y agoAs a long time Cassandra user its easy to forget that some of Cassandra's semantics will be surprising to new users. That being said, if you are considering adopting an AP database it really is important for you to know the details about how write conflicts get resolved. This is perhaps the biggest difference between Cassandra other databases like Riak and ought to be part of your decision making process instead of a surprise you run into later. That being said, using Cassandra for distributed locks is a terrible idea. I can't think of any way in which Cassandra would be better than using {Zookeeper,etcd,consul}. Trying to force a database to do something it really isn't designed for will almost always lead to disappointment (and often resentment) of said database.
- chungtg 10y agoDatastax seems to think that it is a good idea: http://www.datastax.com/dev/blog/consensus-on-cassandra http://www.datastax.com/dev/blog/consensus-on-cassandra
- user5994461 10y agoThat piece was written in 2014, which was a different time ;) Plus, if we were to take the vendor words at face value, we'd all be using Docker and MongoDB in production.
- flurdy 10y ago:) My last client (very, very big client) for the past year ran nearly all their production systems with Docker and MongoDb... Never had any problems with Docker. MongoDB on the other hand... Suffice to say it is really fragile when spread across multiple datacentres, which is probably not a surprise to HN.
- djsumdog 10y agoI work for a large company that used docker in production. A lot of companies are. There are a lot of big names using Kubernetes, Mesos or some other container orchestration system.
- ddorian43 10y agoJust like many companies use mongodb and nanoservices.
- pkolaczk 10y agoDetails matter. All updates in the Datastax post are protected by LWT. That code is correct. The OP's code was wrong, because he was mixing non-LWT updates.
- convolvatron 10y agocassandra never claimed to be a consistent distributed database. its really quite sad that someone had to find that out the hard way.
- lerax 10y agoCassandra is a piece of shit. Negative that comment how you wants, this never change the initial info. I used that BS on a startup on data science on the crawling info level and was on the worst experiences with db I had. CASSANDRA IS JUST PAIN.
- sigy 10y agoSome details besides ad-hominem might actually be helpful. When you make statements like this it just looks like trolling.
- vesak 10y agoCassandra is not a person, so that was not ad hominem.
- stonewhite 10y ago100,000 nodes on Apple and 2,500 nodes on Netflix would like to have a word with you.
- shin_lao 10y agoApple no longer uses Cassandra.
- jeromatron 10y agoI'm not sure how you got that impression but they still have huge deployments of Cassandra.
- zzzcpan 10y agoShouldn't Cassandra be using Lamport timestamps or even vector clocks there? Relying on timer and its resolution sounds strange for a database, especially a distributed one.
- supergirl 10y agohttp://www.datastax.com/dev/blog/why-cassandra-doesnt-need-vector-clocks http://www.datastax.com/dev/blog/why-cassandra-doesnt-need-v...
- zzzcpan 10y ago"Conversely, if there are concurrent changes to a single field, only one will be retained, which is also what we want." They are wrong though. As HN submission illustrates, people want some order and eventual consistency, not a rule to select a single field during concurrent changes. And this is where a Lamport timestamp could help.
- sigy 10y agoNo matter what scheme you use, it is still a rule. Lamport timestamps, whatever are just different ways of resolving a conflict. In general, the fact that you have designed your system to have this conflict in the first place is an easier problem to solve than having users understand more complex resolution methods -- especially those that push a callback into the app design rather than actually just solving the basic problem in a pragmatic way.
- parenthephobia 10y agoCassandra's timestamps can be specified by the client: it's possible to use any integer as your timestamp, such as Lamport timestamps, or even atomic distributed counters (possibly using Cassandra's own counters). Vector clocks are still out, though. Edit: Actually, you can't do Lamport timestamps because you can't query the current value of a timestamp. Re-edit: That's wrong. I shouldn't believe any old blog I find. (I'm leaving it in the comment because it was quoted in a reply.)
- sigy 10y agoI see this as a basic misunderstanding of how LWT works. If you want to ensure serializable operations, then you need to use LWT with preconditions that ensure serializable operations. Even better, stop trying to emulate the old and tired distributed lock methods that have been proven over and over again to be insufficient.
- neeleshs 10y agoAnyone has experience around this on HBase, a CP database?
- linuxhansl 10y agoApache HBase committer here. HBase is strictly CP (except for its geo-replication, and optional timeline-consistent region replicas). It uses MVCC for row "transactions" to always keep rows consistent. HBase also has checkAndPut and checkAndDelete primitives, which are atomic, as well as Increment and Append, which are atomic and serializable. http://hadoop-hbase.blogspot.com/2012/03/acid-in-hbase.html http://hadoop-hbase.blogspot.com/2012/03/acid-in-hbase.html explains it fairly well. Together with Apache Phoenix you have full multi-row transactions, but they come with a price obviously.
- beefsack 10y agoThe post was quite interesting, but I find image macros and GIFs really distracting in technical writings.
- steeve 10y agoUsing Cassandra as a lock is a terrible, terrible idea.
- Animats 10y agoThis: INSERT INTO locks (id, lock, revision) VALUES ('Tom', true, 1) IF NOT EXISTS USING TTL 20; looks like a race condition. The same problem comes up in SQL databases - you can't lock a row that doesn't exist yet. If you write, in SQL: BEGIN TRANSACTION SELECT FROM locks WHERE id = "Tom" AND lock = true AND revision = 1; -- if no records returned INSERT INTO LOCKS locks (id, lock, revision) VALUES ('Tom', true, 1) COMMIT you have a race condition. If two threads make that identical request near-simultaneously, both get a no-find from the select, and both do the INSERT. SELECT doesn't lock rows that don't exist. The usual solution in SQL is to use UNIQUE indices which will cause an INSERT to fail if the record about to be inserted already exists. I ran into this reassembling SMS message fragments, where I wanted to detect that all the parts had come in. The right answer was to do an INSERT for each new fragment, then COMMIT, then do a SELECT to see if all the fragments of a message were in. Doing the SELECT first produced a race condition.
- dhd415 10y agoSome SQL databases (e.g., SQL Server and PostgreSQL) offer key-range locking which, along with a serializable isolation level, will prevent this race condition.
- chillaxtian 10y agothis is not a race in cassandra. IF NOT EXISTS causes replicas to agree on the result using PAXOS, and only responds with success if a certain number of replicas concur + write the transaction to disk. i believe it is a quorum for SERIAL consistency level, and local quorum (quorum of replicas in the local DC) for LOCAL_SERIAL.
- linuxhansl 10y agoON DUPLICATE KEY is another construct that is used in some DB (MySQL or HBase with Apache Phoenix and others).
- im_down_w_otp 10y agoIf you're not writing purely immutable data or can't 100% guarantee a serialized reader/writer, then you're just looking for trouble with Cassandra.
- cbsmith 10y agoIt feels like it is 2013 all over again: https://aphyr.com/posts/294-jepsen-cassandra https://aphyr.com/posts/294-jepsen-cassandra
- aartur 10y agoAnd there's another surprise waiting to be discovered. The execution of a LWT is not guaranteed to return applied/not-applied response [1]. It can raise a WriteTimeout exception that means "I don't know if applied". It looks like in that case it can be worked around by inserting a UUID and in case of a WriteTimeout reading the UUID using SERIAL consistency and checking if it's the inserted UUID. But generally this limitation of LWTs makes implementing some algorithms impossible, e.g. you can't implement a 100% reliable counter. [1] https://issues.apache.org/jira/browse/CASSANDRA-9328 https://issues.apache.org/jira/browse/CASSANDRA-9328
- agentgt 10y agoI guess I'm old or just not hip (most likely both) but I had to google WAT (I know WTF but WAT ... never seen it). Even now I'm still not sure but I presume WAT = what!
- pushrax 10y agoIt is a particular form of "what" that is used when dumbstruck.
- ashmud 10y agoIn case this is not the explanation you found: http://knowyourmeme.com/memes/wat http://knowyourmeme.com/memes/wat
- hesselink 10y agoI know it from this presentation: https://www.youtube.com/watch?v=AU2Rhq5eWa4 https://www.youtube.com/watch?v=AU2Rhq5eWa4
- second_picard 10y agoHazelcast has a distributed lock (see http://docs.hazelcast.org/docs/3.7/manual/html-single/index.html#lock http://docs.hazelcast.org/docs/3.7/manual/html-single/index....) and I've used for more than a year to synchronize jobs across a cluster.
- jbellis 10y agoCassandra developer here. Lots of comments here about how Cassandra is AP so of course you get inconsistent (non-serializable) results. This is true, to a point. I'm firmly convinced that AP is a better way to build distributed systems for fault tolerance, performance, and simplicity. But it's incredibly useful to be able to "opt in" to CP for pieces of the application as needed. That's what Cassandra's lightweight transactions (LWT) are for, and that's what the authors of this piece used. However! Fundamentally, mixing serializable (LWT) and non-serializable (plain UPDATE) ops will produce unpredictable results and that's what bit them here. Basically the same as if you marked half the accesses to a concurrently-updated Java variable with "synchronized" and left it off of the other half as an "optimization." Don't take shortcuts and you won't get burned.
- webmaven 10y agoSo, since their use of UPDATE is problematic, what is the correct way to release the locks?
- pkolaczk 10y agoDELETE ... IF EXISTS DELETE ... IF <condition> UPDATE ... IF <condition>
- verifex 10y agohaha, I was just going to ask you what you thought of this, and I'm glad to see you responded already! Yes! :)
- 4ad 10y agoThis is too subtle. Incompatible operations should be rejected. Reliance on the programmer to correctly use systems that don't enforce consistency to write consistent transactions is a bad strategy.
- sorkin2 10y agoThis kind of nannying can be really expensive to do correctly when trying to build high performance systems -- to the point where it doesn't make sense to punish everyone just because someone who didn't read the manual MIGHT misuse the product. It's not at all unreasonable to mix transactions with different guarantees if they never touch the same data, and tracking that accurately enough without pissing off your customers with performance needs seems like a fool's errand.
- known 10y agoWith the Oracle/PostgreSQL, readers never wait for writers and writers never wait for readers http://philip.greenspun.com/sql/your-own-rdbms.html http://philip.greenspun.com/sql/your-own-rdbms.html using underlying locking mechanism http://www.beej.us/guide/bgipc/output/html/singlepage/bgipc.html#flocking http://www.beej.us/guide/bgipc/output/html/singlepage/bgipc....
- known 10y agoIts' better to implement https://en.wikipedia.org/wiki/Priority_inversion https://en.wikipedia.org/wiki/Priority_inversion in all https://en.wikipedia.org/wiki/NoSQL#Types_and_examples_of_NoSQL_databases https://en.wikipedia.org/wiki/NoSQL#Types_and_examples_of_No...
- seanparsons 10y agoI railed against CQL right from the start and it's precisely because of this kind of thing. Imitating SQL has the side effect of setting certain expectations and drags a certain mental model along with it.
- mydpy 10y agoAs others have noted, this blog uses an approach to data modeling that is considered an anti pattern for an AP data store like Cassandra.
- leastangle 10y agoAuthor here. Great discussion around the CAP theorem but it misses the point. AP vs CP / Cassandra being AP is not relevant to this particular problem: 1) This is not a distributed systems corner case. You will run into this if you are running Cassandra on a single node. A node should be able to guarantee consistency internally during normal operation. If it is not able to do that, there is something wrong with the system. 2) This is a case where queries are being send from the same process/thread and go to exactly the same nodes. Attach a simple, monotonically increasing query counter to each call and you can easily serialize it on the other side.