3 ms·
I had a somewhat similar experience. This, especially, is what my team also found to be true: "The only model that makes sense in cassandra is where you insert
by pixelmonkey 7y ago
I had a somewhat similar experience.
This, especially, is what my team also found to be true: "The only model that makes sense in cassandra is where you insert data, wait for minutes for the propagation, hope it was propagated, then read it, while knowing either the primary/clustering key or the time of the last read data."
We use Cassandra in production at my company, but only after much scope reduction to the use case.
We needed a distributed data store for real-time writes and delayed reads for a 24x7 high-throughput shared-nothing highly-parallel distributed processing plant (implemented atop Apache Storm). Cassandra fit the bill because writes could come from any of 500+ parallel processes on 30+ physical nodes (and growing), potentially across AZs in AWS. And because the writes were a form of "distributed grouping" where we'd later look up data by row key.
At first, we tried to also use Cassandra for counters, but that had loads of problems. So we cut that. We briefly looked at LWTs and determined they were a non-starter.
Then we had some reasonable cost/storage constraints, and discovered that Cassandra's storage format with its "native" map/list types left a lot to be desired. We had some client performance requirements, and found the same thing that hosed storage also hosed client performance. So then we did a lot of performance hacks and whittled down our use case to just the "grouping" one described above, and even still, hit a thorny production issue around "effective wide row limits".
I wrote up a number of these issues in this detailed blog post.
https://blog.parse.ly/post/1928/cass/ https://blog.parse.ly/post/1928/cass/
The good news about Cassandra is that it has been a rock-solid production cluster for us, given the whittled-down use case and performance enhancements. That is, operationally, it really is nice to have a system that thoughtfully and correctly handles data sharding, node zombification, and node failure. That's really Cassandra's legacy, IMO: that it is possible to build a distributed, masterless cluster system that lets your engineers and ops team sleep at night. But I think the quirky/limited data model may hold the project back long-term, when it comes to multi-decade popularity.