4 ms·
Genuinely curious: when you say "it brings enormous benefits to a huge class of queries that are common in timeseries", what are you referring to, exactly? I r
by pixelmonkey 10y ago
Genuinely curious: when you say "it brings enormous benefits to a huge class of queries that are common in timeseries", what are you referring to, exactly?
I run Cassandra in production and I love its operational simplicity, scale-out design, and write performance. But I think its support for time series is perhaps over-hyped. To me, it seems the only queries you can run in Cassandra is a key lookup (partition key row get) and a column slice (partition key row get filtered by an ordered range of columns). This allows for a certain time series use case e.g. where each row represents exactly one series, and where the only thing you want to do with a series is to get its raw values. But it doesn't allow for many of the things I personally think of when I think about "time series queries", e.g. resampling, aggregates, rollups, and the like.
- vegabook 10y agoI am referring to anything that resembles a range query, ie, where you require a bunch of contiguous information queried on a single key. Think "give me all of this person's chat entries from x time to y time", or indeed "give me all this topic's comment entries from x time to y time" (but not both - only one of the above would be efficiently stored - you decide which it would be). Cassandra, as you know, forces a certain amount of "low level awareness" requirement on the programmer because to tap into its uniqueness, you need to know how you will query stuff, so that Cassandra will ensure that the most common range queries are contiguously stored in rows. All other databases hide the on-disk storage order from you in an abstraction, and you can find atomisation causing inefficiency. Cassandra forces you to think about it, and in return, guarantees contiguous storage order on disk along one of your keys so that along that key, retrieval is lightning fast as it requires only one pass. Basically, both spinning disks but also SSDs, are in essence, 1d media (ie, a lot in common with tape) in the sense that along one dimension you can read stuff massively fast, but as soon as you need to seek (ie start using dimension 2), even on an SSD, your performance dramatically declines. Cassandra forces you to think about your queries so that they will be "aligned" along the most efficient direction on disk. Now agreed that if your queries cannot be aligned along said direction, then Cassandra drops to being no better than all the others, and penalises you with some complexity. That includes some examples of aggregrates, resampling etc (though I would argue that the order of magnitude contiguous read still helps these). Some of this can be mitigated with denormalisation ie: storing stuff more than once, in transposed or sub-sampled orders, something that relational DB purists will hate, with some justification (potential for inconsistency). FWIW Riak TS sounds promising with automatic "blob" style storage etc and resampling capabilities which might take Cassandra on quite explicity and in a higher level, more convenient way. I am about to evaluate it because I agree with you that the resampling capability in particular could be better supported in Cassandra, though ultimately, both databases will still be limited by the underlying D1 v D2 "contiguous v seek" capabilities of the storage so I'm not expecting miracles from Riak. By the way, I'm not even touching on Cassandra's scale-out ease. More perf needed? Literally just add boxes though it would be unfair not to comment on the cost of this, which is Cassandra's node-level consistency tradeoffs for very recently added data, and which is, if I recall correctly, why Facebook went to Hbase. You can force consistency at the query level, but performance can suffer.
- _halgari 10y agoAfter some truely horrific experiences with Riak K/V, especially combined with Riak Solr, I won't touch anything from Basho with a ten foot pole. Not sure what's going on over there, but the reality of Riak in production was miles away from what Basho's sales claimed was possible. And yes, we even spent about 4 months working with their tech support. It almost seems that "It's based on Erlang thus it scales" was the entirety of their design work. I've also worked with Cassandra and have nothing but good to say about it, did what we asked it right out-of-the-box. Datastax was really helpful as well. -- And I have no affiliation with either Basho nor Datastax, just really happy with one product and completely blown away with the poor performance of the other.