3 ms·
I know all that, I have extensive experience with NoSQL databases like HBase. But the thing I'm asking about how does Cassandra handle CPU hotspotting on a sing
by bluecmd 11y ago
I know all that, I have extensive experience with NoSQL databases like HBase. But the thing I'm asking about how does Cassandra handle CPU hotspotting on a single row.
If you have timestamp as a column you will never be able to shard a metric across servers, so accessing that metric will be I/O bound to whatever underlying storage.
My hypothesis is that if you save data as "metric.timestamp" as row instead you will get much better scalability.
- PretzelPirate 11y agoThe short answer is that Cassandra doesn't handle hotspotting, its up to you to data model to avoid that. Timestamp is your clustering key, its only the partition key that determines how the data is sharded. If you have a table with firstName, lastName, loginTime and partition key = firstName, clustering key = loginTime, all rows with firstName ="Tim" will live on the same server regardless of loginTime. If you are often querying for logins, you may think that you should instead make the partition key = (firstName, loginTime), but that would leave your data in an almost irretrievable state since you would need to know the exact login times. It would also leave you looking at multiple partitions in every query, which isn't as performance as a single partition. A better partition key would be (firstName, dayAndHourOfLoginTime) so your partitions would be for each first name and each hour. Then your query would still have to know what hours to look for, but its a simple query pattern. If you are often looking at data across more than one hour, you probably want to rethink the key even further because its more efficient to look at a single partition.
- bluecmd 11y agoAha! Re-reading your previous comment I understood what you meant. So Cassandra has a "logical" key that is the primary key, and the actual key from HBase-like DBs is derrived from cluster key and the primary key. So in essence it's "metric.timestamp" as a row key but only schema defined. Nice.