4 ms·
While this is true, for a metrics workload it does not work great I have both seen and heard from others, mainly due to the fact it does not have an inverted in
by roskilli 7y ago
While this is true, for a metrics workload it does not work great I have both seen and heard from others, mainly due to the fact it does not have an inverted index - so finding a small subset of metrics in a dataset of billions of metrics ends up taking significant time due to the scan required to find the timeseries matching the arbitrary number of dimensions specified to find the timeseries you're looking for.
If you're building it with a specific application and a concrete schema you can create which will result in fast queries and don't have requirements for arbitrary dimensions being specified for lookup, then yes it's great as a TSDB.
Prometheus, M3DB, etc all use an inverted index alongside the column store TSDB to help with metrics workloads.
- mbell 7y agoMost practical applications using Clickhouse for metrics data store the metric index separately. What index you want really depends on the metric system, e.g. with graphite data you don't want an inverted index, you want a trie.
- roskilli 7y agoYes I've seen that also work, it's a lot of stitching together things yourself and we had to put a lot of caching in front of the inverted index we were using, however definitely plausible. ClickHouse doesn't do any streaming of data between nodes as you scale up and down which was a big thing for us since we had large datasets and needed to rebalance when cluster expanded/shrunk. With regards to trie vs inverted index for Graphite data, I'd actually still be inclined to say inverted index is better based on the amount of queries I saw at Uber with Graphite where people did `servers.*.disk.bytes-used` type queries which is way faster to do using an inverted index since you have a postings list for each part of the dot-separated metric name, rather than traversing a trie with thousands to tens of thousands of entries in index 1 host part of the Graphite name. This is what M3DB does[0]. [0]: https://github.com/m3db/m3/blob/b2f5b55e8313eb48f023e08f6d53350fabf09338/src/cmd/services/m3coordinator/ingest/carbon/ingest.go#L318-L330 https://github.com/m3db/m3/blob/b2f5b55e8313eb48f023e08f6d53...
- idjango 7y agoJust to point out that there is inverted index implementation of graphite data working on clickhouse. Regarding the auto-rebalance feature, I cannot much more agree with you. It's something that clickhouse definitely need to handle internally.
- roskilli 7y agoThat's interesting, I had not heard of ClickHouse as a backend for Graphite with an inverted index. Let me know if you have any links to that. I'm assuming this is an out of process inverted index used alongside ClickHouse? Or is it more of a secondary table contained by ClickHouse which can be searched to find the metrics, then the data is looked up? The latter scales not as well with billions of unique metrics since it's always a scan across the unique metrics stored in the time window your query searches for (since any arbitrary dimensions can be specified, all must be evaluated). This is the drawback of PromHouse which is an implementation of Prometheus remote storage on top of ClickHouse - and the major reason why PromHouse was only ever a proof of concept rather than a production offering.