4 ms·
Yes I've seen that also work, it's a lot of stitching together things yourself and we had to put a lot of caching in front of the inverted index we were using,
by roskilli 7y ago
Yes I've seen that also work, it's a lot of stitching together things yourself and we had to put a lot of caching in front of the inverted index we were using, however definitely plausible. ClickHouse doesn't do any streaming of data between nodes as you scale up and down which was a big thing for us since we had large datasets and needed to rebalance when cluster expanded/shrunk.
With regards to trie vs inverted index for Graphite data, I'd actually still be inclined to say inverted index is better based on the amount of queries I saw at Uber with Graphite where people did `servers.*.disk.bytes-used` type queries which is way faster to do using an inverted index since you have a postings list for each part of the dot-separated metric name, rather than traversing a trie with thousands to tens of thousands of entries in index 1 host part of the Graphite name. This is what M3DB does[0].
[0]: https://github.com/m3db/m3/blob/b2f5b55e8313eb48f023e08f6d53350fabf09338/src/cmd/services/m3coordinator/ingest/carbon/ingest.go#L318-L330 https://github.com/m3db/m3/blob/b2f5b55e8313eb48f023e08f6d53...
- idjango 7y agoJust to point out that there is inverted index implementation of graphite data working on clickhouse. Regarding the auto-rebalance feature, I cannot much more agree with you. It's something that clickhouse definitely need to handle internally.
- roskilli 7y agoThat's interesting, I had not heard of ClickHouse as a backend for Graphite with an inverted index. Let me know if you have any links to that. I'm assuming this is an out of process inverted index used alongside ClickHouse? Or is it more of a secondary table contained by ClickHouse which can be searched to find the metrics, then the data is looked up? The latter scales not as well with billions of unique metrics since it's always a scan across the unique metrics stored in the time window your query searches for (since any arbitrary dimensions can be specified, all must be evaluated). This is the drawback of PromHouse which is an implementation of Prometheus remote storage on top of ClickHouse - and the major reason why PromHouse was only ever a proof of concept rather than a production offering.