4 ms·
Columnar storage stores data very efficiently, too - because it compresses data of a similar nature (columns). Check e.g. ClickHouse on this matter: https://cl
by isburmistrov 3y ago
Columnar storage stores data very efficiently, too - because it compresses data of a similar nature (columns).
Check e.g. ClickHouse on this matter: https://clickhouse.com/docs/en/about-us/distinctive-features https://clickhouse.com/docs/en/about-us/distinctive-features, https://clickhouse.com/blog/working-with-time-series-data-and-functions-ClickHouse https://clickhouse.com/blog/working-with-time-series-data-an...
So I wouldn't say that events are "expensive" while metrics are "cheap" - both depend on the actual implementation, and events can be cheap too.
And so of course if you have to optimise things, you would need to drop some information you pass to the events, but you would need to do the same for metrics (reduce the number of metrics emitted, reduce the prometheus labels,...).
- flaminHotSpeedo 3y agoThe whole point of wide events is recording an arbitrary set of key value pairs. How do you propose storing that in a columnar datastore?
- phillipcarter 3y agoI can't speak for others, but at Honeycomb that's what we do. There's some details in this blog post that might be interesting: https://www.honeycomb.io/blog/why-observability-requires-distributed-column-store https://www.honeycomb.io/blog/why-observability-requires-dis...
- hagen1778 3y agoStoring telemetry efficiently is only part of what Monitoring is supposed to do. The other part is querying: ad-hoc queries, dashboards, alerting queries executed each 15s or so. For querying to work fast, there has to be an efficient index or multiple indexes depending on the query. Since you referred ClickHouse as efficient columnar storage, please see what makes it different from a time series database - https://altinity.com/wp-content/uploads/2021/11/How-ClickHouse-Inspired-Us-to-Build-a-High-Performance-Time-Series-Database.pdf https://altinity.com/wp-content/uploads/2021/11/How-ClickHou...
- isburmistrov 3y agoAnd yet people use ClickHouse quite effectively for this very problem, see the comment here: https://news.ycombinator.com/item?id=39549218 https://news.ycombinator.com/item?id=39549218 There are also time-series databases out there that are OK with high cardinality: https://questdb.io/blog/2021/06/16/high-cardinality-time-series-data-performance/ https://questdb.io/blog/2021/06/16/high-cardinality-time-ser...
- hagen1778 3y ago> And yet people use ClickHouse quite effectively for this very problem There is no doubt that ClickHouse is a super-fast database. No one stops you from using it for this very problem. My point is that specialized time series databases will outperform ClickHouse. > There are also time-series databases out there that are OK with high cardinality So does this blog say that tolerance to cardinality means that QuestDB indexes only one of the columns in the data generated by this benchmark? TSDBs like Prometheus, VictoriaMetrics or InfluxDB will perform filtering by any of the labels with equal speed, because this is how their index works. Their users don't need to think about the schema or about which column should be present in the filter. But in ClickHouse and, apparently, in QuestDB, you need to specify a column or list of columns for indexing (the fewer columns, the better). If the user's query doesn't contain the indexed column in the filter - the query performance will be poor (full scan). See like this happened in another benchmarketing blogpost from QuestDB - https://telegra.ph/No-QuestDB-is-not-Faster-than-ClickHouse-06-15 https://telegra.ph/No-QuestDB-is-not-Faster-than-ClickHouse-...
- nhourcard 3y agoIn QuestDB, only SYMBOL columns can be indexed. However, sometimes, queries can run faster without indexes. This is because, under the hood, QuestDB runs very close to the hardware and only lifts relevant time partitions and columns for a given query. Therefore table scans between given timestamps are then very efficient. This can be faster than using indexes when the scan is performed with SIMD and other hardware-friendly optimizations. When cardinality is very high, indexes make more sense.
- adql 3y agoIf you have small pre-defined sets of events in data structures that compress well. That is not the case for any real system. > And so of course if you have to optimise things, you would need to drop some information you pass to the events, but you would need to do the same for metrics (reduce the number of metrics emitted, reduce the prometheus labels,...). Those are entirely different orders of magnitude both when it comes to size and how much usefulness you lose. In modern storage backends like Victoriametrics a counter gonna cost you around byte per metric per probe. And as you emit them periodically, that is essentially independent of incoming traffic Capturing the requests into event/trace/whatever other name they gave to logs this month is many times that and is multiplied by traffic.
- isburmistrov 3y ago> Those are entirely different orders of magnitude both when it comes to size and how much usefulness you lose. In modern storage backends like Victoriametrics a counter gonna cost you around byte per metric per probe. And as you emit them periodically, that is essentially independent of incoming traffic I thought this argument was about whether wide events can be used for metrics or metrics is a completely different concept. If we want to emulate metrics in events, we would also make them periodically independently of the traffic. Like emit them once in a while. Pretty much like Prometheus scraping works