4 ms·
Sorry if this is obvious, but... What is ClickHouse?
by andridk 5y ago
Sorry if this is obvious, but... What is ClickHouse?
- dm33tri 5y agoIn a nutshell: like MySQL but for analytic applications
- rapsey 5y agoYou could clink on the link and find out.
- amyjess 5y agoOfficially, it's a column-oriented database. This means that, internally, it stores columns together rather than rows together. In practice, it means that it's optimized for calculating analytics over large datasets. I've found, from personal experience, that it makes a good replacement for time-series databases, even though it's technically not a time-series database. My employer migrated our KPIs and other metrics from InfluxDB to ClickHouse a couple of years ago, and the drastic improvements in performance were well worth the time it took to migrate our data. It also helped that ClickHouse uses a subset of SQL, unlike InfluxDB which uses a superficially SQL-like but practically very different proprietary language.
- andridk 5y agoThanks! Very informative
- ithkuil 5y agoYou may be interested in https://github.com/influxdata/influxdb_iox https://github.com/influxdata/influxdb_iox Column store db, using the DataFusion SQL engine. Persistence using parquet files directly on object stores (e.g. S3) Still under development
- swasheck 5y agoEvery high-volume Influx/Grafana implementation I’ve used has been a disaster. I’m now at a place that uses ClickHouse and I can now see the utility of Grafana
- sofixa 5y agoHow so? I run a moderately sized InfluxDB stack (2x~2TB data) with Grafana, and it works pretty well.
- ithkuil 5y agoThe shape of the data matters. In particular the cardinality of the tags. If clickhouse works well for you, chances are that your use case will be well served by influxdb_iox too
- fnord77 5y agofrom what I've read, it is not a column-oriented db. It's more like druid or pinot.
- manigandham 5y agoClickhouse/Druid/Pinot are all columnstores/column-oriented databases. Clickhouse is a relational engine while Druid/Pinot are a different (and older) design using heavy indexing and pre-aggregation. All of them store table data as per-column segments though which is a defining feature leading to high compression and I/O performance. There's also the badly named wide-column database type like Cassandra, but this is really just advanced or nested key/value rather than what people would consider "columns".