3 ms·
Outside of being open source, how does ClickHouse differ from Snowflake/BigQuery? In what scenarios would I choose ClickHouse over those existing solutions?
by ian-whitestone 4y ago
Outside of being open source, how does ClickHouse differ from Snowflake/BigQuery? In what scenarios would I choose ClickHouse over those existing solutions?
- 62951413 4y agoDruid and Pinot are more likely to be the peer group (e.g. see https://leventov.medium.com/comparison-of-the-open-source-olap-systems-for-big-data-clickhouse-druid-and-pinot-8e042a5ed1c7 https://leventov.medium.com/comparison-of-the-open-source-ol...)
- zX41ZdbW 4y agoClickHouse supports ad-hoc analytics, real-time reporting, and time series workloads at the same time. It is perfectly suited for user-facing analytics services. It supports low-latency (<100ms) queries for real-time analytics as well as high query throughput (500 QPS and more) - all of this with real-time data ingestion of logs, events, and time series. Take some notable examples from the list: https://clickhouse.com/docs/en/about-us/adopters/ https://clickhouse.com/docs/en/about-us/adopters/, something around web analytics, APM, ad networks, telecom data... ClickHouse is perfectly suited for these use cases. But if you try to align these scenarios with, say, BigQuery, they will become almost impossible or prohibitively expensive or just slow. There are specialized systems for real-time analytics like Druid and Pinot, but ClickHouse does it better: https://benchmark.clickhouse.com/ https://benchmark.clickhouse.com/ There are specialized systems for time-series workloads like InfluxDB and TimescaleDB, but ClickHouse does it better: https://gitlab.com/gitlab-org/incubation-engineering/apm/apm/-/issues/4 https://gitlab.com/gitlab-org/incubation-engineering/apm/apm... https://arxiv.org/pdf/2204.09795.pdf https://arxiv.org/pdf/2204.09795.pdf http://cds.cern.ch/record/2667383/ http://cds.cern.ch/record/2667383/ There are specialized systems for logs and APM, but ClickHouse does it better: https://blog.cloudflare.com/log-analytics-using-clickhouse/ https://blog.cloudflare.com/log-analytics-using-clickhouse/ There are specialized systems for ad-hoc analytics, but ClickHouse does it better as well: https://github.com/crottyan/mgbench https://github.com/crottyan/mgbench Well, even if you want to process a text file, ClickHouse will do it better than any other tool: https://github.com/dcmoura/spyql/blob/master/notebooks/json_benchmark.ipynb https://github.com/dcmoura/spyql/blob/master/notebooks/json_... And ClickHouse looks like a normal relational database - there is no need for multiple components for different tiers (like in Druid), no need for manual partitioning into "daily", "hourly" tables (like you do in Spark and Bigquery), no need for lambda architecture... It's refreshing how something can be both simple and fast.