4 ms·
Disclaimer: I'm the product manager for ClickHouse core database. Which version are you working with? I recommend trying the latest version. For the last eight
by SnooBananas6657 1y ago
Disclaimer: I'm the product manager for ClickHouse core database.
Which version are you working with? I recommend trying the latest version. For the last eight months, we have spent considerable time improving our JOINs.
For example, we just merged https://github.com/ClickHouse/ClickHouse/pull/80848 https://github.com/ClickHouse/ClickHouse/pull/80848, which will help a lot with performance in the near future.
- saisrirampur 1y agoSai from ClickHouse here. Adding to above, we just released a blog that presents JOIN benchmarks of ClickHouse against Snowflake and Databricks. This is after the recent enhancements made to the ClickHouse core. https://clickhouse.com/blog/join-me-if-you-can-clickhouse-vs-databricks-snowflake-join-performance https://clickhouse.com/blog/join-me-if-you-can-clickhouse-vs.... The benchmarks is around 2 dimensions of both speed and cost.
- nrjames 1y agoWill Clickhouse spill to disk yet when joins are too large for memory?
- twotwotwo 1y agoThis is really encouraging! Commented elsewhere in the thread but this was one of the main odd points I ran into when experimenting with ClickHouse, and the changes in the PR and mentioned in the recent video about join improvements (https://www.youtube.com/watch?v=gd3OyQzB_Fc&t=137s https://www.youtube.com/watch?v=gd3OyQzB_Fc&t=137s) seem to hit some of the problems. I'm curious whether "condition pushdown" mentioned in the video will make it so "a.foo_id=3 and b.foo_id=a.foo_id" doesn't need "b.foo_id=3" added for optimal speed. I also share nrjames's curiosity about whether the spill-to-disk situation has improved. Not having to even think about whether a join fits in memory would be a game changer.