3 ms·
StarRocks handles latency and concurrency as well as Clickhouse but also does joins. Less denormalization, and you can use the same platform for traditional BI/
by biggestdummy 3y ago
StarRocks handles latency and concurrency as well as Clickhouse but also does joins. Less denormalization, and you can use the same platform for traditional BI/ad-hoc queries.
- riku_iki 3y agoClickhouse also does joins. Somehow StarRocks dudes appear in every relevant post with this false claim.
- biggestdummy 3y agoThere's a difference between "supports the syntax for joins" and "does joins efficiently enough that they are useful." My experience with Clickhouse is that its joins are not performant enough to be useful. So the best practice in most cases is to denormalize. I should have been more specific in my earlier comment.
- riku_iki 3y agoack that anonymous user in internet said he couldn't make CLickhouse joins perform well in his case which he didn't describe
- biggestdummy 3y agoNot "that anonymous user." In my experience, avoiding Join statements is a common best practice for Clickhouse users seeking performant queries on large datasets. A couple examples... https://medium.com/datadenys/optimizing-star-schema-queries-with-in-queries-and-denormalization-cc281bbe19a5 https://medium.com/datadenys/optimizing-star-schema-queries-... https://posthog.com/blog/secrets-of-posthog-query-performance#whats-next https://posthog.com/blog/secrets-of-posthog-query-performanc...
- riku_iki 3y ago> avoiding Join statements is a common best practice for Clickhouse users seeking performant queries on large datasets its common best practice on any database, because if both joined tables don't fit memory, then merge join is O(nlogn) operation which indeed many times slower than querying denormalized schema, which will have linear execution time.
- slotrans 3y agoI wasn't familiar with StarRocks so thanks for calling attention to it. It appears to make very different tradeoffs in a number of areas so that makes it a potential useful alternative. In particular transactional DML will make it much more convenient for workloads involving mutation. Plus as you suggested, having a proper Cost-Based Optimizer should make joins more efficient (I find ClickHouse joins to be fine for typical OLAP patterns but they do break down in creative queries...) It's a bummer though that the deployment model is so complicated. One thing I truly like about ClickHouse is its ability to wring every drop of performance out of a single machine, with a super simple operational model. Being able to scale to a cluster is great but having to start there is Not Great.