4 ms·
[Timescale DevRel here] @zX41ZdbW@ - Thanks for pointing out the various benchmarks that have been run by other companies between Clickhouse and TimescaleDB us
by ryanbooz 5y ago
[Timescale DevRel here]
@zX41ZdbW@ - Thanks for pointing out the various benchmarks that have been run by other companies between Clickhouse and TimescaleDB using TSBS[1]. As we mentioned, we'll dig deeper into a similar benchmark with much more detail than any of those examples in an upcoming blog post.
One notable omission on all of the benchmarks that we've seen is that none of them enable TimescaleDB compression (which also transforms row-oriented data into a columnar-type format). In our detailed benchmarking, queries on compressed columnar data in Timescale outperformed Clickhouse in most queries, particularly as cardinality increases, often by 5x or more. And with compression of 90% or more, storage is often comparable. (Again, blog post coming soon - we are just making sure our results are accurate before rushing to publish.)
The beauty of TimescaleDB columnar compression model is that it allows the user to decide when their workload can benefit from deep/narrow queries of data that doesn't change often (although it can still be modified just like regular row data), verses shallow/wide queries for things like inserting data and near-time queries.
It's a hybrid model that provides a lot of flexibility for users AND significantly improves the performance of historical queries. So yes, we do agree that columnar storage is a huge performance win for many types of queries.
And of course, with TimescaleDB, one also gets all of the benefits of PostgreSQL and its vibrant ecosystem.
Can't wait to share the details in the coming weeks!
[1]: https://github.com/timescale/tsbs https://github.com/timescale/tsbs
- zX41ZdbW 5y agoThank you! Looking forward for a blog post. We need more references for comparison to optimize ClickHouse performance.
- xdanger 5y ago(although it can still be modified just like regular row data) But it can't be updated or deleted, so what do you mean by this?
- ryanbooz 5y agoThat's a great catch @xdanger and you're right, my comment wasn't accurate. Honestly I rewrote the response a few times and this part wasn't cleaned up which is totally on me. The overall concept that I was intending to highlight is that you can benefit from both row & columnar store in TimescaleDB. Chunks that are not yet compressed (row store data) can be modified (INSERT/UPDATE/DELETE) as usual and it's transactional - so you're assured it's been completed. As of TimescaleDB 2.3, compressed chunks (columnar) do allow INSERTS but UPDATES/DELETES on compressed chunks are not yet supported natively. You _can_ decompress any chunk and modify the data (again, transactionally) as needed and recompress.
- stavros 5y agoI have a related question, in case anyone knows: We want to store typical analytics data somewhere (currently in BigQuery) to analyze with Looker. Things like "CI run started", "CI run finished" and then calculate analytics over average CI runtimes. Which database would be a good fit for this? There isn't too much data, maybe tens of thousands of rows eventually. Would Timescale be a good fit? I'd prefer that, due to existing familiarity with Postgres, but if ClickHouse is better, that's good too.
- ants_a 5y agoPostgres has much more featureful query language, and at tens of thousands of rows the performance difference is irrelevant. The story becomes different when answering a query has to touch millions of records and the answer is needed in milliseconds.
- stavros 5y agoThanks!
- saadatq 5y agoWhy move off BigQuery?
- stavros 5y agoIf BigQuery is good, there's no reason to. I just assumed a time-series database would be better suited to the workload.