3 ms·
Slightly off-topic but CedarDB is extremely exciting. It's the commercialization of the widely cited Umbra research DBMS [0] that has been in the works for seve
by refset 2y ago
Slightly off-topic but CedarDB is extremely exciting. It's the commercialization of the widely cited Umbra research DBMS [0] that has been in the works for several years, which benchmarks faster than DuckDB for OLAP [1] whilst simultaneously being really strong for transactional workloads. Also discussed recently here [2].
[0] https://umbra-db.com/ https://umbra-db.com/
[1] https://cedardb.com/blog/ode_to_postgres/ https://cedardb.com/blog/ode_to_postgres/
[2] https://news.ycombinator.com/item?id=40241150 https://news.ycombinator.com/item?id=40241150
- riku_iki 2y ago> which benchmarks faster than DuckDB for OLAP [1] that link doesn't do any performance comparison. They claim that CedarDB executes less "code branches" than duckdb, which may or may not translate to faster performance.
- ChrisWint 2y agoAuthor of that blogpost here. > less "code branches" than duckdb, which may or may not translate to faster performance. In that case it was about 2.5x faster than DuckDB end to end, so a bit less than the difference in branches. If you want to see some independent benchmarks on Umbra, our underlying technology, its currently first place on Clickbench [1]. You can compare against duckdb there as well. [1] https://benchmark.clickhouse.com/ https://benchmark.clickhouse.com/
- riku_iki 2y agoclickbench is a toy benchmark: small dataset, very specific queries. Benchmarking full tcp-h (not just one query like in your post) on sizable dataset (few TBs) would be very good close to real world scenario, but vendors usually avoid this.
- pfent 2y agoThere are TPC-H numbers in another post: https://cedardb.com/blog/simple_efficient_hash_tables/ https://cedardb.com/blog/simple_efficient_hash_tables/
- riku_iki 2y agoits great starting insight, but again its small dataset (100GB) which almost fits memory, and I think many details are missing (for example clickbench publishes all configs and queries, and more detailed report, so vendors can reproduce/optimize/dispute them).
- refset 2y ago> small dataset (100GB) What counts as large or small definitely varies a lot depending on the context of the conversation/analysis. MotherDuck's "Big Data is Dead" post [0] sticks in mind: > The general feedback we got talking to folks in the industry was that 100 GB was the right order of magnitude for a data warehouse. This is where we focused a lot of our efforts in benchmarking. Another point of reference is [1] > [...] Umbra achieves unprecedentedly low query latencies. On small data sets, it is even faster than interpreter engines like DuckDB > TPC-H Small Dataset = 866k tuples, sf 0.1 [0] https://motherduck.com/blog/big-data-is-dead/ https://motherduck.com/blog/big-data-is-dead/ [1] https://db.in.tum.de/~kersten/Tidy%20Tuples%20and%20Flying%20Start%20Fast%20Compilation%20and%20Fast%20Execution%20of%20Relational%20Queries%20in%20Umbra.pdf?lang=de https://db.in.tum.de/~kersten/Tidy%20Tuples%20and%20Flying%2...
- riku_iki 2y ago> The general feedback we got talking to folks in the industry was that 100 GB and then user discovers that DuckDB is plagued with OOMs and dramatic performance degradations when his data is slightly larger than memory.
- jauntywundrkind 2y agoA little sad finding out it's proprietary but to be expected I suppose. It's neat seeing postgres gearing up for async support. There's also folks like OrioleDB doing massive revamps of postgres & doing disaggregated storage in public. https://github.com/orioledb/orioledb https://github.com/orioledb/orioledb https://hn.algolia.com/?query=orioledb&sort=byDate https://hn.algolia.com/?query=orioledb&sort=byDate Oh they were bought by Suprabase two months ago... Fingers crossed! The Suprabase CEO commented at the time, with a nice basic overview, https://news.ycombinator.com/item?id=40039138 https://news.ycombinator.com/item?id=40039138
- deleted 2y ago[deleted]
- pdimitar 2y ago[flagged]