7 ms·
Would be curious how the performance compares to DataFusion[0] as one of the top contenders to DuckDB on this area (albeit they being different in a lot of part
by treesciencebot 3y ago
Would be curious how the performance compares to DataFusion[0] as one of the top contenders to DuckDB on this area (albeit they being different in a lot of parts, I find it one of the closest compared to all others).
ClickBench (from ClickHouse) has some benchmarks[1] where it can be compared, but am not super sure how up to date it is. At least a while back, they were majorly out of date and haven't looked too closely on whether they are keeping it fair for everyone else :)
[0]: https://github.com/apache/arrow-datafusion https://github.com/apache/arrow-datafusion
[1]: https://benchmark.clickhouse.com https://benchmark.clickhouse.com
- jabart 3y agoLooks like a recent PR bumped benchmark.clickhouse.com to DuckDB v0.9 on the 3rd. https://github.com/ClickHouse/ClickBench/pull/141 https://github.com/ClickHouse/ClickBench/pull/141
- leicmi 3y agoA paper on DataFusion is in progress[0]. The draft[1] includes a comparison to DuckDB and preliminary benchmark results. [0]: https://github.com/apache/arrow-datafusion/issues/6782 https://github.com/apache/arrow-datafusion/issues/6782 [1]: https://www.overleaf.com/read/qjhrxqhgksvr https://www.overleaf.com/read/qjhrxqhgksvr
- riku_iki 3y agoWhy do you run benchmarks on such small datasets? It is very hard to judge performance..
- sanderjd 3y agoLooks like DataFusion is included in most of the results in the article?
- treesciencebot 3y agoYou are right! Seems like it is not text-addressable which is why my ctrl+f searches failed.
- slt2021 3y agoquestion about Arrow: the format seems to be not very space efficient. I tried converting one of my parquet files from datalake from parquet to arrow and size difference is staggering. 20mb parquet -> 700mb arrow. doesnt seem fit for datalake at all
- iwd 3y agoDo you have compression enabled? At least from Pandas, Parquet defaults to compressed and Arrow/Feather default to uncompressed. When I enable zstd compression, I get similar file sizes, and sometimes Arrow is smaller.
- slt2021 3y agoI was just trying pandas native .to_parquet and .to_arrow() without any extra config knobs
- alamb 3y agoThe following paper describes some of the tradeoffs between different formats Deep Dive into Common Open Formats for Analytical DBMSs https://www.vldb.org/pvldb/vol16/p3044-liu.pdf https://www.vldb.org/pvldb/vol16/p3044-liu.pdf
- lomereiter 3y agoArrow format is not intended for storage, it's for in-memory data exchange between different libraries and languages.
- ayhanfuat 3y agoArrow is not really designed for storage though. See the "Parquet vs Arrow" section of this post (https://arrow.apache.org/blog/2022/10/05/arrow-parquet-encoding-part-1/ https://arrow.apache.org/blog/2022/10/05/arrow-parquet-encod...): > Parquet and Arrow are complementary technologies, and they make some different design tradeoffs. In particular, Parquet is a storage format designed for maximum space efficiency, whereas Arrow is an in-memory format intended for operation by vectorized computational kernels. > The major distinction is that Arrow provides O(1) random access lookups to any array index, whilst Parquet does not. In particular, Parquet uses dremel record shredding, variable length encoding schemes, and block compression to drastically reduce the data size, but these techniques come at the loss of performant random access lookups.