3 ms·
I submitted this because I thought it was a good, high effort post, but I must admit I was surprised by the conclusion. In my experience, admittedly on differ
by RobinL 2y ago
I submitted this because I thought it was a good, high effort post, but I must admit I was surprised by the conclusion. In my experience, admittedly on different workloads, duckdb is both faster and easier to use than spark, and requires significantly less tuning and less complex infrastructure. I've been trying to transition as much as possible over to duckdb.
There are also some interesting points in the following podcast about ease of use and transactional capabilities of duckdb which are easy to overlook (you can skip the first 10 mins): https://open.spotify.com/episode/7zBdJurLfWBilCi6DQ2eYb https://open.spotify.com/episode/7zBdJurLfWBilCi6DQ2eYb
Of course, if you have truly massive data, you probably still need spark
- IshKebab 2y agoThis guy says "I live and breathe Spark"... I would take the conclusions with a grain of salt.
- mwc360 2y agoAuthor of the blog here: fair point. Pretty much every published benchmark has an agenda that ultimately skews the conclusion. I did my best here to be impartial, I.e I fully designed the benchmark and each test prior to running code on any engine to mimic typical ELT demands w/o having the opportunity to optimize Spark since I know it well.
- titanomachy 2y agoI think you did a good job for these workloads. I did some informal experimenting last year when I had to implement an ELT-type system and I ended up doing it in Spark as well. It was my last choice, because I find operating and debugging Spark to be a huge pain. But everything else I tried was way slower. I didn't think that people used polars a lot for ELT. I've usually seen it used for aggregations with small outputs (which, as you called out, it does a great job at).
- diroussel 2y agoThanks, I'll give that a listen. Here is the Apple Podcast link to the same episode: https://podcasts.apple.com/gb/podcast/the-joe-reis-show/id1676305617?i=1000680142303 https://podcasts.apple.com/gb/podcast/the-joe-reis-show/id16... I've also experimented with duckdb whilst on a databricks project, and did also think "we could do this whole thing with duckdb and a large EC2 instance spun up for an few hours a week". But of course duckdb was new then, and you can't re-architect on a hunch. Thanks for the aricle.
- thenaturalist 2y agoDo you still use file formats at all in your work? I'm currently thinking of ditching Parquet all together, and going all in DuckDB files. I don't need concurrent writes, my data would rarely exceed 1TB and if it were, I could still offload to Parquet. Conceptually I can't see a reason for this not working, but given the novelty of the tech I'm wondering if it'll hold up.
- RobinL 2y agoWe're still using parquet. So we use the native duckdb format for intermediate processing (during pipeline execution) but the end results are saved out as parquet. This is partly because customers often read the data from other tools (e.g. AWS athena) I'd be interested in hearing about experiences of using duckdb files though, i can see instances where it could be useful to us