3 ms·
Yes - agree! I actually wrote a blog about this just two days ago: May be of interest to people who: - What to know what DuckDB is and why it's interesting
by RobinL 2y ago
Yes - agree! I actually wrote a blog about this just two days ago:
May be of interest to people who:
- What to know what DuckDB is and why it's interesting
- What's good about it
- Why for orgs without huge data, we will hopefully see a lot more of 's3 + duckdb' rather than more complex architectures and services, and hopefully (IMHO) less Spark!
https://www.robinlinacre.com/recommend_duckdb/ https://www.robinlinacre.com/recommend_duckdb/
I think most people in data science or data engineering should at least try it to get a sense of what it can do
Really for me, the most important thing is it makes it so much easier to design and test complex ETL because you're not constantly having to run queries against Athena/Spark to check they work - you can do it all locally, in CI, set up tests, etc.
- yakshaving_jgt 2y agoFunny, I read TFA and came to the comments to share exactly this recent blog post of yours. Big fan of your work, Robin!
- RobinL 2y agoAh nice - reading that made me feel good! Appreciate the feedback!
- hn1986 2y agofrom the blog: "This is a very interesting new development, making DuckDB potentially a suitable replacement for lakehouse formats such as Iceberg or Delta lake for medium scale data." I don't think we'll ever see this, honestly. excellent podcast episode with Joe Reis - I've also never understood this whole idea of "just use Spark" or you gotta get on Redshift.
- teruakohatu 2y ago> excellent podcast episode with Joe Reis - I've also never understood this whole idea of "just use Spark" or you gotta get on Redshift. Can you link to the podcast episode?
- hn1986 2y agoepisode and transcript Referenced in the blog: https://www.robinlinacre.com/recommend_duckdb/#:~:text=in%20this%20podcast%2C%20transcribed%20here. https://www.robinlinacre.com/recommend_duckdb/#:~:text=in%20...
- pletnes 2y agoI have the same thoughts. However my impression is also that most orgs would choose eg databricks or something for the permission handling, web ui, ++ so what is the equivalent «full rig» with duckdb and S3 / blob storage?
- RobinL 2y agoYeah I think that's fair, especially from the 'end consumer of the data' point of view, and doing things like row-level permissions. For the ETL side, where often whole-table access is good enough, I find Spark in particular very cumbersome - there's more than can go wrong vs. DuckDB and it's harder to troubleshoot.