4 ms·
Show HN: Easily Convert WARC (Web Archive) into Parquet, Then Query with DuckDB
- wahnfrieden 4y agoHow does this compare with SQLite approaches shared recently?
- llambda 4y agoIt's a great question: fundamentally the Parquet format offers columnar orientation. With datasets like these, there's some research[0] indicating this is a preferable way of storing and querying WARC. DuckDB, like SQLite, is serverless. Duck has a leg up on SQLite though when it comes to Parquet: Parquet is supported directly in Duck and this makes dealing with these datasets a breeze. [0] https://www.researchgate.net/figure/Comparing-WARC-CDX-Parquet-and-Avro-formats_fig1_340332233 https://www.researchgate.net/figure/Comparing-WARC-CDX-Parqu...
- infogulch 4y agoWell there's a virtual table extension to read parquet files in SQLite. I've not tried it myself. https://github.com/cldellow/sqlite-parquet-vtable https://github.com/cldellow/sqlite-parquet-vtable
- 1egg0myegg0 4y agoGood question! As a disclaimer, I work for DuckDB Labs. There are 2 big benefits to working with Parquet files in DuckDB, and both relate to speed! DuckDB can query parquet right where it sits, so there is no need to insert it into the db first. This is typically much faster. Also, DuckDB's engine is columnar (SQLite is row based), so it can do faster analytical queries using that format. I have seen 20-100x speed improvements over SQLite in analytical workloads. Happy to answer any questions!
- arpinum 4y agoDo you see DuckDB as a possible replacement for AWS Athena? Where would Athena still be better than DuckDB + Parquet + Lambda?
- wenc 4y agoDuckDB user here. As far as I can tell, DuckDB doesn’t support distributed computation so you have to set that up yourself, whereas Athena is essentially Presto — it handles that detail for you. It also doesn’t support Avro or Orc yet. DuckDB excels at single machine compute where everything fits in memory or is streamable (data can be local or on S3) — it’s lightweight and vectorized. I use it in Jupyter notebooks and in Python code. But it may not be the right tool if you need distributed compute over a very large dataset.
- youngtaff 4y ago> But it may not be the right tool if you need distributed compute over a very large dataset I’m really interested in what the limits of DuckDB and Parquet are. Can you give me an idea of what size you mean by “a very large dataset”
- wenc 4y agoIn the distributed computing world, the rule is you start to scale horizontally when your compute workload is too large to fit in the memory of a single machine. So it depends on your compute workload and your hardware. (There’s no fixed number for what a large dataset is) DuckDB itself doesn’t have any baked in limits. If it fits in memory, single machine compute is usually faster than distributed compute — and DuckDB is faster than Pandas, and definitely faster than local Spark.
- wenc 4y agoDuckDB has SQLite semantics but is natively built around columnar formats (parquet, in-memory Arrow) and strong types (including dates). It also supports very complex SQL. SQLite is a row store built around row based transactional workloads. DuckDB is built around analytics workloads (lots of filtering, aggregations and transformations) and for these workloads DuckDB is just way way faster. Source: personal experience.
- 1vuio0pswjnm7 4y agoSQLite3 is 1.6MB duckdb is 41MB (q/k, another columnar SQL database, is less than a MB)
- mritchie712 4y agoNice! I've been considering using DuckDB for our product (to speed up join's and aggregates of in-memory data), it's an incredible technology.