5 ms·
DuckDB is one of the things I've been most excited about in a long time. Introduced it to projects at 3 companies since 2023, greatly lowering resource requirem
by jtbaker 1mo ago
DuckDB is one of the things I've been most excited about in a long time. Introduced it to projects at 3 companies since 2023, greatly lowering resource requirements and running it in a variety of environments. Just having the ability to do out of core bigger than memory data processing on lower end consumer grade hardware is remarkable.
Thanks to the team for everything!
- throwaw12 1mo agoCurious to learn more about how people are using it? Are they downloading parquet files and running analyses locally, or are they connecting to Iceberg-like data lake and leveraging DuckDBs query engine capabilities or have you exposed an interface (REST, UI) to query your data?
- arealaccount 1mo agoWe use DuckDB WASM with parquet to build dashboards in-browser. It's cool to be able to write SQL directly in a browser and not have to rely on REST/Graphql/etc to access the data layer.
- jayct 1mo agocurious if you're using something mostly-out-of-the-box to layer on visualizations for your dashboards? relatively new to duckdb, love it so far, looking at alternatives for downstream visualization. so far just exporting datasets and piping into python scripts.
- mediaman 1mo agoI do something similar and just use echarts. Very happy with it.
- jtbaker 1mo agoFor a schema-first (vs. code first) approach (which I think would be a sweet spot for agent driven dashboarding), I'd suggest looking at https://vega.github.io/vega-lite/ https://vega.github.io/vega-lite/ or https://vega.github.io/vega/ https://vega.github.io/vega/. A little higher level than full D3 but gives you a little higher level approach.
- drums8787 1mo agoWe use WASM DuckDB as the target for an in-browser agentic feature. Generated SQL runs against the user's individual tables that then feed in-browser dashboards. Excellent performance.
- throw1234567891 1mo ago[flagged]
- fg137 1mo agoRunning duckdb as wasm in browser for dashboards is a very common use case. Does that make this account an alias as well?
- jagenabler2 1mo agoWhere are you storing and fetching the users data to be able to query it in-browser? Are you pointing it at S3? Investigating a similar use case and trying to figure out what our infra could look like!
- jtbaker 1mo agoI've got a couple of different use cases: - ETL pipelines running on K8s nodes. Using their streaming processing engine means I can run smaller pods/nodes if needed, for datasets that may have required large dataframe-like transformations that may have buffered a big dataset into memory previously. - A CLI distributed to an internal team to do a postprocessing step on a large modeling dataset - to get it into a consumable format and upload it to a bucket as a .db file. - A SvelteKit app that used the node duckdb bindings to attach to the .db on the bucket and explore the results through a suite of BI tools. These tables have millions of rows, and would be pretty heavy to store in PG. The DuckDB version works really, really well.
- tccole 1mo agoHell yeah; a fellow sveltekit fan.
- staticautomatic 1mo agoSimilar here. Lots of places where we replaced Pandas with DuckDB for transformations. Also have scriptable custom dashboards running on top of BigQuery data pre-aggregated and extracted to parquet on GCS. It's way faster and the only limiting factor is your viz library. It was pretty easy to build and the only big gotcha I encountered was finding, somewhat counter-intuitively, that it's often best minimize partitioning.
- jtbaker 1mo ago> the only big gotcha I encountered was finding, somewhat counter-intuitively, that it's often best minimize partitioning. For parquet, I think with partitioning, it's really important to be mindful of the ordering of the data within the parquet file and also the query patterns of the main use cases. A little hard to generalize well to every pattern I guess.
- peesem 1mo agomaybe a niche use case but i've found it's perfect to store & query random trivia/gameshow questions based on filters for my personal clones of things like Family Feud and Jeopardy
- tccole 1mo agoYes. I have used it with WASM for some web applications for web use. I have also used with locally for querying 100 gigs of data. And I have used it in the cloud as the serverless gold layer for Apache superset.
- arpinum 1mo agoETL from DynamoDB into Ducklake
- phqb 1mo agoI'm using duckdb/duckdb-go as query engine for my Go services: moving hot data from Postgres to Parquet files on S3 or to Iceberg; querying cold data on Iceberg, ... instead of using different Go libraries.
- malshe 1mo agoI use it locally with parquet files
- AlfeG 1mo agoRealtime full MSSQL database mirroring into DuckDb to do a complex reporting. Everything is in-process. DuckDb database mapped to temp storage and recreated on app restart. Still order of magnitude faster then doing a direct query over MSSQL Server (2ms vs 40+ seconds on same query). Some devs in team still cannot believe that there is no cheating, that it's possibe, that some 60Mb DB can do queries faster then MSSQL Server with just around 250Mb+ of memory overhead. (.Net 10 + DuckDB.NET package)
- raihansaputra 1mo agohow are you doing the etl from mssql to duckdb? are you doing naive snapshots? or ducklake?
- markus_zhang 1mo agoI actually used it in an interview. I downloaded the csv/parquet file and load it into DuckDB. In the future, I also plan to use it for testing production data pipelines: imagine you have a streaming cdc pipeline running in development env, and at the end you can dump every parquet into DuckDB as a verification — the end result should be the same. I could also use the same Database for testing, but I like DuckDB somehow.
- fifilura 1mo agoMy favourite is AWS Athena (backed by Trino). "If we use this we get indefinite RAM indefinite CPU and do not need to host a server". I had an impression that DuckDB was not great at distributing work to other machines, but good at doing it locally? Am I wrong?
- ericpauley 1mo agoAthena + Clickhouse has been an absolute game changer for us. Perfect combo for OLAP + deeper filtering that we can’t necessarily pre-index for.
- abirch 1mo agoDuckDB out of the box may not be great. But you have DuckLake, Quack, and even DeepSeek made their own distributed DB based on DuckDB: https://github.com/deepseek-ai/smallpond https://github.com/deepseek-ai/smallpond
- jtbaker 1mo agoI don't think DuckDB itself can coordinate work across multiple nodes. But you could put it behind an HTTP layer and scale horizontally based on resource utilization?