8 ms·
Smallpond – A lightweight data processing framework built on DuckDB and 3FS
- fastasucan 2y agoWhat does this do - what is the benefit over DuckDB, Polers etc?
- mritchie712 2y agoI don't think you get any really benefits over duckdb unless your data is 10tb+ or you spin up 3FS (which seem challenging).
- articsputnik 2y agoMehdi just wrote about this. Mainly starting DAGs parallelism using Ray (core) and their filesystem 3FS. See https://mehdio.substack.com/p/duckdb-goes-distributed-deepseeks https://mehdio.substack.com/p/duckdb-goes-distributed-deepse....
- ilove196884 2y agoAny benchmark and comparisons?
- HackerThemAll 2y agoDuckDB itself is cool enough, especially when combined with SQLite and/or PostgreSQL, and now this. Thanks DeepSeek!
- dcreater 2y agoHow is duckdb combined with SQLite? Aren't they alternatives to each other?
- jitl 2y agoNot sure what the poster meant but DuckDB is an analytics DB, it doesn’t have a btree index - at least not last time I looked. You could consider it the OLAP embedded DB to SQLite’s OLTP embedded db. DuckDB can read SQLite so you can even imagine using them side by side in the same system, serving point reads and writes from SQLite and using DuckDB for stuff like aggregates and searches that SQLite is slower at.
- HackerThemAll 2y agoThey are complementary to each other. There's an SQLite extension for use within DuckDB [1], which gives you a power of great transactional capabilities of SQLite and speed of analytical queries within DuckDB's columnar storage engine, all within a single database. [1] https://duckdb.org/docs/stable/extensions/sqlite.html https://duckdb.org/docs/stable/extensions/sqlite.html
- orlp 2y agoOne thing I found peculiar is that for the GraySort benchmark it dispatches to Polars by default to do the actual sorting, not DuckDB: https://github.com/deepseek-ai/smallpond/blob/ed112db42af4d006a80861d1305a1c22cabdd359/benchmarks/gray_sort_benchmark.py#L274 https://github.com/deepseek-ai/smallpond/blob/ed112db42af4d0....
- tomnipotent 2y agoThe function argument defaults to polars, but the actual benchmark code sets duckdb by default. https://github.com/deepseek-ai/smallpond/blob/ed112db42af4d006a80861d1305a1c22cabdd359/benchmarks/gray_sort_benchmark.py#L366 https://github.com/deepseek-ai/smallpond/blob/ed112db42af4d0...
- orlp 2y agoI see, confusing multiple layers of defaults :)
- shipp02 2y agoIs the code written by the deepseek model? I should probably give up on being a software engineer if it is.
- breadwinner 2y agoGive up and become what? Most white collar jobs will be automated in the coming years. You think doctors' jobs are safe?
- ezst 2y agoNot OP, but, anything that actually physically affects the real world for the better? For instance, large infrastructure engineering and construction projects are not going to run themselves any time soon. The world doesn't revolve around ad and fin tech.
- rscho 2y agoYes, doctors are safe. Because they do things. With their hands. That no one else does.
- delfinom 2y agoNope. Healthcare megacorps are buying up independent practices like crazy. All because doctors can't keep up with the bullshit IT required for insurance, state mandates, etc and that's in addition to the insanity of even renting commercial real estate for an office these days. These megacorps set quotas and push doctors to nickel and dime like crazy. They sure as shit will spend the money to find robots that can give you a prostate exam with a robot dildo.
- mdaniel 2y agoSounds good; if all these pro-AI folks could get it to complete the insurance paperwork that'd be swell. Actually, come to think of it, do that for the paperwork from both sides, doctor and patient, and eliminate and entire class of leaches upon humanity I'm going to laugh if DOGE eliminates the IRS, but also might be thankful
- rubenvanwyk 2y agoMay Data Engineering content keep on hitting front page HN!
- RyanHamilton 2y agoIf you want to checkout duckdb try QStudio. It's a free sql client with duckdb integrated: https://www.timestored.com/qstudio/help/duckdb-sql-editor https://www.timestored.com/qstudio/help/duckdb-sql-editor. Disclaimer: I'm the main author.
- maximilianroos 2y agoBig fan of QStudio! Thanks for building it!
- dcreater 2y agoWhat's with the win95 ui?
- RyanHamilton 2y agoThere are many themes to choose from. I recorded the demo on that page and I like windows 95. I concede it may not be pretty but I've always found it functional. The default is darcula theme like shown on the main page: https://www.timestored.com/qstudio/ https://www.timestored.com/qstudio/
- lvl155 2y agoLooking forward to next few years when we can finally abstract away all the back-end techs.
- threeseed 2y agoWe've had this for at least a decade now. If you use a cloud provider there are managed solutions for data engineering pipelines.
- BobbyJo 2y agoWe ain't even solved garbage collection yet, and you think "back end systems" are going to abstracted away in the next few years?
- tarruda 2y ago> We ain't even solved garbage collection yet Can you elaborate on that?
- BobbyJo 2y agoPeople still write in languages that force you to manage your own memory. Once performance starts to matter (either due to scale or time requirements) abstractions always have tradeoffs you can't accept.
- dang 2y agoRelated ongoing thread: Understanding Smallpond and 3FS - https://news.ycombinator.com/item?id=43232410 https://news.ycombinator.com/item?id=43232410 also: DuckDB goes distributed? DeepSeek's smallpond takes on Big Data - https://news.ycombinator.com/item?id=43206964 https://news.ycombinator.com/item?id=43206964 (no comments there, but some people have been recommending that article)
- jamesblonde 2y agoWe are seeing more and more specialized query engines. This is a query engine specialized for training pipelines. It is not general purpose - it is for providing batches of training data at workers. It uses Ray for parallelization. The kind of queries you need are random reads (to implement shuffling across epochs), arrow support (zero copy to Pandas DataFrames), and efficient checkpointing.
- auxten 2y agoData operations are increasingly happening near the GPU side to boost efficiency—especially for compute-heavy workflows. Talking about Arrow file processing and zero-copy queries on DataFrames, which are becoming crucial for modern data pipelines. I think another option worth considering is chdb, which supports these features and fits well with this shift. (author of chdb here)
- agilob 2y agoI'm super impressed how much effort DeepSeek did and how much of it they opensourced.
- nyrikki 2y agoSome of what they are doing is simply what was lost due to the ubiquitous nature of the relational model. The hierarchical model is applicable to many problems and actually in part why moving off mainframes is challenging because IMS is so much more efficient than the relational model for applications like airline tickets. There have been several efforts to leverage object stores in the way they did that I am aware of but it was a hard sell. The hierarchical model really only works for many to one relationships, and it's integrity model differs and is not as DRY. There are lessons to learn here but it requires some relearning. When you have a shopping cart, having data local to the server handling the transaction is also a benefit. Codd's relational model has advantages, but has held back some efforts because we are just use to dealing with the painful parts that we often don't consider other options.
- HackerThemAll 2y agoDuckDB is specialized in efficient storage and fast query for analytics (OLAP), using a columnar storage (in contrast to row storage, used by usual RDBMSes doing OLTP processing). It's nothing new, it's been there for couple decades already. But this "distributed" DuckDB can indeed be beneficial for training.
- dcreater 2y agoConfused by the example in the repo? What is the use case for this? Is it a replacement for dask, ray etc? (Not a professional swe)