5 ms·
Building Databases over a Weekend
- 01HNNWZ0MV43FF 2y ago> In this post we take you on a walkthrough on how you can use DataFusion Thought it was gonna be a "build your own SQLite" or something
- dangoodmanUT 2y agothis post feels like it's skipping over a lot of code that could be included
- JoeOfTexas 2y agoStep 1. Choose a color Step 2. Finish the database
- ztratar 2y agoI almost thought the opposite, but im no db guy.
- ambrood 2y agothanks for the feedback! the first version had a lot more detailed code but decided to go with linking to our GitHub than copying all the code. Wanted to illustrate the core touch points involved in extending DF.
- deleted 2y ago[deleted]
- Gepsens 2y agoI remember 2 years ago someone proposed adding stream processing in datafusion and PRs followed. But IMO stream processing is an entirely different beast, some people could use the sql engine of df for it though. There are rust projects like Arroyo
- necubi 2y agoCreator of Arroyo here—we agree that stream processing is a different beast and needs different infrastructure from a batch engine like DataFusion. Our approach has been to take pieces of DF (including the SQL frontend and expression engine) but embedding them in our own dataflow and operators. This allows us to support low latency, distribution, watermark processing, and consistent checkpointing. But the great thing about DF is that it’s designed as a toolkit for SQL-oriented data processing, so it’s relatively easy to pick and use just the pieces you need.
- knuckleheads 2y agoI’ve been messing around with sql and stream processing off and on the last few months via https://github.com/zmaril/bpfquery https://github.com/zmaril/bpfquery and then https://github.com/zmaril/zquery https://github.com/zmaril/zquery, so I very much feel this comment. I didn’t want to build out my own stream processing architecture in bpfquery, it was getting pretty tough pretty fast, so I switched over to a datafusion backend in zquery in the hopes that it could do stream processing well. It can handle static data really well, much better the home grown half engine I made in bpfquery, but streaming sql isn’t easily possible at the moment, everybody is building their own implementations and trying to upstream what they need, no coherent whole from data fusion. I was looking into making an attempt with arroyo sometime, but I think the authors want that code to be used as a standalone binary and not as a library in something else, based on my last impression of it a while back. So, maybe in a few years it’ll be as easy to make a streaming database as it is now to make a normal one, but that’s not the case currently.
- hantusk 2y agoI agree. So many disparate solutions. The streaming sql primitives are by themselves good enough (e.g. `tumble`, `hop` or `session` windows), but the infrastructural components are always rough in real life use cases. crossing fingers for solutions like `https://github.com/feldera/feldera https://github.com/feldera/feldera` to be wrapped in a nice database, `https://materialize.com/ https://materialize.com/` to solve their memory issues, or `https://clickhouse.com/docs/en/materialized-view https://clickhouse.com/docs/en/materialized-view` to solve reliable streaming consumption. Various streaming processing frameworks often have domain specific languages with a lot of limitations of how to express aggregations and transformations.
- maximus93 2y agoGreat discussion here! At AI Squared, we have also been exploring the evolving landscape of stream processing and SQL engines. While batch engines like DataFusion excel at handling static data, we recognize the challenges around integrating streaming capabilities and infrastructure seamlessly. Our focus has been on simplifying data activation pipelines with tools like Multiwoven, which aims to bridge the gap between static and dynamic data needs by supporting connectors for both traditional databases and real-time platforms like Kafka. However, the need for more embedded, developer-friendly streaming solutions is clear, and it’s exciting to see the progress in projects like Arroyo, Materialize, and ClickHouse. For us, the balance lies in usability and flexibility—how can we empower teams to embed robust data capabilities (whether streaming or batch) into their workflows without overloading on infrastructure complexity? As this ecosystem evolves, we’re optimistic about collaborating and contributing to solutions that make streaming SQL as accessible as traditional SQL. Looking forward to seeing how this space develops—and kudos to the teams pushing boundaries! https://github.com/Multiwoven/multiwoven/ https://github.com/Multiwoven/multiwoven/
- deleted 2y ago[deleted]
- alamb 2y agoBTW here is a fun exercise that takes this idea to the extreme. Who can build a custom file format that gets the best ClickHouse performance (on DataFusion): https://github.com/apache/datafusion/issues/13448 https://github.com/apache/datafusion/issues/13448 Disclaimer I am on the PMC of Apache DataFusion, so am totally a fan boy.