8 ms·
Building SQLite with a small swarm
Hope some find this post interesting on my experience with parallel coding agents.
- deleted 8mo ago[deleted]
- scirob 8mo agoDid they pass all unit tests in the end ?
- comex 8mo agoIf it works, then it’s impressive. Does it work? Looking at test.sh, the oracle tests (the ones compared against SQLite) seem to consist in their entity of three trivial SELECT statements. SQLite has tens of thousands of tests; it should be possible to port some of those over to get a better idea of how functional this codebase is. Edit: I looked over some of the code. It's not good. It's certainly not anywhere near SQLite's quality, performance, or codebase size. Many elements are the most basic thing that could possibly work, or else missing entirely. To name some examples: - Absolutely no concurrency. - The B-tree implementation has a line "// TODO: Free old overflow pages if any." - When the pager adds a page to the free list, it does a linear search through the entire free list (which can get arbitrarily large) just to make sure the page isn't in the list already. - "//! The current planner scope is intentionally small: - recognize single-table `WHERE` predicates that can use an index - choose between full table scan and index-driven lookup." - The pager calls clone() on large buffers, which is needlessly inefficient, kind of a newbie Rust mistake. However… It does seem like a codebase that would basically work. At a large scale, it has the necessary components and the architecture isn't insane. I'm sure there are bugs, but I think the AI could iron out the bugs, given some more time spent working on testing. And at that point, I think it could be perfectly suitable as an embedded database for some application as long as you don't have complex needs. In practice, there is little reason not to just reach for actual SQLite, which is much more sophisticated. But I can think of one possible reason: SQLite has been known to have memory safety vulnerabilities, whereas this codebase is written in Rust with no unsafe code. It might eat your data, but it won't corrupt memory. That is impressive enough for now, I think.
- wedog6 8mo agoSQLite is tested against failure to allocate at every step of its operation: running out of memory never causes it to fail in a serious way, eg data loss. It's far more robust than almost every other library.
- sigmoid10 8mo agoUnfortunately it is not so easy. If rigorous tests at every step were able to guarantee that your program can't be exploited, we wouldn't need languages like Rust at all. But once you have a program in an unsafe language that is sufficiently complex, you will have memory corruption bugs. And once you have memory corruption bugs, you eventually will have code execution exploits. You might have to chain them more than in the good old days, but they will be there. SQLite even had single memory write bugs that allowed code execution which lay in the code for 20 years without anyone spotting them. Who knows how many hackers and three letter agencies had tapped into that by the time it was finally found by benevolent security researchers.
- gzread 8mo agoassuming your malloc function returns NULL when out of memory. Linux systems don't. They return fake addresses that kill your process when you use them. Lucky that SQLite is also robust against random process death.
- formerly_proven 8mo agoThat's not how Linux memory management works, there are no poison values. Allocations are deferred until referenced (by default) and when a deferred allocation fails that's when you get a signal. The system isn't giving you a "fake address" via mmap.
- mcculley 8mo agoMy interpretation of the GP comment is that you are saying the same thing. Linux will return a pointer that is valid for your address space mappings, but might not be safe to actually use, because of VM overcommit. Unixes in general have no way to tell the process how much heap can be safely allocated.
- losteric 8mo agoThis blog post doesn't say anything about your experience. How well does the resulting code perform? What are the trade-offs/limitations/benefits compared to SQLite? What problems does it solve? Why did you use this process? this mixture of models? Why is this a good setup?
- kyars 8mo agothe code has not been rigorously tested in all honesty, (this is mainly an experiment on agent orchestration as opposed to building a viable sqlite in rust) - The choice of two workers per model is purely pragmatic: I can't afford more. - I chose heterogeneous agents because it has not been done yet. There is no performance justification for this choice.
- mrmrcoleman 8mo agoTake a look at SQLite’s test coverage. It’s impressive: https://sqlite.org/testing.html https://sqlite.org/testing.html 590x the application code
- echelon 8mo agoThe fact that AI agents can even build something that purports to be a working database is also impressive. A small, highly experienced team steering Claude might be able to replicate the architecture and test suite reasonably quickly. 1-shotting something that looks this good means that with a few helping hands, small teams can likely accomplish decades of work in mere months. Small teams of senior engineers can probably begin to replicate entire companies worth of product surface area.
- ncruces 8mo agoThe other day I asked AI to one-shot an implementation of hyperbolic trig functions for double-double floats. I provided a repo (mine) that already implemented double-double arithmetic, trigonometry, and logarithms/exponentials, with plenty of tests. It produced something that looked this good. It had tests, it followed the style of the existing code base, etc. But it was full of shit and outright lies. After I reviewed it to fix deficiencies, I don't think there was anything left of the original. I had much more success the previous week using an AI to rubber duck the algorithms to implement trig. I am incredibly sceptical that just adding more loops — and less critical thinking/review — to brute force through a solution, is a good idea.
- kyars 8mo agoI push back on loops being insufficient because algorithms such as alpha evolve have already proved very effective.
- gf000 8mo agoApologies for the snark, but are you also impressed by `git clone` downloading a repository that is openly available on the internet? It can even do that in a loss-less way, instead of burning a bunch of tokens to get a bad, barely working half-copy. Don't get me wrong, I'm no AI hater, they are an impressive technology. But both AI-deniers and hypers need a reality check.
- samrus 8mo agoI cant quite tell if the tests that passed were sqlites own famously thorough test suite, or your own. If its sqlites suite then its great the models managed to get there, but one issue (without trying to be too pessimistic) is that the models had the test suite there to validate against. Sqlites devs famously spend more of their time making the tests than building the functionalities. If we can get AI that reliably defines the functionality of such programs by building the test suite over years of trial and error, then we'll have what people are saying
- kyars 8mo agoSorry for the ambiguity, it's not the test suite and I have updated the blog post to make that clear. I agree that building software where you do have this oracle is much easier than not having it and expecting the AI to build it.
- comrade1234 8mo agoWhat's the point of building something that already exists in open source. It's just going to use code that already exists. There's probably dozens of examples written by humans that it can pull from.
- embedding-shape 8mo agoWhat do you suggest we build instead, that hasn't already been done? I've been developing for decades, and I can't think of a single thing that hasn't already been kind of done either in the same or other language, or at least similar.
- prmph 8mo agoI want a language with: - the memory, thread safety, and build system of Rust - the elegant syntax of OCaml and Haskell - the expressive type system of Haskell and TypeScript - the directness and simplicity of JavaScript Think coding agents can help here?
- gcr 8mo agoNim comes close to what you want. Looking a bit further out, F# and Swift also come close.
- viraptor 8mo agoYou have conflicting requirements there - expressive type systems are not direct and simple. And elegant is subjective. But seriously though: have you tried to see how far you can get with the design right now? You can start iterating on it already, even if the implementation will lag.
- prmph 8mo agoI do not have conflicting requirements. Expressive type system ARE direct and simple. Expressive power is the ratio how strongly/clearly you can encode invariants to how complex and ceremonious the syntax of it needs to be. See how JS, a language usually seen as a middling/mediocre language, can distill the basic good parts of OOP into very direct and clear idioms? I can just create an object literal and embed simple methods on them that receive the "this" pointer and use it. The constructor would be just a regular function. None of the cruft of standard OOP. See how you define an enumerable union in TypeScript? Very simple. And yet I can think of many major languages that do not have this, certainly not with a lot of ceremony and complexity. And I can go on.
- bob1029 8mo ago> 84 / 154 commits (54.5%) were lock/claim/stale-lock/release coordination. Parallelism over one code base is clearly not very useful. I don't understand why going as fast as possible is the goal. We should be trying to be as correct as possible. The whole point is that these agents can run while we sleep. Convergence is non linear. You want every step to be in the right direction. Think of it more as a series of crystalline database transactions that must unroll in perfect order than a big pile of rocks that needs to be moved from a to b.
- philipswood 8mo agoI think we can now begin to experimentally test Conway's law and corollaries. Agreed, a flat set of workers configured like this is probably not the best configuration. Can you imagine what an all human team configured like this would produce?
- CuriouslyC 8mo agoOrchestration and autonomy are the things people get hyped about, but validation is the real bottleneck, and I'm pretty sure it's not amenable to complete automation. The people pushing orchestration the hardest are trying to get their users to validate for them, which taints the AI related open source ecosystem for everyone (sorry Steve/Peter!). I wrote a rant about this a while back to try and encourage people to be more responsible: https://sibylline.dev/articles/2026-01-27-stop-orchestrating-and-start-validating/ https://sibylline.dev/articles/2026-01-27-stop-orchestrating...
- k33n 8mo ago> There isn’t a great way to record token usage since each platform uses a different format, so I don’t have a grasp on which agent pulled the most weight lol
- kyars 8mo agoClaude code token tracking doesn't even work, for example. And Gemini also doesn't provide statistics, so I'm just being honest here.
- gmerc 8mo agoWhy do people fall for this. We're compressing knowledge, including the source code of SQLite into storage, then retrieve and shift it along latents at tremendous cost in a while loop, basically brute forcing a franken version of the original.
- foo42 8mo agoI agree. While I'm generally sympathetic to the idea that humans and LLM creativity is broadly similar (combining ideas absorbed elsewhere in new ways), when we ask for something that already exists it's basically just laundering open source code
- pjc50 8mo agoLicense laundering and the ability to not credit or pay the original developers.
- formerly_proven 8mo agoLaundering public domain code no less
- viraptor 8mo agoBecause virtually all software is not novel. For each single partially novel thing, there are tens of thousands of crud apps with just slightly different flow and data. This is what almost every employed programmer does right now - match the previous patterns and produce a solution that's closer to the company requirements. And if we can brute force that quickly, that's beneficial for many people.
- mpalmer 8mo ago> Because virtually all software is not novel. That isn't true, not by a long shot. Improvements happen because someone is inspired to do something differently. How will that ever happen if we're obsessed with proving we can reimplement shit that's already great?
- delegate 8mo agoGreat work! Obviously the goal of this is not to replace sqlite, but to show that agents can do this today. That said, I'm a lot more curious about the Harness part ( Bootstrap_Prompt, Agent_Prompt, etc) then I am in what the agents have accomplished. Eg, how can I repeat this myself ? I couldn't find that in the repo...
- kyars 8mo agohello, thanks! all of the harnessing is in this repo: https://github.com/kiankyars/parallel-ralph/ https://github.com/kiankyars/parallel-ralph/
- khazhoux 8mo agoI'm a heavy Cursor user (not yet on Claude) and I see a big disconnect between my own experience and posts like this. * After a long vibe-coding session, I have to spend an inordinate amount of time cleaning up what Cursor generated. Any given page of code will be just fine on its own, but the overall design (unless I'm extremely specific in what I tell Cursor to do) will invariably be a mess of scattered control, grafted-on logic, and just overall poor design. This is despite me using Plan mode extensively, and instructing it to not create duplicate code, etc. * I keep seeing metrics of 10s and 100s of thousands of LOC (sometimes even millions), without the authors ever recognizing that a gigantic LOC is probably indicative of terrible heisenbuggy code. I'd find it much more convincing if this post said it generated a 3K SQLite implementation, and not 19K. Wondering if I'm just lagging in my prompting skills or what. To be clear, I'm very bullish on AI coding, but I do feel people are getting just a bit ahead of themselves in how they report success.
- viraptor 8mo ago> cleaning up what Cursor generated What model? Cursor doesn't generate anything itself, and there's a huge difference between gpt5.3-codex and composer 1 for example.
- khazhoux 8mo agoWell, I've got it as Auto (configured by my company and I forget to change it). The list of enabled models includes claude-4.6-opus-high, claude-4.5-sonnet, gpt-5.3-codex, and a couple more.
- Philpax 8mo agoThat is probably Composer-1, which is their in-house model (in so much a fine-tune of an open-weights model can be called in-house). It's competent at grunt work, but it doesn't compare to the best of Claude and Codex; give those a shot sometime.
- viraptor 8mo agoAuto is not likely to choose the high quality models unless you really try for complex plans. Give the explicit models a try instead. It really makes a difference.
- small_model 8mo agoWould be better to choose a small subset of functionality and get that working as well as sqlite (or better) Then iterate that way. Context size is too small to work on such a large system.
- kyars 8mo agoI agree with you that choosing a subset and iterating to get an optimized version as good as SQLite or better is a better way to test and achieve more useful results. But with respect to the context size, Cursor has made agentic projects with over one million lines of code, so I would push back on that: https://cursor.com/blog/scaling-agents https://cursor.com/blog/scaling-agents
- diimdeep 8mo agoWhy do you think that it is a good idea to make it public ? It is obviously half hallucinated mostly broken unusable piece of low effort (on human part), with as much value as blurry image generated with stable diffusion that people now widely consider bad taste and slop.
- kyars 8mo agoI hope I did not give the impression that I wanted people to actually use this. I'm just using this as a test bench similar to how Anthropic made a C compiler with Claude, which of course they do not recommend you use.
- marxisttemp 8mo agoWho cares?
- kyars 8mo agoI hope someone
- rco8786 8mo agoNot a single comment about whether it actually works or not?
- gcr 8mo agoIt largely doesn't. The authors didn't attempt to run against SQLite's open test suite.
- kyars 8mo agoYou are right, I'm rectifying that now
- kyars 8mo agoSorry full transparency, I put my confidence in the fact that the model said it was passing all tests and had implemented most SQLite operations, but that was a mistake, so now I'm independently running tests.
- randomifcpfan 8mo agoInteresting to compare this to the in-progress project https://github.com/Dicklesworthstone/frankensqlite https://github.com/Dicklesworthstone/frankensqlite Which aims to match SQLite quality and provide new features (free encryption, multiple simultaneous writers, and bitflip resistance.)
- kyars 8mo agoThat project is definitely of higher quality than this one. For instance, this project does not have concurrency.
- simonw 8mo ago"Implements + tests against sqlite3 as oracle" That's the real unlock in my opinion. It's effectively an automated reverse engineering of how SQLite behaves, which is something agents are really good at. I did a similar but smaller project a couple of weeks ago to build a Python library that could parse a SQLite SELECT query into an AST - same trick, I ran the SQLite C code as an oracle for how those ASTs should work: https://github.com/simonw/sqlite-ast https://github.com/simonw/sqlite-ast Question: you mention the OpenAI and Anthropic Pro plans, was the total cost of this project in the order of $40 ($20 for OpenAI and $20 for Anthropic)? What did you pay for Gemini?
- kyars 8mo agoyes, in the order of $50 let's say, although with api I believe it would be in the hundreds Gemini is free, I don't even know if they have a paid plan?
- tonetheman 8mo ago[dead]
- MagicMoonlight 8mo agoWhy would you need 6 different models running across three providers? Just have a single one running, then you avoid all this nonsense around locking. And this is ultimately pointless, because it’s just a shitter SQLite. It’s nothing new. If you’re going to build something big like this, there needs to be a real business case You could already slop out a replica of SQLite if you wanted. But you don’t, because of the effort it would take to test and maintain it.
- kyars 8mo agoUltimately, this was an experiment with no intent to migrate to a production environment. Regarding single agents, for large projects one agent is too slow. So that's why developing multi-agent paradigms is compelling. I view SQLite as just an objective to attain and optimize for, but nothing more. I agree 100% that this is just a shittier SQLite.
- cadamsdotcom 8mo agoIf anyone is looking for ideas for these projects - it’d be great to be able to run macos applications on linux… Someone could have a swarm of agents build “wine for macos apps”.
- hashmak_jsn 8mo agoI discourage coding sqlite in Rust, Here are the reasons that sqlite developers mentioned: - Rust needs to mature a little more, stop changing so fast, and move further toward being old and boring. - Rust needs to demonstrate that it can be used to create general-purpose libraries that are callable from all other programming languages. - Rust needs to demonstrate that it can produce object code that works on obscure embedded devices, including devices that lack an operating system. - Rust needs to pick up the necessary tooling that enables one to do 100% branch coverage testing of the compiled binaries. Rust needs a mechanism to recover gracefully from OOM errors. - Rust needs to demonstrate that it can do the kinds of work that C does in SQLite without a significant speed penalty. https://sqlite.org/whyc.html#why_isn_t_sqlite_coded_in_a_safe_language_ https://sqlite.org/whyc.html#why_isn_t_sqlite_coded_in_a_saf...