4 ms·
If it works, then it’s impressive. Does it work? Looking at test.sh, the oracle tests (the ones compared against SQLite) seem to consist in their entity of th
by comex 8mo ago
If it works, then it’s impressive. Does it work? Looking at test.sh, the oracle tests (the ones compared against SQLite) seem to consist in their entity of three trivial SELECT statements. SQLite has tens of thousands of tests; it should be possible to port some of those over to get a better idea of how functional this codebase is.
Edit: I looked over some of the code.
It's not good. It's certainly not anywhere near SQLite's quality, performance, or codebase size. Many elements are the most basic thing that could possibly work, or else missing entirely. To name some examples:
- Absolutely no concurrency.
- The B-tree implementation has a line "// TODO: Free old overflow pages if any."
- When the pager adds a page to the free list, it does a linear search through the entire free list (which can get arbitrarily large) just to make sure the page isn't in the list already.
- "//! The current planner scope is intentionally small: - recognize single-table `WHERE` predicates that can use an index - choose between full table scan and index-driven lookup."
- The pager calls clone() on large buffers, which is needlessly inefficient, kind of a newbie Rust mistake.
However…
It does seem like a codebase that would basically work. At a large scale, it has the necessary components and the architecture isn't insane. I'm sure there are bugs, but I think the AI could iron out the bugs, given some more time spent working on testing. And at that point, I think it could be perfectly suitable as an embedded database for some application as long as you don't have complex needs.
In practice, there is little reason not to just reach for actual SQLite, which is much more sophisticated. But I can think of one possible reason: SQLite has been known to have memory safety vulnerabilities, whereas this codebase is written in Rust with no unsafe code. It might eat your data, but it won't corrupt memory.
That is impressive enough for now, I think.
- wedog6 8mo agoSQLite is tested against failure to allocate at every step of its operation: running out of memory never causes it to fail in a serious way, eg data loss. It's far more robust than almost every other library.
- sigmoid10 8mo agoUnfortunately it is not so easy. If rigorous tests at every step were able to guarantee that your program can't be exploited, we wouldn't need languages like Rust at all. But once you have a program in an unsafe language that is sufficiently complex, you will have memory corruption bugs. And once you have memory corruption bugs, you eventually will have code execution exploits. You might have to chain them more than in the good old days, but they will be there. SQLite even had single memory write bugs that allowed code execution which lay in the code for 20 years without anyone spotting them. Who knows how many hackers and three letter agencies had tapped into that by the time it was finally found by benevolent security researchers.
- gzread 8mo agoassuming your malloc function returns NULL when out of memory. Linux systems don't. They return fake addresses that kill your process when you use them. Lucky that SQLite is also robust against random process death.
- formerly_proven 8mo agoThat's not how Linux memory management works, there are no poison values. Allocations are deferred until referenced (by default) and when a deferred allocation fails that's when you get a signal. The system isn't giving you a "fake address" via mmap.
- mcculley 8mo agoMy interpretation of the GP comment is that you are saying the same thing. Linux will return a pointer that is valid for your address space mappings, but might not be safe to actually use, because of VM overcommit. Unixes in general have no way to tell the process how much heap can be safely allocated.
- deleted 8mo ago[deleted]
- camgunz 8mo agoI'm not impressed: - if you're not passing SQLite's open test suite, you didn't build SQLite - this is a "draw the rest of the owl" scenario; in order to transform this into something passing the suite, you'd need an expert in writing databases These projects are misnamed. People didn't build counterstrike, a browser, a C compiler, or SQLite solely with coding agents. You can't use them for that purpose--like, you can't drop this in for maybe any use case of SQLite. They're simulacra (slopulacra?)--their true use is as a prop in a huge grift: tricking people (including, and most especially, the creators) into thinking this will be an economical way to build complex software products in the future.
- gf000 8mo agoAlso, the very idea is flawed. These are open-source projects and the code is definitely part of the training data.
- tux3 8mo agoThat's why our startup created the sendfile(2) MCP server. Instead of spending $10,000 vibe-coding a codebase that can pass the SQLite test suite, the sendfile(2) MCP supercharges your LLM by streamlining the pipeline between the training set and the output you want. Just start the MCP server in the SQLite repo. We have clear SOTA on re-creating existing projects starting from their test suite.
- viraptor 8mo agoThis would be relevant if you could find matching code between this and sqlite. But then that would invalidate basically any project as "not flawed" really - given GitHub, there's barely any idea which doesn't have multiple partial implementations already.
- criemen 8mo agoEven if was copying sqlite code over, wouldn't the ability to automatically rewrite sqlite in Rust be a valuable asset?
- IshKebab 8mo ago> I think the AI could iron out the bugs, given some more time spent working on testing I would need to see evidence of that. In my experience it's really difficult to get AI to fix one bug without having it introduce others.
- simonw 8mo agoHave it maintain and run a test suite.
- olmo23 8mo agoIIRC the official test-suite is not open-source, so I'm not sure how possible this is.
- SQLite 8mo agoYou do not recall correctly. There is more than 500K SLOC of test code in the public source tree. If you "make releasetest" from the public source tarball on Linux, it runs more than 15 million test cases. It is true that the half-million lines of test code found in the public source tree are not the entirety of the SQLite test suite. There are other parts that are not open-source. But the part that is public is a big chunk of the total.
- FeistySkink 8mo agoOut of curiosity, why aren't all tests open source?
- graemep 8mo agoOne set of proprietary tests is used in their specialist testing service that is a paid for service.
- FeistySkink 8mo agoWhat is that service used for besides SQLite?
- embedding-shape 8mo agoOne could assume also for Fossil.
- tonyarkles 8mo agoIt's still SQLite, they just need to make money: https://sqlite.org/prosupport.html https://sqlite.org/prosupport.html Edit: also this: > TH3 Testing Support. The TH3 test harness is an aviation-grade test suite for SQLite. SQLite developers can run TH3 on specialized hardware and/or using specialized compile-time options, according to customer specification, either remotely or on customer premises. Pricing for this services is on a case-by-case basis depending on requirements.
- alt187 8mo ago> But I can think of one possible reason: SQLite has been known to have memory safety vulnerabilities, whereas this codebase is written in Rust with no unsafe code. I've lost every single shred of confidence I had in the comment's more optimistic claims the moment I read this. If you read through SQLite's CVE history, you'll notice most of those are spurious at best. Some more context here: https://sqlite.org/cves.html https://sqlite.org/cves.html
- ii41 8mo agoI am using sqlite in my project. It definitely solves problems, but I keep seeing overly arrogant and sometimes even irresponsible statements from their website, and can't really appreciate much of their attitude towards software engineering. The below quote from this CVE page is one more example of such statements. > All historical vulnerabilities reported against SQLite require at least one of these preconditions: > 1. ... > 2. The attacker can submit a maliciously crafted database file to the application that the application will then open and query. > Few real-world applications meet either of these preconditions, and hence few real-world applications are vulnerable, even if they use older and unpatched versions of SQLite. This 2. precondition is literally one of the idiomatic usage of sqlite that they've suggested on their site: https://sqlite.org/appfileformat.html https://sqlite.org/appfileformat.html
- rstuart4133 8mo ago> That is impressive enough for now, I think. There are lot of embedded SQL libraries out there. I'm not particularly enamoured with some of the design choices SQLite made, for example the "flexible" approach they take to naming column types, so that isn't why I use it. I use it for one reason: it is the most reliable SQL implementation I know of. I can safely assume if file corruption, or invariants I tried to keep aren't there, it isn't SQLite. By completely eliminating one branch of the failure tree, it saves me time. That one reason is the one thing this implementation lacks - while keeping what I consider SQLite's warts.