9 ms·
Tracking down the 16-year-old WAL-reset SQLite bug
- tyho 2mo agoWhat a brutal bug. I'd never entertain a that bug in SQLite could be causing problems in code I wrote.
- hedgehog 2mo agoThe fun part is unless your app is pretty high traffic you might see the bug only once, or twice, and never have a reasonable way to even get close to a fix or know what your real exposure is.
- sgarland 2mo agoThat is the correct mentality, but sometimes, you’re wrong. I found (not entirely true; someone else [0] found it first, but my report from my company got traction [1] on it) a weird bug in ProxySQL several months ago involving mirroring and fast routing, where if you had a configuration that was illogical from the standpoint of documented behavior, it would duplicate queries, which led to super fun times for writes. I spent hours empirically trying various scenarios before concluding that no, it was a ProxySQL bug. [0]: https://github.com/sysown/proxysql/issues/2233 https://github.com/sysown/proxysql/issues/2233 [1]: https://github.com/sysown/proxysql/pull/5385 https://github.com/sysown/proxysql/pull/5385
- zbentley 2mo agoI mean, ProxySQL and SQLite are at very different tiers of reliability. I encountered multiple unreported data-corrupting (and some resource exhausting/connection mis-pinning) bugs in ProxySQL within a few months of using it for the first time, and I wasn’t using it for anything particularly complex or advanced—just a basic connection pool, no failover or caching/rewriting/replica awareness, but a lot of frontends and QPS.
- simonw 2mo ago> We funded the open-source SQLite VFS shim that helped isolate the race condition almost immediately, and will help track down similar bugs in the future. Interesting example of a company funding open source - in this case paying for the development of a new and very specific debugging tool.
- deleted 2mo ago[deleted]
- binhex 2mo agoYeah, tailscale seems to have leadership with their head on right, I agree with the way they handle a lot of things.
- devmor 2mo agoTheir CEO is a very nice and personable guy too. Has given me and others advice on random topics of his interest with no nonsense plenty of times.
- saghm 2mo agoYeah, this part also stuck out to me: > Because this wouldn’t be a quick or easy fix, we reached out to the SQLite developers for a professional support contract. This was a great decision. It gave us direct access to their deep expertise and experience, and we had many detailed technical conversations about our architecture and our incidents. They were willing to pay to get help solving the problem, and then pay again to make sure that the problem is easier to avoid in the future! That kind of long-term thinking seems pretty rare nowadays...
- dolmen 2mo agoWhich SQLite driver for Go does Tailscale use?
- sethops1 2mo agohttps://github.com/tailscale/sqlite https://github.com/tailscale/sqlite
- bobtheborg 2mo agoGreat read. So glad they took the time to tell this story. (And glad they, as a for profit corporation, took out a support contract with SQLite. I hope they continue to do so even though this problem is resolved.)
- calmingsolitude 2mo agoWell written post, really enjoyed reading it. > A single Go process exclusively accesses that database, and serves the control plane for those tailnets. This single-writer design is exactly how SQLite is meant to be used. This line led me to believe that the writer and checkpointing logic lived on the same database connection, so I was curious to find out how the data race occurred. However, the bug details on the SQLite page[0] outline that it can only ever occur if there are multiple connections open, so the writer and the checkpointer must have been on different threads. [0] https://sqlite.org/wal.html#the_wal_reset_bug https://sqlite.org/wal.html#the_wal_reset_bug
- maitrungduc 2mo ago[flagged]
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- pseudohadamard 2mo agoI don't know enough about the scale of Tailscale's operations to comment strongly on this, but if they're fairly significant shouldn't that have read "is exactly how MariaDB is meant to be used" or "exactly how Postgres is meant to be used"? SQLite has a "lite" in the name for a reason, but it's often pushed into places where it's being asked to do things it was never really designed for.
- zbentley 2mo agoThis particular bug doesn’t seem to arise from SQLite’s “lite” nature. It’s a TOCTOU inside the DB when applying WAL segments in a checkpoint, which is a pattern used in extremely similar ways by Postgres and MySQL. They don’t seem to have similar bugs, but I don’t think there’s any reason to believe that this is due to their being client/server rather than coordinated-file databases.
- ec109685 2mo agoWhile technically true as written, it seems to downplay the significance: > The bug is a data race with tight timing constraints. It is unlikely to occur in common use. A large customer did experience this corruption, so it's important for people with tailscale's setup update immediately. > The developers have never been able to reproduce the bug organically and had to add special testing logic to SQLite that deliberately triggers the circumstances of the the bug in order to verify that the issue has been fixed.
- Ariarule 2mo agoOdd not to highlight the sentence where they answer the obvious question "Why Tailscale in particular?": > They also explained why we were more likely to hit the bug than other SQLite users: we take manual control of the checkpointing process, and we checkpoint very aggressively. Even a bug triggered by a rare condition was bound to hit us eventually.
- dboreham 2mo agoQuick note that data corruption bugs that are impossible to reproduce are not uncommon (perhaps they're the norm). So some amount of head scratching trying to figure out a plausible scenario by which the system could get into the state represented by the smoking remains is often required. Then you attempt to force it into the supposed bad state by modifying code paths accordingly. So the approach used in this case is clever, but it's not particularly unusual in the world of data stores.
- 2mo ago
- riknos314 2mo ago> Whenever corruption occurred, we had to stop the control plane process on the shard while we repaired or restored the database. This was painful for tailnets on that shard, because their entire control plane disappeared during that recovery window. Gotta love single points of failure...
- kccqzy 2mo agoWhat are some solutions to avoid database corruption being single points of failure? I can’t think of any off the top of my head. I don’t think people typically consider database corruption to be a kind of failure common enough to design for, unless you have unusual requirements.
- AlotOfReading 2mo agoThe general answer to this is Byzantine consensus, which cryptocurrency blockchains are designed to solve. If your nodes are willing to fail a little more politely (e.g. no lying, immediately crashing, etc) you can use something cheaper like raft/paxos. But yeah, it's a lot cheaper to build a reliable system than it is to be resilient.
- spockz 2mo agoThe shard was already a way to make it not a single point of failure.
- Spivak 2mo agoThis is a great example of outages looking different from the perspective of the operator vs the user. Because there's many shards the blast radius of failure is contained to a small subset of users but for those users it's an outage. The way it's designed you can't lose any shards without impacting users. Compare to say Elasticsearch where it's possible to lose nodes and lose shards without the user noticing. One approach isn't universally better than the other.
- spockz 2mo ago
- deleted 2mo ago[deleted]
- pstuart 2mo agoI imagine the SQLite eschews AI generated code, but using it for testing (vulnerability, performance, etc) would seem like an easy win. I know their proprietary testing framework is their secret sauce so we may never know...
- d-us-vb 2mo agoRichard Hipp's recent talk at Software Should Work explains that AI agents have been testing SQLite and they've gotten a deluge of new bug reports from the fuzz-like testing they can do. But they do not do this in house; hobbyists and other organizations do this in their own internal agent-driven fuzzing.
- declan_roberts 2mo ago> Now we’re in summer, we’re confident that we’ve found the bug, that we understand it—and more importantly, that we’ve fixed it. This is the feeling I chase as a software engineer. It's the greatest motivator.
- jeffbee 2mo agoBlock device upfuckery layers are powerful against databases. Years ago some colleagues wrote one that provides most of the hazards described by "Parity Lost and Parity Regained"[1] to test FoundationDB, which immediately uncovered several flaws in a project that described itself as well-tested. It's easy to do this with all the probing features that Linux (and others) provide today. 1: https://www.usenix.org/legacy/event/fast08/tech/full_papers/krioukov/krioukov.pdf https://www.usenix.org/legacy/event/fast08/tech/full_papers/...
- jnwatson 2mo agoI'm definitely going to use the word "upfuckery" instead of fault injection the next time I need it.
- dilyevsky 2mo agoWith opus 4.7-ish to fable 5, immediately after release - before they locked it down, it was shockingly easy to find crash bugs in a lot of very heavily used DBs and other software.
- jeffbee 2mo agoI'll accept crashing in preference to a database acknowledging transaction that weren't durably committed.
- dilyevsky 2mo agoFor sure! I haven't had the opportunity to burn many tokens on this but given how quickly it was able to fuzz it into crashing i'm sure it can manage.
- inigyou 2mo agoDid you personally find them?
- Zenul_Abidin 2mo agoSimilar bug to the one that plagued Codex until 3 months ago.
- ameliaquining 2mo agoDo you happen to have a link?
- sandeepkd 2mo ago>In our control plane, we take manual control of the checkpoint process so we can run fast and consistent backups. > running boring technology in a non-standard way is a risk. It was a good read and reminder that the industry is loosing experts gradually. I am not a DBA and yet I have heard about this behavior at least couple times in the past as something to avoid. Its just one of those things which didnt get a chance to be documented cause experts avoided it and regulars didn't get into
- sandeepkd 2mo agoIn other words there exists a concept of HOT and COLD backups for this reason only.
- inigyou 2mo agoI've never heard of any reliability reason you shouldn't run checkpoints whenever you want - only performance reasons. Can you elaborate? Backing up sqlite by copying the file (e.g. rsync) while it's open is a surefire way to eventually get corruption caused by a race condition, but it seems like tailscale wasn't doing that. They were probably using the proper sqlite backup API. But you don't need checkpoints for consistency and I think the backup API will not copy both the old and new versions of pages just because they're in the WAL, in other words I think checkpointing then backing up should give you the same pages as backing up without checkpointing. So the whole thing seems unnecessary.
- sandeepkd 2mo agoMy generalist reading of this (which is applicable in lot of cases) is that a dedicated flow is handling the house keeping part of the job, aka resetting the checkpoint after writing the WAL to disk. These one off house keepers are common in a lot of softwares and they expect to work alone, they are tested to work alone. They always have a set of ritual, rules and order in which they do all the things. Now if some other thread takes off some of those jobs then they break this routine for the house keeping job and this new thread/person may not always know what else has to be done before and after this one particular job for the sake of completeness. On second question of why the tailscale developers did it, its possibly for the same reason why they invested this much into debugging this issue. Some one believed the current behavior did not fit into their architecture, they want to be more performant and take control over things. A big part of me considers this is a required exercise to try, grow and learn. The only thing they could have for improvement would be to have these old hands on architect kind of folks on their team who might have hinted/pointed them to the problem a long before. Challenge/chances are that these older folks would have even stopped them from going in this direction in the design phase itself.
- myshapeprotocol 2mo ago[flagged]
- LgWoodenBadger 2mo agoMaybe it's just me, but the explanations of the cause don't align. One clue was that during corruption incidents, our metrics showed that SQLite would report copying more pages from the WAL file than were actually available. If there are 10 pages in the WAL file and 20 pages get copied to the database, something is clearly wrong. vs it thinks some of the pages have been copied from the WAL into the main database file, but they haven’t. Those pages never get written to the database file, and that data is permanently lost. The first says "more were copied than existed" but the second says "fewer were copied than should have been." Like I said, it's probably just me interpreting something incorrectly.
- surgical_fire 2mo agoMy interpretation is that they haven't been copied because they didn't exist? If you have 10 pages and it tries to copy 20, either those 10 pages wouldn't really be copied, or bogus data would be written. That's how I read at least. Those things are not mutually exclusive.
- ball_of_lint 2mo agoOr you could have 10 pages, it actually copies 9, and reports 20 anyways.
- surgical_fire 2mo agoThat's also possible
- nearlyepic 2mo agoI haven't looked into the actual code fix, but given the "reset" name I have to think it has to do with SQLite "thinking" it has copied more pages than it actually did. i.e. The checkpoint starts, and a write hits after the modifications to data structures have been done but before the data has actually been put in the database. The process starts over again, but doesn't undo the changes it made to indexes etc. Hence the db thinks it holds pages that don't exist. That's my interpretation, anyways.
- deleted 2mo ago[deleted]
- andai 2mo agoSQLite: 92 million lines of tests Dijkstra: Tests can only prove the presence of bugs, never their absence!
- otterley 2mo agoEveryone knows that tests don't prevent all bugs. But they are very good at preventing known bugs from recurring in the future.
- layer8 2mo agoEveryone doesn’t seem to know that, because tests are often cited as a way to ensure that AI-generated code is correct.
- andai 2mo agoI was very excited about formal proofs, which are now very cheap to produce, in service of validating AI generated code. But I had a funny experience recently where an agent implemented an entire feature completely wrong (exactly backwards, actually, in a way that defeated the purpose, introduced security issues etc.). It happily supplied tests for the new functionality, and all the tests passed. What I realized was, even formal verification wouldn't have helped here -- it would have just written a mathematical proof that the incorrect functionality was correctly implemented! So there's a gap here, where first, the human's intention needs to be formally specified (by the human, or at least the human needs to be able and willing to verify it), and then the slopswarm can hack away at it...
- otterley 2mo ago> first, the human's intention needs to be formally specified This never changed with AI; in fact, I think it made this need more visible than it ever had been before. You can't get away with not being able to describe in detail what you want. As with working with humans, any ambiguities will be interpreted, and not always in the way you hoped.
- 2mo ago
- hn3ufz62f7 2mo agoLearned something new today, thanks
- bch 2mo agoThis was really, really interesting - what a triumphant adventure. A few (very, very, very pedantic) things that stood out: > We wanted a way to restore service that didn’t involve rolling back to the last known-good backup (which would lose a lot of data) or repairing the known-corrupted database (which was potentially risky). (Emphasis mine) - it would be "risky", not "potentially risky" - then the "calculated risk period" starts and it's "potentially problematic". In the SQLite report[0] (11.2) I wish they downplayed this less - a mention of the rarity, then technical details - I'm friendly with a few of the devs/previous-devs, have the utmost respect for their skill and accomplishments (and by extension, faith that the developers I do not personally interact with are also excellent), appreciation and fondness for the huge accomplishment that is SQLite, and on and on... this is world-class work. Maybe section 11.2 wasn't really aimed at me, or I'm too critical. To be fair to all involved, what a minor quibble for such an interesting problem/fix. I hope my comment isn't a fly in the ointment. Last bugfix point[1] - ugh. What a sinking feeling that must've been to deploy a fix then be flooded with not-green - and a lesson[2] against smuggling other changes in a changeset "just because we're already here"? Happy it turned out non-catastrophic, but did result in a rare (not remembering other instances of top of head) recall[3] from SQLite. That it was throwing errors at the same time SQLite and Tailscale were testing the other WAL-issue bug must've upset some stomachs for a moment. [0] https://sqlite.org/wal.html#the_wal_reset_bug https://sqlite.org/wal.html#the_wal_reset_bug [1] https://tailscale.com/blog/sqlite-wal-reset-bug#fixed-with-a-false-alarm https://tailscale.com/blog/sqlite-wal-reset-bug#fixed-with-a... [2] Nobody conceptually learned anything here - we're all just reminded of what we know: that sometimes "perfect storms" do actually occur. [3] https://sqlite.org/releaselog/3_52_0.html https://sqlite.org/releaselog/3_52_0.html
- dzonga 2mo agoyou gotta admire the power of using json/b and simple KV stores. so many people sleep on that.
- w10-1 2mo agoThe irony is that the SQLite developers get a support contract iff someone runs off the path in anger and finds an ancient bug. But perhaps that's part of what make it a quality team: devotion thriving without adverse incentives.
- rmunn 2mo agoWhile you might be correct, I wouldn't necessarily assume the "and only if" part of your statement. There might be companies that choose to proactively purchase support contracts. And there are companies that have paid $150K/year for https://sqlite.org/consortium.html https://sqlite.org/consortium.html access. Which means, among other things, that they get first priority for any needs they have: > Consortium members have the guaranteed, undivided attention of the SQLite developers for 23 staff-days per year and for as much additional time above and beyond that amount that the core developers have available. There are no arbitrary limits on contact time. The consortium will never be over-subscribed. New SQLite developers will be recruited and trained as necessary to cover the 23 day/year support commitment. The SQLite home page lists five companies that have paid for consortium access. I can easily imagine that there are more who don't want to pay $150K/year but would pay $1.5k/year, proactively, to get "private, expert email advice from the developers of SQLite" when they need it.
- ChuckMcM 2mo agoGreat writeup, and it was great to see them step in an pay the developers of SQLite to help them fix the bug. I get tired of corporations asking open source authors to fix problems that affect the corporation for free. And while I'm sure it was frustrating for folks to have these outages, I find such puzzles pretty fun to get to the bottom of.
- manoji 2mo agoSuch a good write up . Having explored a little bit of sqlite internals for a codecrafters challenge i was mildly happy i could follow along what was happening .
- procflora 2mo agoVery nice article, and I appreciate SQLite's explanation of the bug too. And how extremely cool Tailscale appears to have been about it (paying for the VFS shim, etc.). I'd have liked to have heard more about the decision to checkpoint so frequently that put them on this path though. Presumably that's to keep the WAL tiny for very fast recovery. Trying to mitigate some of the deleterious effects of inserting a DBMS into your network layer, I suppose? Tricky stuff. Wonder how that compares to typical etcd snapshot frequencies too.
- buggymcbugfix 2mo agoAgreed, I was also wondering about this. Maybe the aggressive checkpointing was for preventing WAL-overflow?
- nujabe 2mo agoWas wondering the same, doesn’t mention if they tried checkpointing less frequently
- wanderr 2mo agoThis was a great technical writeup and very interesting to read, but it's not clear to me why once the suspected source of the bug was identified, they seemingly didn't build a automated way to trigger the condition? It seems like that could have cut down on the uncertainty of whether the fix worked over a painfully long period of time.
- Neywiny 2mo agoThe sqlite dev team did. It's in the article. > It could exist that long because it was rare—so rare, the SQLite developers had to add code to deliberately trigger it in their testing environments.
- wanderr 2mo agoYes of course I mean the tailscale team. Having an independent way to repro the bug besides waiting for it to happen in prod seems pretty basic eng best practice.
- asveikau 2mo ago> SQLite corruption is possible, but it’s highly unusual and not something you should encounter in normal operation If there's a hardware failure, for example a flaky SD card, it's not out of the question. A mobile app with a lot of usage will see it. (Yes, I know this appears to be a server use case.)
- devmor 2mo agoI have encountered this exactly once, and it was in fact a flaky SD card - running a small web service off of a raspberry pi, with the SQLite DB stored on the SD.
- e28eta 2mo agoThe sentence you’ve quoted actually links to a list of ways to corrupt SQLite, which I think is interesting in its own right: https://www.sqlite.org/howtocorrupt.html https://www.sqlite.org/howtocorrupt.html I believe your flaky SD card is category 4, Disk Drive and Flash Memory Failures.
- jbs789 2mo agoAs a simple user of SQLite, I think this level of debugging is incredible and appreciate being a beneficiary of the ecosystem and hard work of others. Thank you!
- danpalmer 2mo agoGlad this got found and fixed, but I continue to be astounded at the amount of work people put into making SQLite do things that would be much simpler with other systems.
- antonvs 2mo agoYeah it’s a pity that wasn’t addressed in the article. This is a little like “we shot ourselves in the foot and then performed surgery on our foot, and everything is resolved now.”
- bluehatbrit 2mo agoThat assumes that they do actually believe it's a mistake. They didn't explain the reasons they've gone for this architecture in much/any detail. I'd be interested in hearing them talk more about that in the future. Perhaps you or I would make a different decision based on the aims that lead them there. But there isn't enough information to say whether or not their decision was a mistake, even if it has lead to a peculiar bug. It may be perfectly legitimate and we just don't know some of the constraints they had.
- antonvs 2mo agoRight, the point is it's a kind of elephant in the room - they even allude to it not being a common use case - but they say nothing about the rationale. That leads me to suspect it was a kind of "seemed like a good idea at the time" situation. They also touch on this when they say that it worked for them for a long time. This is a pretty classic symptom of a system that was designed a certain way early on and then runs into issues as the system grows. One can argue that this was due to a bug, but it's a bug that they shouldn't really have had to deal with - a consequence that the design opened them up to.
- waterTanuki 2mo agoTailscale likely deals with a lot of security-sensitive traffic, given the nature of the service. I'm guessing one of the requirements they were given was to have a tiny blast radius in case encryption keys got leaked, and that meant isolating each customer's tailnet (meta)data to it's own sqlite db rather than letting everyone share a postgres cluster.
- gwking 2mo agoI wonder if this also affected litestream disproportionately, because litestream also inserts itself into the checkpoint process.
- why-el 2mo agoI would assume they actually moved off of litestream here no? otherwise how can their frequent manual checkpointing even succeed when litestream locks for the same behavior.
- gwking 2mo agoI meant does the rare SQLite bug tend to manifest in litestream for the same reason it showed up in Tailscale.
- wwilson 2mo agoWas curious so we checked and yep, Antithesis finds this bug in about 15 minutes. Will post a repro/writeup here soon.
- minimaltom 2mo agoI'm equal parts intrigued and skeptical- I guess if the prompt doesn't lead on there is a bug there then I'm impressed.
- wwilson 2mo agoStay tuned!
- wwilson 2mo agoRECEIPTS: https://antithesis.com/blog/2026/wal-reset-bug/ https://antithesis.com/blog/2026/wal-reset-bug/
- truetraveller 2mo agoOh wow. I checked your site, but I (still) don't see the ability for average joes (and their pet AI agents) to signup. Could be huge, even if just offered in 5 minute slots. Is your email contact still valid in your profile? Want to reach out (Moose is the name). FYI, got a job at Antithesis a while ago after some rigorous interviewing.
- wwilson 2mo agoWe are working on getting self-serve signups to you ASAP. For now, just email us. We don’t bite.
- wwilson 2mo agoDiscussion here: https://news.ycombinator.com/item?id=49277799 https://news.ycombinator.com/item?id=49277799
- inigyou 2mo ago
- grahar64 2mo agoAwesome write up. Finding these bugs in such a well used piece of software is like donating to humanity
- khernandezrt 2mo agoShow us the $$ numbers! Or an estimate at least.
- password4321 2mo agoI lost my Android SMS DB way back in the day (20+ years? dang I'm getting old) because the SMS app's fix when opening the DB detected any issue was to delete it and start fresh.
- buggymcbugfix 2mo agolol good old ostrich algorithm
- throw0101a 2mo agoSee perhaps recent video "Reliability Lessons From SQLite - Richard Hipp | SSW 2026": > Abstract: SQLite is a C-language library that implements a self-contained, in-process relational database engine supporting full-featured SQL, an advanced query planner, and ACID transactions. By many estimates, SQLite is the most widely used software library in the world today. > Over its 26-year history, SQLite has gained a reputation as software that "just works". This talk goes over the design choices and development practices that have, at least in the opinion of the lead developer, resulted in that reputation. * https://www.youtube.com/watch?v=V_qzqY1bb7I https://www.youtube.com/watch?v=V_qzqY1bb7I
- positive-spite 2mo ago"Tailscale Traces Database Corruption to 16y/o ..." Would have been a superior headline *sarcasm
- positive-spite 2mo agoThe title changed... Now my joke works even worse :(
- klaas- 2mo ago> This investigation is a useful reminder: running boring technology in a non-standard way is a risk. The common paths and standard configurations are incredibly well-tested and reliable. Most people use SQLite in a standard configuration and never face this sort of issue. Everything we were doing was a public, documented, supported configuration—but by taking manual control of the checkpointing process and running at our own aggressive pace, we stepped off the well-trodden operational path. I feel like they missed a key takeaway from their own argument here, they should not be running a non-standard configuration :)
- perching_aix 2mo agoThis reminds me why RFC 2119 includes the description it has regarding SHOULD. Like this is a great practical takeaway, sure, but eh.
- raggi 2mo agoWe have very good reasons for our checkpointing model, related to our backup + disaster recover strategy, along with resource cost. It might be worth writing about one day, so I'll not give away all the details, but in very short form, we organize a backup strategy that has minimal pause time, avoids doubling the page cache cost of the database, and enables extremely fast byte-copy restores in disaster recovery.
- inigyou 2mo agookay but explain why you are using sqlite and copying the file to S3 instead of using any client/server DB and its online backup feature?
- XorNot 2mo agoI see tailscale also use the pure Go SQLite conversion so I hope this fix will land there soon, it's rapidly become one of my favorite packages for self-contained tools.
- catapart 2mo agoAs others have said: great article! I did find myself wanting them to get to the point, but once they started describing the bug and the fix, it was very satisfying. I'm very happy there are companies out there on the frontiers of functionality not only funding fixes and debugging measures, but taking the time to write up the details so we can all benefit. Tailscale just moved up in my priorities list. Was going to host my next website with hostinger, but now I'm going to at least try to run a personal server with tailscale to make it public. I might not be able to figure it all out, and may end up going with the VPS route, but this gave me some appreciation for the company that makes me willing to try the less familiar method.
- inigyou 2mo agoI don't think Tailscale can do that though? It can't open a port from the internet to your private network.
- theMonsterpus 2mo agoIt can, but with some specifics about DNS names and supported ports: https://tailscale.com/docs/features/tailscale-funnel https://tailscale.com/docs/features/tailscale-funnel
- shye 2mo agohttps://tailscale.com/docs/reference/tailscale-cli/funnel https://tailscale.com/docs/reference/tailscale-cli/funnel
- deepsun 2mo agoNormal code has 50% to 90% ratio of code coverage by unit-tests. Dynamic-typed languages (Python, Ruby) usually require more, like 100% - 120%. SQLite has 59,000% ratio [1] Yet it didn't help for a bug to left unnoticed for 16 years :( I don't know what we can do for the industry. I doubt one can formally verify a project like SQLite, and keep it maintainable. https://sqlite.org/testing.html https://sqlite.org/testing.html
- inigyou 2mo agoOne might be able to formally verify the sqlite pager layer, where this problem occurred. It's just writing pages to disk and reading pages from disk, and providing ACID guarantees.
- inigyou 2mo agoThis reaffirms my belief that SQLite is not well suited for systems with significant concurrency. It replaces fopen, not postgres. Although this corruption is a rare bug and sqlite is usually extremely stable, it's usually not worth it from a performance and features standpoint either. Here they were trying to do a backup by forcing a checkpoint and then copying the file. Systems like postgres let you do online continuous backups.
- q3k 2mo agoYeah, I also feel 'just use SQLite [no matter what]' is just the pendulum swinging hard after the 'just use mongodb [no matter what]' of yesteryear. It's so sometimes just performative. I remember when Tailscale had a similar performative approach with 'just use a JSON file on disk'. Then etcd. Then SQLite. Like sure, you can keep picking the absolite mininum technology for your needs and then change it every couple of years... Or you could just immediately go with a solid Postgres (or Postgres-like setup, eg. yugabyte) and save yourself a bunch of faffing about with weird solutions and migrating between them. But I guess that doesn't drive engagement on your blog.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- sealeck 2mo agoPostgres has also had pretty serious bugs, eg fsyncgate
- lima 2mo agoThere's also the more fundamental issue that Postgres - unlike something like Yugabyte - does not use a distributed consensus algorithm for writes and can lose committed writes during network partitions. But of course, neither does Tailscale's weird DIY contraption.
- getnormality 2mo agoI wonder what happens when you give Claude the old version and the logs and ask it to find the bug.
- nujabe 2mo agoWhat logs though? They had to make multiple deployments to get the right log traces in the first place.
- tizerluo 2mo ago[flagged]
- JohnBooty 2mo agoExcellently-written article. Bravo!
- vluft 2mo agohuh. ran into almost-certainly this, but blamed it on litestream and rearchitected a bit as a result. will have to see if I can reproduce the issue as we were using with the patched sqlite
- chewbacha 2mo agoWild to have worked in the industry long enough that 16 years doesn’t feel that long ago.
- quarkcarbon279 2mo agolove the simplicity of this article. Reminds me so much of foundational software engineering. i love databases def not cosmosdb
- deleted 2mo ago[deleted]
- anitil 2mo agoIt says a lot about sqlite that a bug becomes front-page news on HN. I'm impressed that Tailscale took this seriously enough to engage with a commercial support contract. I'd love to work for a company that cared so much about correctness.
- Lio 2mo agoIndeed, SQLite has got to be one of the best tested pieces of software with famously 100% test coverage. https://sqlite.org/testing.html https://sqlite.org/testing.html It's amazing that a bug could exist for 16 years but it is sobering.
- sorry_i_lisp 2mo agoI think FoundationDB has that crown. https://antithesis.com/company/about/ https://antithesis.com/company/about/
- stillpointlab 2mo agoIt gives me a warm feeling when companies invest in open source support in this way. Helping great projects get even better is somehow better than releasing yet another project.
- madhu_ghalame 2mo ago[dead]
- punnerud 2mo agoChecked and this 100% compatible SQLite3 database (also C-API) did not contain the bug: https://github.com/punnerud/mpedb https://github.com/punnerud/mpedb (Disclaimer: My own project) And this statement is wrong in the article: “ Because SQLite is a single-writer database with serialisable transactions, our transaction history was completely linear and deterministic. (This wouldn’t be true in a multi-writer database like Postgres or MySQL.)” Actually possible in mpedb to replay multi-writer, and actually better than SQLite3. Try to reply now() in a statement, that is not deterministic in SQLite but is in MPEdb.
- looperhacks 2mo agoAt least mention that it's your own project ...
- punnerud 2mo agoAdded, thanks
- sdiazthomas 2mo ago[flagged]
- fenestella 2mo ago[dead]
- colomo 2mo agoFor this particular category of bug, SQLite's existing testing methodology is demonstrably outclassed by modern deterministic concurrency testing. https://antithesis.com/blog/2026/wal-reset-bug/ https://antithesis.com/blog/2026/wal-reset-bug/
- Kwantuum 2mo agoIf you already know the bug is caused by a race between writing and checkpointing then it's easy. I'm unsure how the linked article makes your point at all though.
- jkwang 2mo ago[flagged]
- tangsoupgallery 2mo ago[flagged]
- RinatNabiev 2mo ago[dead]
- miki123211 2mo ago> Nobody wanted us to spend six months looking for bugs in SQLite Most companies wouldn't. Instead, they'd fire the weirdo that came up with the idea of using some weird db, and switch to Postgres like God intended. I'm not saying either is right, this is not a criticism of Tailscale and their approach, paying the Sqlite maintainers to fix a real bug is commendable, but it certainly doesn't inspire confidence in the "Sqlite in production" hype train. While Sqlite is indeed boring technology for single-user SQL DBs, for traditional CRUD and network services, Postgres seems to be a much better trodden and much safer path.
- opengrass 2mo agoI wish Microsoft wrote an apology like that.
- richard_gg 2mo ago[dead]
- five9llc 2mo ago[flagged]
- peter_d_sherman 2mo ago>"...and then we discovered an unexpected clue. We wanted a way to restore service that didn’t involve rolling back to the last known-good backup (which would lose a lot of data) or repairing the known-corrupted database (which was potentially risky). To do this, we built a transaction logging pipeline. We streamed every SQL statement that modified the database to a separate log file. Because SQLite is a single-writer database with serializable transactions, our transaction history was completely linear and deterministic. (This wouldn’t be true in a multi-writer database like Postgres or MySQL.) Replaying those transactions against the latest known-good backup should restore the database to its most recent state, safely bypassing the corruption. [...] This pipeline worked, but then it did something even better: it gave us a clue. [...] To understand what was happening during these faulty checkpoints, the SQLite developers created a new debugging tool for the virtual filesystem layer. [...] To help diagnose our problem, the SQLite developers created a wrapper around the virtual filesystem that writes additional tracing information and logs about changes to the database. [...] After our next corruption incident, the additional logs from the new tmstmpvfs shim allowed the SQLite developers to find and fix the bug: a rare data race in the SQLite source code between a checkpoint and a write transaction." Great article! Software Engineering lessons (that repeat in this article!): So called "Heisenbugs" (bugs that make it past developer test harnesses and a company's Quality Assurance (QA) team) that show up post-deployment intermittently and can't be reproduced locally, occur because one or more of the following factors: 1) The lack of Determinism in a software process or processes. 2) The lack of appropriate logging. 3) The lack of the ability to replay a software process, step by exact step, state by exact state, as it has occurred in the field (occurs as an effect of #1 and/or #2). 4) Multi-threaded code; i.e., multiple threads giving rise to race conditions or other very specific intermittent combinatorial/permutational conditions caused by multiple threads and specific sections of code, which due to very large numbers of permutational timing possibilities, were not or could not be exactly tested for in development... Anyway, great article! A must-read for any Sr. Software Engineer, or any developer that wrestles with hard-to-find-and-fix bugs in the field...
- Ruca_AI 2mo ago[flagged]
- arendtio 2mo agoEverybody here knows that SQLite is not production-ready and does not scale. This is just proof that it can't be used in real-world applications. SCNR
- maxman88 2mo ago[flagged]