5 ms·
Honker – Durable queues, streams, pub/sub, and cron scheduler in a SQLite file
- mergisi 5mo ago[flagged]
- EvanAnderson 5mo agoPrior discussion a few days ago: https://news.ycombinator.com/item?id=47874647 https://news.ycombinator.com/item?id=47874647
- itopaloglu83 5mo agoIt’s an interesting approach and can be quite fun to use for new projects. > How it works: honker polls SQLite’s PRAGMA data_version every millisecond. That’s a monotonic counter SQLite increments on every commit from any connection, journal mode, or process — a ~3 µs read for a precise wake signal.
- arlobish 5mo agoAt the end it says: "pg-boss and Oban are the Postgres-side gold standards" -- but Oban supports SQLite now too https://github.com/oban-bg/oban https://github.com/oban-bg/oban
- odie5533 5mo agoThere's also Graphile Worker. https://github.com/graphile/worker https://github.com/graphile/worker
- tptacek 5mo ago"Idle cost is that one lightweight SELECT per millisecond per database — no page-cache pressure, no writer-lock contention, no kernel file watcher in the mix." I think (respectfully) the LLM that probably wrote this overshot the mark here because busy-polling a select does not actually sound better to me than a "kernel file watcher".
- felooboolooomba 5mo ago"one lightweight SELECT per millisecond" This reminds me of the teenager who told her dad that she was just a tiny little bit pregnant.
- giraffe_lady 5mo ago[flagged]
- tptacek 5mo agoYeah, again, to be clear: I get how SQLite works and I'm not dunking on the design, I'm just saying the comparison set up on this page snags. It's a classic LLM negated triptych, but "one of these things is not like the other": cache pressure: bad, writer contention: bad, kernel file watcher: ... good, actually? Intuitively seems better than this design?
- rv64imafdc 5mo agoHold on -- if it really is "one lightweight SELECT per millisecond", and you're saying a select is "a couple hundred microseconds", say generously 200us?, then you're spending 200us out of every 1000us just selecting. That's a lot of polling!
- giraffe_lady 5mo agoI mean only in the same sense that you spend 1 second per second doing something. Time is probably not the best way to evaluate the resources this consumes and I doubt it takes much of anything else either. It does seem weird though even for sqlite. I wonder how oban does it. I also wonder if OP knows oban can run on sqlite.
- sroussey 5mo agoThing of the battery! (read that in the way of "think of the children!")
- nine_k 5mo ago
- vmsp 5mo agoReminds me of Litestack for Rails. Eventually, it was abandoned because Rails itself started going all out on SQLite. https://github.com/oldmoe/litestack https://github.com/oldmoe/litestack
- nop_slide 5mo agoAll in*
- canadiantim 5mo agoCould this work with Turso, the SQLite rust rewrite?
- russellthehippo 5mo agoAuthor here. Yeah doesn’t depend on the underlying db if it speaks SQLite.
- enduku 5mo agoI think this is interesting too sqlite a as the coordination boundary: business state, queue state, stream offsets, retries, and acks all sharing one transactional substrate. The 1ms polling is getting a lot of weight in the thread though :)
- maxdo 5mo agoAlmost feels like someone is trying to joke about similar postgres application . To make it look even more absurd . SQLite is not concurrent and you’ll have tons of problems using it practically .
- andrewstuart 5mo agoI can’t see any benchmarks or performance stats. I’d like to see messages per second.
- deferredgrant 5mo ago[flagged]
- andrewstuart 5mo agoSuggestion for the author wind back the polling to once a second when nothing is happening.
- wmanley 5mo agoI've implemented something similar in the past, but using inotify. You need to watch the -wal file for IN_MODIFY. To make it work reliably I found I had to run: BEGIN IMMEDIATE TRANSACTION; ROLLBACK; Otherwise the new changes weren't guaranteed to be visible to the process. I'm sure there's a more targetted approach that would work instead - maybe flock on a particular byte in the `-shm` file.
- codedokode 5mo ago> Once real work flows through a SQLite-backed app, you need a queue. The usual answer is “add Redis + Celery.” Are they joking? SQLite is usually used for single-process (mutliple threads) applications. The proper way to communicate between threads/processes is a ring buffer, where you allocate structs (allocation typically is incrementing a pointer), and futex/eventfd for notifications (+ some spinlocking to avoid going to kernel when the tasks arrive quickly). Why do you need redis for that? If you need persistent tasks, then you can store them in the table, and still use futex for notifications. This polling is inefficient and they should not make it a library which will cause other lazy developers add it to their app. > honker polls SQLite’s PRAGMA data_version every millisecond. That’s a monotonic counter SQLite increments on every commit from any connection, journal mode, or process — a ~3 µs read for a precise wake signal That's 3 ms per second = 0.3% CPU time wasted for every waiting thread. Like Electron, this feels like written by a web developer and not a real programmer.
- deepsun 5mo agoNevertheless, expect articles like "We replaced our redis cluster with this simple extension and got it N times faster".
- Groxx 5mo ago>That's 3 ms per second = 0.3% CPU time wasted for every waiting thread. I suspect that's actually "per process, per database (usually 1)", and not based on number of threads or tables. `data_version` semantics mean there's no need for more than one connection polling it, and it's being used as a relatively lightweight "DB has changed, check queues" check (that's pretty much its whole purpose). Also I believe this is mostly intended for multi-process use, e.g. out-of-process workers, so an in-process dirty tracker (e.g. just check after insert/update/delete) isn't sufficient. So I do think it's somewhat crazy, but it is at least very simple. fsnotify-like monitoring seems like a fairly obvious improvement tho, not sure why that isn't part of it. Maybe it's slower? I haven't tried to do anything actually-performant-or-reliable with fs notifications, dunno what dragons lie in wait.
- russellthehippo 5mo agoAuthor here - previously posted here: https://news.ycombinator.com/item?id=47874647 https://news.ycombinator.com/item?id=47874647 Key difference vs SQL polling is that we’re touching metadata instead of data pages. I have work in process to make this work without any polling (innotify, kqueue, mmap’d shm file check) after the original stat(2) direction proved unreliable if lightweight. Would love your feedback and or contributions in the repo - still figuring out the end shape.
- deleted 5mo ago[deleted]
- opiniateddev 5mo agoWhy not just use https://github.com/conductor-oss/python-sdk https://github.com/conductor-oss/python-sdk provide durability, distributed and orchestration.
- sexylinux 5mo agoA good reason: you do not want npm AND docker AND java just for your queue.
- kweiza 5mo agoOn edge this misses Durable Objects + alarms — same primitives, no polling, no Redis to skip in the first place.
- ghm2180 5mo agoCan this work with lightstream?
- sparkbyte 5mo ago[flagged]
- neocron 5mo agoNo maven package for java? Guess this isn't a serious project
- Jaden688 5mo ago[dead]
- tengbretson 5mo agoI'm a big fan of SQLite and all that, but if SQLite constrains you to a single writer process, why not do this in your application layer anyway?
- SQLite 5mo agoSQLite allows multiple writers. The constraint is that only one of how writers can be actively writing at any moment in time. If there are multiple processes wanting to write, they take turns. SQLite prevents two or more writes from running concurrently, so there is nothing the application needs to do to implement this, other than responding to SQLITE_BUSY replies from failed (concurrent) write attempts and retrying after a short delay. Why this constraint? Because SQLite is serverless. There is no central server available to coordinate concurrent writes. At the lowest level of the stack, every database engine has this same constraint, as there is only one wire connecting the CPU to the SSD, and you cannot send multiple writes over the same wire at the same time. But in a client/server database, the server (in cooperation with the filesystem) is at hand to serialize the writes and prevent problems in ways that are not possible without a server. The server creates the illusion of concurrent writes by multiplexing the single write wire efficiently and making that multiplexing transparent to the application.
- dimitrismrtzs 5mo ago[flagged]