7 ms·
Making Postgres queues scale
- dewey 2mo ago> The conventional wisdom around Postgres-backed queues is that they don't scale. That might have been the case 10 years ago. In the past years there have been many Postgres powered queueing systems and even Rails switched to Postgres powered queues by default (SolidQueue) more than 3 years ago.
- mjfisher 2mo agoI see a lot of back and forth about postgres' suitability as a queuing system. I wonder if there's a couple of separable problems here. Postgres backed queues - even very scalable ones - might work well for background jobs in a monolithic app backed by a single DB. Things like backgrounding sending an email etc. But usually when I reach for a queuing system, it's because I want to decouple a part of the architecture. And in that case, it's probably better to use a dedicated queueing system instead of postgres. I wonder if the two cases are conflated in a lot of online discussion.
- dewey 2mo agoWe are very heavily using Postgres as a queuing system in production for many years already, not just some quiet background tasks but with millions of tasks in the queue at any given moment. I have yet so see any issues with that so I'm always a bit suspicious when people say they had to reach for something else unless you are at a crazy scale. I know it's very hard to compare workloads, but famous recent example: https://openai.com/index/scaling-postgresql/ https://openai.com/index/scaling-postgresql/ > It may sound surprising that a single-primary architecture can meet the demands of OpenAI’s scale; however, making this work in practice isn’t simple.
- tomnipotent 2mo agoThat post is about how Postgres doesn't scale for write-heavy workloads and that they had to move those workloads to Cosmos DB. For the rest of the remaining mostly-read workload they have a single primary with 50 read replicas.
- dewey 2mo agoThe point is that Postgres scales a very long way. Once you arrive at OpenAI / ChatGPT scale there's no shame in reaching for a dedicated queuing system.
- tomnipotent 2mo ago> scales a very long way This ignores the fine print. It scales a very long way under specific circumstances with specific workloads. The more write-heavy your workload the less eloquently Postgres scales. For OpenAI's use case you could swap Postgres with MySQL and it would scale just as well.
- ballon_monkey 2mo agoIf your DB is write heavy, tune it for writes...
- tomnipotent 2mo agoYou can't tune your way out of this constraint. The heap table and MVCC implementation put a hard ceiling on what can be accomplished without reaching for another tool.
- dewey 2mo agoWe are talking about Postgres as a queue here, which is heavy on tiny writes, reads and churning tables. The point is just that a RDMS like Postgres scales very far as a queue, it’s not a fight if Postgres or MySQL. Nobody is arguing for using it for everything and forever but for most company sizes it’s perfectly fine to not reach for a dedicated queuing tool if you already have PG running.
- bel8 2mo agoI have the impression that SolidQueue uses whatever your rails app is using by default. That can be SQLite, PostgreSQL, MySQL, etc. So I don't think they changed to PostgreSQL per se.
- dewey 2mo agoCorrect, what I meant was that SolidQueue was the first database-backed default adapter over the usually common Redis + Sidekiq combo for ActiveJob.
- sorentwo 2mo agoThey certainly do, and I don't think it's a controversial take at this point. Shameless link to an older article about throughput with Oban (https://oban.pro/articles/one-million-jobs-a-minute-with-oban https://oban.pro/articles/one-million-jobs-a-minute-with-oba...), and in follow-up research we've sustained 12k/s with a p99 under ~100ms.
- sebmellen 2mo agoVery interesting, I’d never heard of Oban before, and that’s after a fair amount of research into background job queueing systems. Thanks for posting!
- richwater 2mo agoReading Postgres queuing posts always seem like deja vu. People love to write about them
- mlnj 2mo agoThese days, most of them are from the DBOS folks. :)
- tonyhb 2mo agoindeed, every ~week with essentially the same thing: skip locked.
- keeganpoppen 2mo agoeveryone sure acts like postgres is some sort of thing that we discovered next to the antikythera mechanism on the sunken isle of atlantis. the "just get one big machine to do all of the stuff" scaling strategy remains undefeated, though. even if it is presented like it is anything but.
- 0x457 2mo agoAnd every time it is "FOR UPDATE SKIP LOCKED" or something else super obvious that should have been a starting point when you pick PG for your queue.
- everfrustrated 2mo ago[flagged]
- atombender 2mo agoA performance pitfall that isn't addressed in the DBOS article at all is the bloat problem: If you update or delete rows that you consume, dead tuples start to accumulate due to Postgres' way of doing MVCC. This is a serious problem because it affects the planner's ability to make good choices. Dead tuples are still indexed and the need to skip them isn't accounted for by the query planner, so a table with lots of dead tuples may perform really badly. The autovacuum process will be constantly chasing dead tuples, and you'll want to set the autovacuum settings to be very aggressive to be able to keep up. My team has been starting to use PgQue [1] for a new application, and it seems really well-designed. PgQue is explicitly designed to solve the bloat problem, by avoiding tuple deletion. Instead of deleting processed tuples, it will periodically TRUNCATE the entire table. It uses two tables that it "flips" between so TRUNCATE can run on the inactive table while the active on is used for queuing. PgQue also uses a snapshot approach to avoid row-level locks. PgQue also stands out in that its queue model is position-based, so it can implement nice features like collaborative consumers, fan-out, atomic batches, and "recover from last good" behaviour. It makes some compromises (no priority support, slightly higher latency), but they're fine for most use cases. Previously discussed on HN here [2]. [1] https://github.com/NikolayS/pgque https://github.com/NikolayS/pgque [2] https://news.ycombinator.com/item?id=47817349 https://news.ycombinator.com/item?id=47817349
- sgt 2mo agoNot a problem unless your system has busy queue processing 24/7 though, which I bet is pretty rare for most companies.
- atombender 2mo agoNot true, unfortunately. The dead tuple build-up can happen in a very short amount of time. I speak from having had to deal with this in a production environment that used the SKIP LOCKED method used in the article.
- sgt 2mo agoWhat were the volumes involved, how many jobs per second and so on?
- rcleveng 2mo agoAre these posted manually or is there a scheduled claude co-work task that does it?
- rtpg 2mo agowe had a whole discussion at work around whether to not to use pg for queues at a reasonable size (in particular for doing some notion of fair queueing distribution across tenants). I ended up finding a good number of HN comments like "we were doing this and regretting it". So here's my ask: anybody here use PG for queues at a system with reasonable throughput, without regretting it? Like where there might be some contention
- hmaxdml 2mo agoWhat's reasonable? DBOS has users running queues at millions of tasks per hour. For fair queuing you can have partitioned queues where only active partitions consume resources
- mattlong 2mo agoCan you say more on how you support fair queueing with partitioned queues? Very top of mind for me right now!
- shakow 2mo agoRunning it at a couple dozens jobs per second, and we're happy with it now that we polished the cutting edges (virtually the same thing as in TFA).
- hiyer 2mo agoIn a recent interview I was asked to design a job queue and I went with postgres with the first and third optimizations mentioned here. For the scale of the question - 1000 concurrent jobs - I argued that postgres would easily scale. But the interviewer - maybe because they were from aws - felt it wouldn't and wanted me to go with sqs instead.
- LtdJorge 2mo agoIsn’t "for no key update skip locked" better than "for update skip locked"? Most of the times there will be no improvement, but it’s a good option when you don’t need a stronger lock (e.g., for DELETE).
- NightMKoder 2mo agoYou can go deeper and model a queue as a ring-esque buffer with a write head (can just be a serial id) and a read head. The read head starts at the same place as the write head and advances only up to the write head and no further via nextval(). The main benefit is you now remove the lock contention as many workers attempt to dequeue at once. The super advanced version of this is pgque - https://pgque.dev/ https://pgque.dev/ - but that’s more like Kafka in Postgres. I wouldn’t go there if you don’t know the Kafka model already and you want it.