9 ms·
Show HN: An SQS Alternative on Postgres
- conroy 2y agoI'm curious how this performs compared to River https://riverqueue.com/ https://riverqueue.com/ https://news.ycombinator.com/item?id=38349716 https://news.ycombinator.com/item?id=38349716
- chuckhend 2y agoI think it would be tough to compare. There are client libraries for several languages, but the project is mostly a SQL API to the queue operations like send, read, archive, delete using the same semantics as SQS/RSMQ. Any language that can connect to Postgres can use PGMQ, whereas it seems River is Go only?
- thangngoc89 2y agoI’m wondering if there are language agnostic queues where the queue consumers and publishers could be written in different languages?
- rco8786 2y agoThat's exactly what this is. You write your own consumer/publisher code however you want, and interact with the queues via SQL queries.
- SideburnsOfDoom 2y agoThat's normal, yes. Name a queuing system, and with very few exceptions it will have clients for a variety of languages. It is also normal to exchange messages with contant as json, protobuf or similar format, which again can be processed by any language that aims to be widely used. In fact, are there any queues that aren't language-agnostic? I had the idea that ZeroMQ was a C/C++ only thing, but I checked the docs and it's got the usual plethora of language bindings https://zeromq.org/get-started/ https://zeromq.org/get-started/ So right now I can't name any queue systems that are single language. They're all aimed at interop.
- thangngoc89 2y agoInteresting. I've been looking at a much more simpler system than Celery queue to publish jobs from Golang and consume jobs from Python side (AI/ML stuff). This threads gave a lot of names and links for further investigation.
- anamexis 2y agoAs mentioned, there are a plethora of queuing systems that have cross platform clients. If you’re interested specifically in a background job system, you may want to check out Faktory. It’s Mike Perham’s language-agnostic follow-up to Sidekiq. https://github.com/contribsys/faktory https://github.com/contribsys/faktory
- thangngoc89 2y agoThanks for the recommendation. I've checked this out and looks like a good alternative with much simpler configuration and less foot-gun than Celery.
- justinclift 2y ago> Name a queuing system, and with very few exceptions it will have clients for a variety of languages. RQ seems to be one of those exceptions, as even after many years it looks to be Python only. :( https://python-rq.org https://python-rq.org
- thangngoc89 2y agoRQ uses Python's Pickle and it doesn't have any serialization protocol so it stays Python only.
- justinclift 2y agoYeah, it's a right pain. :(
- SideburnsOfDoom 2y ago
- bastawhiz 2y ago> Guaranteed "exactly once" delivery of messages to a consumer within a visibility timeout That's not going to be true. It might be true when things are running well, but when it fails, it'll either be at most once or at least once. You don't build for the steady state, you build against the failure mode. That's an important deciding factor in whether you choose a system: you can accept duplicates gracefully or you can accept some amount of data loss. Without reviewing all of the code, it's not possible to say what this actually is, but since it seems like it's up to the implementor to set up replication, I suspect this is an at-most-once queue (if the client receives a response before the server has replicated the data and the server is destroyed, the data is lost). But depending on the diligence of the developer, it could be that this provides no real guarantees (0-N deliveries).
- MuffinFlavored 2y ago> That's not going to be true. It might be true when things are running well, but when it fails, it'll either be at most once or at least once. Silly question as somebody not very deep in the details on this. It's not easy to make distributed systems idempotent across the board (POST vs PUT, etc.) Distributed rollbacks are also hard once you reach interacting with 3rd party APIs, databases, cache, etc. What is the trick in your "on message received handler" from the queue to achieve "exactly once"? Some kind of "message hash ID" and then you check in Redis if it has already been processed either fully successfully, partially, or with failures? That has drawbacks/problems too, no? Is it an impossible problem?
- samtheprogram 2y agoYou don’t achieve exactly once at the protocol/queue level, but at the service/consumer level. This is why SQS guarantees at-least-once. It’s generally what you described, but the term I’ve seen is “nonce” which is essentially an ID on the message that’s unique, and you can check against an in-memory data store or similar to see if you’ve already processed the message, and simply return / stop processing the message/job if so.
- 2y ago
- pyuser583 2y agoWhat advantages does this have over RabbitMQ? My experience is Postgres queuing makes sense you must extract or persist in the same Postgres instance. Otherwise, there’s no advantage over standard MQ systems. Is there something I don’t know.
- vbezhenar 2y agoOne big advantage of using queue inside DB is that you can actually use queue operations in the same transaction as your data operations. It makes everything incredibly simpler when it comes to failure modes. IMO 90% of software which uses external queues is buggy when it comes to edge cases.
- zbentley 2y agoA fair point. If you do need an external queue for any reason (legacy/already have one, advanced routing semantics, integrations for external stream processors, etc.) the "Transactional Outbox" pattern provides a way to have your cake and eat it too here--but only for produce operations. In this pattern, publishers write to an RDBMS table on publish, and then best-effort publish to the message broker after RDBMS transaction commit, deleting the row on publish success (optionally doing this in a failure-swallowing/background-threaded way). An external scheduled job polls the "publishes" table and republishes any rows that failed to make it to the message broker later on. When coupled with inbound message deduplication (a feature many message brokers now support to some degree) and/or consumer idempotency, this is a pretty robust way to reduce the reliability hit of an external message broker being in your transaction processing path. It's not a panacea, in that it doesn't help with transactional processing/consumption and imposes some extra DB load, but is fairly easy to adopt in an ad-hoc/don't-have-to-rewrite-the-whole-app way. https://microservices.io/patterns/data/transactional-outbox.html https://microservices.io/patterns/data/transactional-outbox....
- ilkhan4 2y agoSure, not having to spin up a separate server is nice, but this aspect is underappreciated, imo. We eliminated a whole class of errors and edge cases at my day job just by switching the event enqueue to the same DB transaction as the things that triggered them. It does create a bottleneck at the DB, but as others have commented, you probably aren't going to need the scalability as much as you think you do.
- seveibar 2y agoPeople considering this project should also probably consider Graphile Worker[1] I've scaled Graphile Worker to 10m daily jobs just fine The behavior of this library is a bit different and in some ways a bit lower level. If you are using something like this, expect to get very intimate with it as you scale- a lot of times your custom workload would really benefit from a custom index and it's handy to understand how the underlying system works. [1] https://worker.graphile.org/ https://worker.graphile.org/
- philefstat 2y agohave also used/introduced this to several places I've worked and it's been great each time. My only qualm is it's not particularly easy to modify the exponential back off timing without hacky solutions. Have you ever found a good way to do that?
- valenterry 2y agoIs there something like that in the jvm world?
- RedShift1 2y agoYou don't need anything specific. SELECT ... FROM queue FOR UPDATE SKIP LOCKED is the secret sauce.
- valenterry 2y agoWould be nice to get some goodies for free, like overview, pausing, state, statistics etc. :-)
- chuckhend 2y agoThis is exactly how pgmq is implemented, + the usage of VT.
- wmfiv 2y agohttps://www.jobrunr.io//en/ https://www.jobrunr.io//en/ seems to be popular at the moment.
- rco8786 2y agoThis is neat. Would be cool if there was support for a dead letter or retry queue. The idea of deleting an event transactionally with the result of processing said event is pretty nice.
- jpambrun 2y agoRetries are baked in as message re-becomes visible after a configurable time. For dead letter you can move to another queue on the nth retry.
- chuckhend 2y agoThis is exactly how we do this in our SaaS at Tembo.io. We check read_ct, and move the message if >= N. I think it would be awesome if this were a built-in feature though.
- deleted 2y ago[deleted]
- ltbarcly3 2y agoAnother one of these! It's interesting how many times this has been made and abandoned and made again. https://wiki.postgresql.org/wiki/PGQ_Tutorial https://wiki.postgresql.org/wiki/PGQ_Tutorial https://github.com/florentx/pgqueue https://github.com/florentx/pgqueue https://github.com/cirello-io/pgqueue https://github.com/cirello-io/pgqueue Hundreds of them! We have a home grown one called PGQ at work here also. It's a good idea and easy to implement, but still valuable to have implemented already. Cool project.
- mattbillenstein 2y agoHa, I wrote one too - https://github.com/mattbillenstein/pg-queue/blob/main/pg-queue.sql https://github.com/mattbillenstein/pg-queue/blob/main/pg-que... Roughly follows the semantics of beanstalkd which I used once upon a time and quite liked.
- arecurrence 2y agoYeah, I've written a few of these and should probably release a package at some point but each version has been somewhat domain specific. The last time we measured an immediate 99% performance improvement over SNS+SQS. It was so dramatic that we were able to reduce job resources simply due to the queue implementation change. There's a lot of useful and almost trivial features you can throw in as well. SQS hasn't changed much in a long time.
- deleted 2y ago[deleted]
- dostoevsky013 2y agoI’m not sure what are the benefits for the micro service architecture. Do you expect other services/domains to connect to your database to listen for events? How does it scale if you have several micro services that need to publish events? Or do you expect to a dedicated database to be maintained for this queue? Worth comparing it with other queue systems that persist messages and can help you to scale message processing like kafka with topic partitions. Found this article on how Revolut uses Postgres for events processing: https://medium.com/revolut/recording-more-events-but-where-will-we-store-them-4b1dad457cf5 https://medium.com/revolut/recording-more-events-but-where-w...
- chuckhend 2y agoWe talk a little bit in https://tembo.io/blog/managed-postgres-rust https://tembo.io/blog/managed-postgres-rust about how we use PGMQ to run our SaaS at Tembo.io. We could have ran a Redis instance and used RSMQ, but it simplified our architecture to stick with Postgres rather than bringing in Redis. As for scaling - normal Postgres scaling rules apply. max_connections will determine how many concurrent applications can connect. The queue workload (many insert, read, update, delete) is very OLTP-like IMO, and Postgres handles that very well. We wrote some about dealing with bloat in this blog: https://tembo.io/blog/optimizing-postgres-auto-vacuum https://tembo.io/blog/optimizing-postgres-auto-vacuum
- solatic 2y agoWho said anything about microservice architecture? If you're building a new MVP that needs background processing, you have a queue, and your monolith listens to the queue. Sticking to one database during the MVP keeps things simple until you validate both the product and expected scale. You can migrate to something heavier later if the product gets traction.
- airocker 2y agoI think it will be better if you create events automatically based on commit events in wal.
- bdcravens 2y agoNote there are a number of background job processors for specific languages/frameworks that use Postgresql as the broker. For example GoodJob and the upcoming SolidQueue in Ruby and Rails.
- cynicalsecurity 2y agoBut why? Why not have a proper SQS service? What's the obsession with Postgres?
- poisonborz 2y agono dependence on a third party
- chuckhend 2y agoIMO, it is most valuable when you are looking for ways of reducing complexity. For a lot of projects, if you're already running Postgres then it is maybe not worth the added complexity of bringing in another technology.
- cryptonector 2y agoSee https://news.ycombinator.com/item?id=40307454#40311843 https://news.ycombinator.com/item?id=40307454#40311843
- jpambrun 2y agoWhy use a service that comes with lock-in and poor developer experience when I can use the database I already have?
- redact207 2y agoI agree in general, but there will always be certain requirements and team structures where stuff like this makes sense. For me, I work in a small team of 6 devs on an ever growing app and feature set. I 100% will leverage managed services where cost and complexity allow. SQS is one of the most stable and cheapest AWS service, and the ability to just use it and not have to sysops it means we can spend more time building features.
- mlhpdx 2y agoIndeed. I’ve relied heavily on SQS for years and never regretted it. I question the comparison to SQS for this add on — it’s not really in the same ballpark.
- 2y ago
- RedShift1 2y agoThis seems like a lot of fluff for basically SELECT ... FROM queue FOR UPDATE SKIP LOCKED? Why is is the extension needed when all it does is run some management type SQL?
- acaloiar 2y agoA similar argument can be made of many primitives and their corresponding higher order applications. Why build higher order concept Y when it's simply built on primitive X? Why build the C programming language when C compilers simply generate assembly or machine code? --- A sensible answer is that new abstractions make lower level primitives easier to manage.
- selcuka 2y agoSQS API compatibility?
- andrewstuart 2y agoSeems an unusual choice that this does not have an HTTP interface. HTTP is really the perfect client agnostic super simple way to interface with a message queue.
- solatic 2y agoOn the contrary, creating new HTTP connections introduces an irreducible source of latency compared to establishing and reusing a persistent connection. You may end up building a single-tenant architecture where each tenant gets its own database and there are relatively few consumers that are able to respond quicker due to sticking with a long-lived connection model.
- andrewstuart 2y agoI don’t understand your point. These are long polled connections.
- solatic 2y agoIf you look at other HTTP-based databases, like DynamoDB or S3, the latency involved in setting up new connections is a downside of those databases (not arguing that it's never worth it, architectural decisions are all trade-offs, but that is a trade-off).
- andrewstuart 2y agoHttp doesn’t set up a new connection each time it supports persistent connections. And in the case of message queue long polling the http connection stays open and waits for a message to become available.
- solatic 2y agoTrying to build on persistent HTTP connections that don't close is a recipe for frustration when scaling horizontally, which is something you plan to do, since the reason why not to go with Postgres is so you can have more than one instance, right? You can't juggle the connection between different servers. If a server drops out (because it's being scaled in), then you lose the connection, and you have to establish a new one, which introduces a hiccup/latency. No free lunches.
- deleted 2y ago[deleted]
- ComputerGuru 2y agoThis is strictly polling, no push or long poll support?
- chuckhend 2y agoThere is a long poll, https://tembo-io.github.io/pgmq/api/sql/functions/#read_with_poll https://tembo-io.github.io/pgmq/api/sql/functions/#read_with... We have been talking about a push using Postgres 'notify', or even via an http, but we don't have a solid design for it yet.
- jilles 2y agoCan someone tell me what the usefulness of this is compared to RabbitMQ or Kafka?
- justinclift 2y agoAs a data point, there's a similar Go based project called Neoq: https://github.com/acaloiaro/neoq https://github.com/acaloiaro/neoq
- deepsun 2y ago> TIMESTAMP WITH TIME ZONE I'm yet to find a use case for "WITH TIME ZONE", in all cases it's better to use "WITHOUT TIME ZONE". All it does is displays the date in sql client local timezone, which should never matter for well done service. Would be glad to learn otherwise.
- boromisp 2y agoTimestamp with time zone is the type for an "absolute timestamp". Timestamp without time zone is for local time, and sometimes abused as "time in utc without being explicit about it". The naming describes the expected input, not what is stored. The time zone name or offset is not stored with timestamptz. Always use timestamptz, unless you have a specific use case for local time.
- globular-toast 2y agoYes, it is quite confusing and I dread to think how many have got it wrong and store local times like the GP. But it's also not as simple as "always use WITH TIME ZONE". That also leads to a mistake. The reason is (just to reiterate what you said) the TIMESTAMP WITH TIME ZONE does not store the time zone! If you ever want to get local time back (e.g. ask a question like "how many users log on before lunch time") then you need to store either local time in a TIMESTAMP WITHOUT TIME ZONE field, or the time zone, and get local time like: SELECT recorded_at AT TIME ZONE time_zone AS local_time ... (I prefer the latter).
- ddorian43 2y agoActually, it's the complete opposite. You always want WITH TIMEZONE.
- globular-toast 2y agoLook carefully. The SQL client does not just "display it in local time", it displays it with a UTC offset. You can be sure whenever you see a UTC offset that UTC is fully recoverable. In this way the TIMESTAMP WITH TIME ZONE field is context independent. It's just UTC. Conversely, if there is no offset, time zone, or something to distinguish it as UTC (like the Z in ISO8601) then you are just storing "local time", that is the time on the clock in someone's kitchen, somewhere. This is the TIMESTAMP WITHOUT TIME ZONE field and is highly context dependent (in particular, what clock was used?)
- kelnos 2y agoPerhaps we were all just not good at database'ing, but at a previous job, "using RDBMS as a queue" became a meme/shorthand for "terrible idea that needs to be stamped out immediately". Does Postgres have some features that make it not entirely unsuitable to use for queuing?
- ttymck 2y agoFor update skip locked
- dewey 2y agoI think this is one of the cases where "You don't have Google problems, so you don't need Google solutions" applies. Using Postgres or a RDBMS for queuing is perfectly fine and it'll get you a long way before you have to worry about scaling or optimizing it. The benefits are easy to see: You already know how to operate a database, you can easily see what's the in the queue, you can easily insert items in the queue with a simple "insert" query, you can use triggers to enqueue items that got changed etc. In the end, a queue is relatively simple to switch out later.
- anentropic 2y agoOne of the appeals of doing MQ in Postgres is being able to submit events atomically in same db transaction as the stuff that raised the event Looking at https://github.com/tembo-io/pgmq/tree/main/tembo-pgmq-python https://github.com/tembo-io/pgmq/tree/main/tembo-pgmq-python ...how do I integrate the queue ops with my other db access code? Or is it better not to use the client lib in that scenario and use the SQL functions directly? https://github.com/tembo-io/pgmq?tab=readme-ov-file#send-two-messages https://github.com/tembo-io/pgmq?tab=readme-ov-file#send-two...
- chuckhend 2y agoThe client libs are a nice convenience, but most users write the sql directly when integrating with other SQL statements, something like: begin; select * from my table group by... select pgmq.send(); commit;
- lloydatkinson 2y agoFor the longest time the common advice was that using a database as a message queue/broker was a bad idea. Now, everyone seems keen to use a DB for this instead of tools dedicated to this purpose. Why?
- sethammons 2y agoThe circle of life. We used to track jobs in a db, but that would eventually hit performance issues (locks, noisy neighbors, etc). So more nosql solutions and services showed up. New devs gobbled up the concepts and APIs but were frustrated that distributed nosql solutions were SaaSified and wanted to bring control back in house, in a single db, and we are back to square one. But with better tooling and hardware in theory. Hopefully you have a dedicated db at least to avoid noisy neighbor workloads but you may still have task queue work that hits scaling limits on the single db causing tasks to interfere with one another--so you can shard the db and isolate task streams, and the circle continues
- pulkitsh1234 2y agoA client is supposed to poll the queue for new items (i.e. issue pop requests in a loop), or is there some better event-oriented approach for this (via pg notify ?)
- biehl 2y agoI've thought about this too. But I can't even tell what would be the good default. At low load events seem nicer, but at high load polling seems necessary?
- chuckhend 2y agopop() or read() in a loop, yes. can read 1 message or many messages at a time. what we do at Tembo in our infrastructure is pause for up to a few seconds if a read() returns no messages. when there are messages, then we read() with no pause in between. when the queues are empty it amounts to less than one query per second. there is not much cost to reading frequently if you use a client side connection pool, or a server side pool like pgbouncer.
- junail 2y agosounds interesting, is there a timeout? or batch processing?
- chuckhend 2y agoyou can send or read a single message at a time or as many as you want in a batch. https://github.com/tembo-io/pgmq?tab=readme-ov-file#read-messages https://github.com/tembo-io/pgmq?tab=readme-ov-file#read-mes...