8 ms·
In summary -- their RabbitMQ consumer library and config is broken in that their consumers are fetching additional messages when they shouldn't. I've never seen
by mark242 3y ago
In summary -- their RabbitMQ consumer library and config is broken in that their consumers are fetching additional messages when they shouldn't. I've never seen this in years of dealing with RabbitMQ. This caused a cascading failure in that consumers were unable to grab messages, rightfully, when only one of the messages was manually ack'ed. Fixing this one fetch issue with their consumer would have fixed the entire problem. Switching to pg probably caused them to rewrite their message fetching code, which probably fixed the underlying issue.
It ultimately doesn't matter because of the low volume they're dealing with, but gang, "just slap a queue on it" gets you the same results as "just slap a cache on it" if you don't understand the tool you're working with. If they knew that some jobs would take hours and some jobs would take seconds, why would you not immediately spin up four queues. Two for the short jobs (one acting as a DLQ), and two for the long jobs (again, one acting as a DLQ). Your DLQ queues have a low TTL, and on expiration those messages get placed back onto the tail of the original queues. Any failure by your consumer, and that message gets dropped onto the DLQ and your overall throughput is determined by the number * velocity of your consumers, and not on your queue architecture.
This pg queue will last a very long time for them. Great! They're willing to give up the easy fanout architecture for simplicity, which again at their volume, sure, that's a valid trade. At higher volumes, they should go back to the drawing board.
- stolsvik 3y agoTheir solution does not preclude fanout. Fetching work items from such a work queue by multiple nodes/servers should be no problem. One solution that also would be good for monitoring would be to have a state column, and a "handled by node" column. So, in a transaction, find a row that is not taken, take it by setting the state to PROCESSING, and set the handled_by_node to the nodename. Shove in a timestamp for good measure. When it is done, set the state to DONE, or delete the row. Monitor by having some health check evaluating that no row stays in PROCESSING for too long, and that no row stays NOT_STARTED for too long, etc. Introspect by making a nice little HTML screen that shows this work queue and its states. As I wrote in another comment, this is somewhat similar to a "work queue pattern" I've described here: https://mats3.io/patterns/work-queues/ https://mats3.io/patterns/work-queues/
- datavirtue 3y agoThis.
- dpflan 3y agoYour last comment is the key, they had an issue and not the scale so a simpler approach works, but then I imagine that this company, which is a new company and growing, will have a future blogpost about switching from pg queue to something that fits their scale...
- SahAssar 3y agoSo they are picking the right tool (and a tool that they know) for their problem.
- hospadar 3y agoAlso when they have two many jobs for their one table - partition the table by customer, when that's still somehow too big - shard the table across a couple DB instances. Toss in some beefy machines that can keep the tables in memory and I suspect you'd have a LOOOONG way to go before you ever really needed to get off of postgres. In my experience, the benefits of a SQL table for a problem like this are real big - easier to see what's in the queue, manipulate the queue, resolve head-of-queue blocking problems, etc.
- dpflan 3y agoThere is another variable of experience with the technology which seems to be high for Postgres, low for RabbitMQ…
- seunosewa 3y agoNot exactly. For performance (I guess) each worker fetches an extra job while it's working on the current job. If the current job happens to be very long, then the extra job it fetched will be stuck waiting for a long time. Your multiple queue solution might work but it is most efficient to have just one queue with a pool of workers where where each worker doesn't pop a job unless it's ready to process it immediately. In my experience, this is the optimal solution.
- mannyv 3y agoIt's actually a misconfiguration (see the comment below with the documentation).
- yarg 3y agoBetter to post a link - the order of comments changes with votes and time.
- quintes 3y agoI used rabbit many many years ago but agree, scale consumers and only pop when ready to actually process
- whakim 3y agoI think this is definitely optimal if your jobs take a long time (and if you add in worker autoscaling of some type, you can deal with quick jobs reasonably well too). But if you have a very large number of quick jobs, setting a high prefetch limit can increase throughput by a tremendous amount. It's in this context that multiple queues makes a lot of sense; you can optimize some workers to read from one queue and others to read from another.
- echelon 3y agoI'm at a point where I built a low volume queue in MySQL and need to rip it out and replace it with something that does 100+ QPS, exactly once dispatch / single worker processing, job priority level, job topics, sampling from the queue without dequeuing, and limited on failure retry. I can probably bolt some of these properties onto a queue that doesn't support all the features I need.
- ryanjshaw 3y ago> I've never seen this in years of dealing with RabbitMQ. Did you do long running jobs like they did? It's a stereotype, but I don't think they used the technology correctly here -- you're not supposed to hold onto messages for hours before acknowledging. They should have used RabbitMQ just to kick off the job, immediately ACKing the request, and job tracking/completion handled inside... a database.
- xahrepap 3y agoI’ve used RabbitMq to do long running jobs. Jobs that take hours and hours to complete. Occasionally even 1-2 days. It did take some configuring to get it working. Between acking appropriately and the prefetch (qos perhaps? Can’t remember, don’t have it in front of me). We were able to make it work. It was pretty straightforward it never even crossed my mind that this isn’t a correct use case for RMQ. (Used the Java client.)
- mark242 3y agoThe short answer is "yes" but the questions that you should be asking are: A) How long am I willing to block the queue for additional consumers, B) How committed am I to getting close to exactly-once processing, and C) how tolerant of consumer failure should I be? Depending on the answer to those three questions is what drives your queue architecture. Note that this has nothing to do with time spent processing messages or "long running jobs". Assume that your producers will be able to spike and generate messages faster than your consumers can process them. This is normal! This is why you have a queue in the first place! If your jobs take 5 seconds or 5 hours, your strategy is influenced by the answers to those three questions. For example -- if you're willing to drop a message if a consumer gets power-cycled, then yeah, you'd immediately ack the request and put it back onto a dead letter queue if your consumer runs into an exception. Alternatively, if you're unwilling to block and you want to be very tolerant of consumer failure, you'd fan out your queues and have your consumers checking multiple queues in parallel. Etc etc etc, you get the drift. Keep in mind also that this isn't specific to RabbitMQ! You'd want to answer the same questions if you were using SQS, or if you were using Kafka, or if you were using 0mq, or if you were using Redis queues, or if you were using pg queues.
- aidos 3y agoIt may be a misconfiguration but I’m fairly sure you couldn’t change this behaviour in the past. Each worker would take a job in advance and you could not prevent it (I might be misremembering but I think I checked the source at the time). In my experience, RabbitMQ isn’t a good fit for long running tasks. This was 10 years ago. But honestly, if you have a short number of long running tasks, Postgres is probably a better fit. You get transactional control and you remove a load of complexity from the system.
- whakim 3y agoI don't think this behavior has changed significantly. The key issue is that they seem to have correctly identified that they wanted to prefetch a single task, but they didn't recognize that this setting is the count of un-ACK'ed tasks. If you ACK upon receipt (as most consumers do by default), then you're really prefetching two tasks - one that's being processed, and one that's waiting to be processed. If you ACK late, you get the behavior that TFA seems to want. I've seen this misconfiguration a number of times.
- skrtskrt 3y ago> RabbitMQ isn’t a good fit for long running tasks yeah I've seen 3 different workplaces run into this exact issue, usually when they started off with a default Django-Celery-Redis approach all of those cases were actually easily fixed with Postgres SELECT FOR UPDATE as a job queue
- phamilton 3y agoI'll add another to the anecdata. We saw this issue with RabbitMQ. We replaced it with SQS at the time but we're currently rebuilding it all on SELECT FOR UPDATE. Our problem was that when a consumer hung on a poison pilled message, the prefetched messages would not be released. We fixed the hanging, but hit a similar issue, and then we fixed that, etc. We moved to SQS for other reasons (the primary being that we sometimes saturated a single erlang process per rabbit queue), but moving to the SQS visibility timeout model has in general been easier to reason about and has been a better operations experience. However, we've found that all the jobs are in postgres anyway, and being able to index into our job queue and remove jobs is really useful. We started storing job metadata (including "don't process this job") in postgres and checking it at the start of all our queue workers and we've decided that our lives would be simpler if it was all in postgres. It's still an experiment on our part, but we've seen a lot of strong stories around it and think it's worth trying out.
- ftkftk 3y agoMy thoughts exactly half way through the article.
- tracker1 3y agoYeah, my first thought was curiosity about their volume needs. DB based queues are fine if you don't need more than a few messages a second of transport. For that matter, I've found Azure's Storage Queues probably the simplest and most reliable easy button for queues that don't need a lot of complexity. Once you need more than that, it gets... complicated. Also, sharing queues for multiple types of jobs just feels like frustration waiting to happen.
- TexanFeller 3y ago> DB based queues are fine if you don't need more than a few messages a second of transport I'd estimate more like dozens to hundreds per second should be pretty doable, depending on payload, even on a small DB. More if you can logically partition the events. Have implemented such a queue and haven't gotten close to bottlenecking on it.
- tracker1 3y agoHad meant few hundred... :-)
- eastern 3y agoA column in the database identifying which 'queue' is all you need for that. The table is not a queue, it's a structure to store queues.
- whalesalad 3y agoSounds like a prefetch issue. Or auto ack. Rabbit is a phenomenal tool but you need to know how to use it.
- Xenoamorphous 3y agoPretty sure they just needed to disable autoack and manually ack after the long running task was done.
- btilly 3y agoIf they knew that some jobs would take hours and some jobs would take seconds, why would you not immediately spin up four queues. Two for the short jobs (one acting as a DLQ), and two for the long jobs (again, one acting as a DLQ). Your DLQ queues have a low TTL, and on expiration those messages get placed back onto the tail of the original queues. Here is why I would not recommend that. Do that and you have to rewrite your system around predictions how long each job will take, deal with 4 sources of failure, and have more complicated code. All this to maintain the complication of an ultimately unneeded queueing system. You call it an "easy fanout architecture for simplicity." I call it, "an entirely unnecessary complication that they have no business having to deal with." If they get to a volume where they should go back to the drawing board, then they can worry about it then. This would be a good problem to have. But there is no need to worry about it now.
- pushedx_ 3y agoAgreed. I previously worked for a well known service provider that relied on a single SQL database for queueing hundreds of jobs per second across all regions. The system was well understood by everyone at the company and the company had a surprisingly high market cap. You can get very far with the simple thing.
- noisy_boy 3y agoWith the additional advantage that you can run adhoc queries against the SQL database which isn't an option with RMQ.
- mekoka 3y ago> In summary -- their RabbitMQ consumer library and config is broken in that their consumers are fetching additional messages when they shouldn't. I've never seen this in years of dealing with RabbitMQ. What do you mean by "broken"? Are you implying that the behavior they're describing is not the way the consumer library is supposed to work? They linked to RabbitMQ's documentation basically saying that's exactly how it works. Also, where do you get the sense that they've misconfigured it? You've made these statements, but did not exactly enlightened us as to how one should set things up to have consumers handle exactly one job at a time. That was their only problem (Edit: an answer by @whakim https://news.ycombinator.com/item?id=35530108 https://news.ycombinator.com/item?id=35530108 is providing more light on this). The rest of your answer sanctimoniously presumes that they don't know how to use the tool, but your own proposed solution is moot, as it seems to address a different problem, not the one that they have (1 job per consumer max).
- tuyiown 3y agoThe architectural error was to have messages acknowledged at end of long processing, probably to avoid to handle messages produced from the worker, instead of having two messages with quick ack, one message at start, one message at end from worker. RabbitMQ jobs is to handle message transfert, if you tie business logics (jobs completion) to its state, you have a problem. So basically, producer and consumer MUST have their own storage for tracking processes, typically by ID, RabbitMQ jobs being the message handler for synchin' states.
- whakim 3y agoI wish you wouldn't speak so authoritatively because much of this is just one way (and hardly the only way) to implement such a system. If you read the RabbitMQ docs (see https://www.rabbitmq.com/confirms.html https://www.rabbitmq.com/confirms.html) you'll see that ACK'ing the messages after processing is explicitly described as a way to handle consumer failures.
- tuyiown 3y agoDepends on your definition of processing. In the case I describe the processing is receiving and recording the job id sent by the producer. Ack should be sent asap so that all non-business issues on message transmissions should be handled by messaging logics. Having the jobs processing itself being independent of the messaging solution seems like good way to go as the actual message implementation, might have to change. I really think that you don't want ossification of your messaging implementation detail on your business logics. I did not pretend that it was the only solution, merely just something I know would be reliable, an implicit «one way to it».
- mastermedo 3y agoWhat about messages that uncover a bug in the consumer/message format? After X failed attempts, the data shouldn’t be enqueued anymore, but rather logged somewhere for engineers to inspect.
- raverbashing 3y agoIt might be because RabbitMQ libraries and tooling is awfully unintuitive and confusing (without some higher level library on top). At some point using what you understand is easier
- rawoke083600 3y agoThis! In all my years of working with RabbitMQ is been super solid! Throwing all that out in exchange for a db queue, honestly doesnt seem smart (with respect)