4 ms·
Celery's pretty nice, I've used it in several projects with rabbitMQ and it's useful for IO bound workloads. The main drawback with it is the overhead of dispat
by denom 10y ago
Celery's pretty nice, I've used it in several projects with rabbitMQ and it's useful for IO bound workloads. The main drawback with it is the overhead of dispatching the job and waiting for the result. With threads, there's no overhead of starting up a separate process. That said, I'd recommend celery to anyone where latency is not a problem.
- kerkeslager 10y agoCould you explain what you mean by "The main drawback with it is the overhead of dispatching the job and waiting for the result."?
- denom 10y agoWell there are a couple drawbacks. The author is describing celery as a solution to the GIL thread lock problem. However the time it takes to dispatch the job[1] and fire up the worker is much greater than the time it takes to create a thread. So while celery is _a_ solution to the problem, it's not an equivalent solution. Just something to be aware of. The type of computation you're modeling is limited by the async/dispatched nature of celery. Threads can communicate data structures to each other. With celery, that's awkward with the latency and async semantics. [1] I've used rabbitmq to dispatch jobs over to celery in the past, here's a good overview of latency in that process. TLDR requests that hit the wire are on the order of milliseconds: https://www.rabbitmq.com/blog/2012/05/11/some-queuing-theory-throughput-latency-and-bandwidth/ https://www.rabbitmq.com/blog/2012/05/11/some-queuing-theory...
- kerkeslager 10y agoI guess I'm just not sure what people are expecting here. The point of Celery isn't to parallelize and synchronize, it's to fire off a task, probably on an entirely different machine, and forget about it. If you're trying to sync up afterward, you're going to run into problems because it's not designed for that.
- crucialfelix 10y agoThe overhead is messaging to rabbit, having that create ids for it and queue it; then celery pulls the next job from the exchange, unpacks the job description with its arguments, and then calls the function (wrapped in exception handlers etc. etc.) Celery is suitable for "send an email" or "rebuild a value and put in memcache" but not suitable for "when the OS finally returns the file handle I requested, could you read the file into memory please" Totally different time scales and use-cases.
- kerkeslager 10y agoI guess what I'm confused about is the "waiting for the result" in the parent comment. Celery is pretty much entirely fire-and-forget, and you're usually firing to an entirely different server. If this is being used to achieve parallelization instead of asynchronous processing, then it's no wonder that the parent is having trouble, because they're using Celery for a task for which it's wholly inappropriate.
- crucialfelix 10y agoAgreed. There are a lot of bells and whistles with celery and I've tried them and now avoid them all. I don't nest or chain tasks. I don't use periodic tasks because they can pile up badly and block everything. I have my own simple loop that calls the most urgent periodic-task. I built an angular button that would call a task and would show spinning in-progress and then show when the task completes. Works on development, didn't work on production. why why why ? well, no time to fiddle with celery so I moved on. over the years I've come to mistrust its complexity.
- kerkeslager 10y ago> I don't nest or chain tasks. Yeah, I think nested/chained tasks are pointless. The point of celery is to offload the task to an asychronous worker (probably on another machine) so if you're already on the worker there's no point sending another task to the same worker--it doesn't achieve anything above doing the task inline. > I don't use periodic tasks because they can pile up badly and block everything. This sounds like you're either running the periodic task too frequently, or you need more workers. I've seen a few cases where people tried running a 20-minute task every second on a single worker--of course this is going to pile up. Periodic tasks are among the most predictable, so they're easiest to prevent from piling up: simply make sure that the following relation holds true: w > t / p w = number of workers t = how long it takes for the task to run p = the period (a task that runs ever 20 minutes has period 20 minutes) The > should be greater than by an order of magnitude to be safe in case t increases.
- brianwawok 10y agoI think the idea is for many things you don't sit around and wait for Celery to complete. User clicks "send email"? You throw a send email task in Celery. If it takes .1 seconds or 12 seconds, the user doesn't have to wait for it, because they don't care. This is perfect Celery use. Now when you are loading a page to display to the user, and that page makes a bunch of celery tasks that you NEED To complete.. this is more iffy.. in most cases, you may just want to run those inline in the webrequest. Needing a bunch of tasks to complete to show a user a result is not an ideal use case of Celery.