5 ms·
I've heard some great things about Celery, and I think it may be better at handling massive amounts of tasks (10s of thousands) without using SQS like Norc. Ho
by darrellsilver 17y ago
I've heard some great things about Celery, and I think it may be better at handling massive amounts of tasks (10s of thousands) without using SQS like Norc.
How does it handle logs, resources and changing trees?
Gonna have to look into it more!
- timf 17y agoInstead of SQS it uses RabbitMQ. I don't know about it's limits but since a serious RabbitMQ installation can handle something like millions of messages per second I assume the limits you will hit are in the persistence solution on celeryd. It can use a Tokyo Tyrant or MongoDB backend instead of a RDBMS, as well as memcached support. Those seem like they would help in that department. I don't think you can change trees after they are sent (you mean subtasks, right?). You can see logging example at "defining and executing tasks" section here: http://ask.github.com/celery/introduction.html#usage http://ask.github.com/celery/introduction.html#usage Sorry, I am only a light user at this point in time.
- darrellsilver 17y agoI've been hearing really good things about RabbitMQ, and haven't been particularly impressed with SQS. It works at scale (we're currently at 10s of thousands a day, but it's probably the same performance at 100 or 1000x that...) but tasks take many seconds (like 10-30) to show up in the queue, which absolutely kills us. We'll be adding a RabbitMQ plugin to Norc, just like SQS, when we can.
- vegai 17y agohttp://www.rabbitmq.com/faq.html#performance http://www.rabbitmq.com/faq.html#performance "From our testing, we expect easily-achievable throughputs of 4000 persistent, non-transacted one-kilobyte messages per second (Intel Pentium D, 2.8GHz, dual core, gigabit ethernet) from a single RabbitMQ broker node writing to a single spindle." Or what do you mean by serious?
- timf 17y agoI was thinking more about multiple brokers working together. Also, hmm, I didn't think rabbit was so significantly behind other AMQP implementations like zeromq: "4,100,000 messages a second" - http://www.zeromq.org/ http://www.zeromq.org/
- asksol 17y agoJust a small note: You can also use CELERY_BACKEND="amqp" to send back the result as a message, it's the most efficient way, but then you can only look up the result once (unless you send another message with the same result).
- asksol 17y agoNot sure what you mean by logs, resources and changing trees. But there is logging support (the python logging module). Resources: There is AMQP QoS which makes sure it only receives as many tasks as it can handle. Task hard and soft time-limits is coming in 1.0 (patch ready). If the soft timeout is exceeded an exception is raised which the task can catch to do any clean up before the hard time limit is exceeded and the task is forcefully killed. Rate limit (per task type or global) using the token bucket algorithm (which allows for bursts of data). For 1.0 (patch ready and tested) Otherwise you have to OS process resource limits (cpu/memory etc). Monitoring is coming in 1.0 as well, someone is working on a monitoring system with a web-frontend where you can see the current state of the system (support for deleting already published tasks might be added, but then on an opt-in basis) The current scheduling system is flawed (it uses the database, which is a dead end in my opinion), a new solution is almost ready which uses a separate centralized service that works like a clock sending out messages at schedule time: http://wiki.github.com/ask/celery/rewriting-the-periodic-task-service http://wiki.github.com/ask/celery/rewriting-the-periodic-tas... Now for changing trees, I'm not sure what you mean here, please correct me if I misunderstood. Messages can not be changed once they have been published, so the task itself is responsible for changing the execution order. You can chain tasks, so say TaskA launches another task. You can retry tasks if they fail. Oh and there's the message routing features made available by AMQP, which means you can have different servers/instances handle different tasks. Celery has a lot of features, and even more is under development, so I don't think I can list them all here. I can't see anything hindering a Celery implementation of Norc, but as I read it you started working on this before celery started. Bad luck when we could have shared a lot of work :(
- darrellsilver 17y agoThat sounds super interesting! This sentence raises a question: >You can chain tasks, so say TaskA launches another task. You can retry tasks if they fail. How does this chaining work? Does each Task define its children? I've found that its cleaner to separate tasks from their place in the tree. For example, a script that downloads a CSV file each hour shouldn't care how that file is used.
- 17y ago