3 ms·
There are some tasks, such as asynchronously-generated web content, where a message queue is the obvious solution. But there are a class of problems where eith
by mattrepl 16y ago
There are some tasks, such as asynchronously-generated web content, where a message queue is the obvious solution. But there are a class of problems where either a message queue or batch processing system (e.g., Hadoop) could be used. Consider the summation of user mention counts from the Twitter streaming API for an example; either a job is put on a queue for each tweet containing a mention, or each of the mentioning tweets are thrown into a bucket that is used as the source for a batch job that recurrently executes.
From what's mentioned in the article, it sounds like some of the tasks that Beanstalk is being used for at PostRank are the same type of tasks that other companies, such as FlightCaster, are doing with Hadoop/Cascading.
The trade-off seems to be that message queues are more flexible and can offer lower latency of job completion but batch processing systems provide better support for admin concerns like adding worker nodes, debugging, and reporting.
- igrigorik 16y ago"From what's mentioned in the article, it sounds like some of the tasks that Beanstalk is being used for at PostRank are the same type of tasks that other companies, such as FlightCaster, are doing with Hadoop/Cascading." Hmm, not at all. A message queue and a job queue are not necessarily one and the same. What we need is real-time scheduling, with up to the second resolution for each job. We don't run a batch "go fetch all of these pages" jobs, rather our system is always running, always fetching content. For that, a heap/work queue is required.
- mattrepl 16y agoI've enjoyed reading your posts over the years, thanks for them. You're correct that a message queue isn't necessarily a job queue, I was referring to job queues.