4 ms·
> We use Redis to send task (update / discovery) to our crawlers. Some kind of queue implemented with Redis? How does it work?
by frik 9y ago
> We use Redis to send task (update / discovery) to our crawlers.
Some kind of queue implemented with Redis? How does it work?
- CGamesPlay 9y agoProbably not what the GP uses, but Resque does this in Ruby land.
- bdcravens 9y agoSidekiq has emerged as a better option to Resque
- thibaut_barrere 9y agoSee https://sidekiq.org https://sidekiq.org for instance.
- samtc 9y agoIt's a simple redis list containing JSON task. We have a custom Scrapy Spider hooked to next_request and item_scraped [1]. It check (lpop) for update/discovery tasks in the list and build a Request [2]. We only crawl max ~1 request per second, so performance is not an issue. For every website we crawl we implement a custom discovery/update logic. Discovery can be, for example, crawl a specific date range, seq number, postal code.... We usually seed discovery based on the actual data we have, like highest_company_number + 1000, so we get the newly registered companies. Update is to update a single document. Like crawl document for company number 1234. We generate a Request [2] to crawl only that document. [1] https://doc.scrapy.org/en/latest/topics/signals.html https://doc.scrapy.org/en/latest/topics/signals.html [2] https://doc.scrapy.org/en/latest/topics/request-response.html https://doc.scrapy.org/en/latest/topics/request-response.htm...