3 ms·
It's a simple redis list containing JSON task. We have a custom Scrapy Spider hooked to next_request and item_scraped [1]. It check (lpop) for update/discovery
by samtc 9y ago
It's a simple redis list containing JSON task. We have a custom Scrapy Spider hooked to next_request and item_scraped [1]. It check (lpop) for update/discovery tasks in the list and build a Request [2]. We only crawl max ~1 request per second, so performance is not an issue.
For every website we crawl we implement a custom discovery/update logic.
Discovery can be, for example, crawl a specific date range, seq number, postal code.... We usually seed discovery based on the actual data we have, like highest_company_number + 1000, so we get the newly registered companies.
Update is to update a single document. Like crawl document for company number 1234. We generate a Request [2] to crawl only that document.
[1] https://doc.scrapy.org/en/latest/topics/signals.html https://doc.scrapy.org/en/latest/topics/signals.html
[2] https://doc.scrapy.org/en/latest/topics/request-response.html https://doc.scrapy.org/en/latest/topics/request-response.htm...