4 ms·
Hey, author here! Feel free to ask any questions you have.
by stummjr 11y ago
Hey, author here! Feel free to ask any questions you have.
- piroux 11y agoHere is a first one : What are the best ways to detect changes in html sources with scrapy, thus giving missing data in automatic systems that need to be fed ?
- deleted 11y ago[deleted]
- stummjr 11y agoHey, not sure if I understood what you mean. Did you mean: 1) detect pages that had changed since the last crawl, to avoid recrawling pages that hadn't changed? 2) detect pages that have changed their structure, breaking down the Spider that crawl it.
- grantbachman 11y agoAs someone who does a fair amount of scraping at his job, I'd like to hear what you have to say regarding both questions :)
- mitchtbaum 11y ago> 1) detect pages that had changed since the last crawl, to avoid recrawling pages that hadn't changed? Usually web clients use https://en.wikipedia.org/wiki/HTTP_ETag https://en.wikipedia.org/wiki/HTTP_ETag , afais. If a web app\server lacks that skill, then you could compute your own hash and check it yourself, instead of processing that condition at the network layer.
- stummjr 11y ago1) detect pages that had changed since the last crawl, to avoid recrawling pages that hadn't changed? You could use the deltafetch[1] middleware. It ignores requests to pages with items extracted in previous crawls. 2) detect pages that have changed their structure, breaking down the Spider that crawl it. This is a tough one, since most of the spiders are heavily based on the HTML structure. You could use Spidermon [2] to monitor your spiders. It's available as an addon in the Scrapy Cloud platform [3], and there are plans to open source it in the near future. Also, dealing automatically with pages that change their structure is in the roadmap for Portia [4]. [1] https://github.com/scrapinghub/scrapylib/blob/master/scrapylib/deltafetch.py https://github.com/scrapinghub/scrapylib/blob/master/scrapyl... [2] http://doc.scrapinghub.com/addons.html?highlight=monitoring#monitoring http://doc.scrapinghub.com/addons.html?highlight=monitoring#... [3] http://scrapinghub.com/scrapy-cloud/ http://scrapinghub.com/scrapy-cloud/ [4] http://scrapinghub.com/portia/ http://scrapinghub.com/portia/
- eliasdorneles 11y agoWell, missing data can happen from problems in several different levels: 1) site changes caused the items that were scraped to be incomplete (missing fields) -- for this, one approach is to use an Item Validation Pipeline in Scrapy, perhaps using a JSON schema or something similar, logging errors or rejecting an item if it doesn't pass the validation. 2) site changes caused the scraping the items itself to fail: one solution is to store the sources and monitor the spider errors -- and when there are errors, you can rescrape from the stored sources (it can get a bit expensive store sources for big crawlers). Scrapy doesn't have a complete solution for this out-of-the-box, you have to build your own. You could use the HTTP cache mechanism and build a custom cache policy: http://doc.scrapy.org/en/latest/topics/downloader-middleware.html#scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware http://doc.scrapy.org/en/latest/topics/downloader-middleware... 3) site changed the navigation structure, and the pages to be scraped from were never reached: this is the worst one, it's similar to the previous one, but it's one that you want to detect earlier -- saving the sources doesn't help much, since it happens at an early time during the crawl, so you want to be monitoring it. One good practice is to split the crawl in two: one spider does the navigation and push the links of the pages to be scraped into a queue or something, and another spider reads the URLs from that and just scrape the data.
- mryan 11y agoGreat article, thanks for sharing. What would you recommend for people who need to operate a cluster of scrapyd instances? Are there any recommended ways of managing the distribution of tasks to multiple scrapyd instances? I have looked at scrapyd-cluster, but I would prefer not to add a Zookeeper cluster to my stack. Currently I'm thinking of modifying scrapyd-cluster so that it uses AMQP (via Celery) to handle task distribution/retries/etc. I appreciate this conflicts with the scrapinghub business model, so any tips you can offer would be greatly appreciated :-)
- asibiryakov 11y agoHi, mryan! I'm the core Frontera developer. The precise answer heavily depends on your use case (what "task" is? scalability requirements, data flow), so please ask your question in Frontera google groups, and we will try to address it directly. First you could try making use of Frontera, here are the different distribution models it provides out of the box http://frontera.readthedocs.org/en/latest/topics/run-modes.html http://frontera.readthedocs.org/en/latest/topics/run-modes.h.... Frontera is web crawling framework made in Scrapinghub, providing crawl frontier and scaling/distribution capabilities. Along with flexible queue and partitioning design, you will get also document metadata storage (HBase or RDBMS of your choice) with simple revisiting mechanism. Second, we have a simple redis-based solution for scaling spiders https://github.com/rolando/scrapy-redis https://github.com/rolando/scrapy-redis. It's only dependency is Redis, so it's easy to quick start, but it has only one queue shared between spiders, hard-coded partitioning, and Redis limiting scalability. You mentioned a scrapy-cluster (not scrapyd-cluster probably). It provides a more sophisticated distribution model, allowing you to separate crawls with jobs concept within the same service, maintains separate per-spider queues (allowing to crawl politely, I hope you plan to do so? :), and forcing you to use it's prioritization model. Also it allows to use spiders of different types sharing the same Redis instance, and prioritize requests on cluster level. BTW, I haven't found any dependencies on Zookeeper. None of the solutions provides provisioning out of the box. If spider was killed by OOM, or consume too much resources (open file descriptors, memory) you have to take care of it by yourself. You could use supervisord, or upstart or some custom process management solution. It all depends on you monitoring requirements. Good luck choosing the right solution! A.