8 ms·
I love ScrapingHub (and use them) but these tips go completely against my own experience. Whenever I've tried to extract data like that inside Spiders I would
by kami8845 11y ago
I love ScrapingHub (and use them) but these tips go completely against my own experience.
Whenever I've tried to extract data like that inside Spiders I would invariably (and 50,000 URLs later) come to the realization that my .parse() ing code did not cover some weird edge case on the scraped resource and that all data extracted was now basically untrustworthy and worthless. How to re-run all that with more robust logic? Re-start from URL #1
The only solution I've found is to completely de-couple scraping from parsing. parse() captures the url, the response body, request and response headers and then runs with the loot.
Once you've secured it though, these libraries look great.
PS: If you haven't used ScrapingHub you definitely should give it a try, they let you use their awesome & finely-tuned infrastructure completely for free. One of my first spiders ran for 180,000 pages and 50,000 items extracted for $0.
- stummjr 11y agoWhat do you mean by "data like that"? Metadata in Microdata format? Btw, nice to hear your own experience here. :)
- kami8845 11y agoMicrodata is indeed awesome but it's not always there. One website had location tags for me to extract from 97% of pages. Except for a few where the <meta> tags were just missing. Wound up using AlchemyAPI's entity extraction to try and get the location out of the text that way. Thanks for the blogpost! I actually did not know about any of the 3 libraries and am gonna start using them.
- ddebernardy 11y agoWe actually do what you describe as well sometimes. In particular when we scrape sites with robust bot counter-measures to save on Crawlera [1] usage, or on crawls that take long enough that there's a genuine possibility that the site might change before you're done. [1]: http://crawlera.com http://crawlera.com
- JoshTriplett 11y agoCrawling sites actively hostile to your crawler seems like it has both useful and shady applications. What kinds of things do you use that for?
- ddebernardy 11y agoOn the one hand side there's no shortage of users who want to crawl popular sites to monitor e.g. search engine ranking or prices. Which is kind of shady in some sense, or not - when there's no API there's no other way... On the other there are also areas of the web where crawlers are simply not welcome. For instance, DARPA uses a number of our technologies to monitor the dark web for criminal activities: http://opencatalog.darpa.mil/MEMEX.html http://opencatalog.darpa.mil/MEMEX.html
- shostack 11y agoAs an "early" programmer playing with web scraping with the Nokogiri gem, I've been wondering about this aspect (although haven't encountered it yet). Are there legal implications to scraping a site that actively tries to prevent bots from scraping it? I mean, if the data is publicly accessible on the web, could they go after you? I don't plan on doing this for any malicious reasons or anything, and like I said, I haven't encountered it yet. Just having the "what if" thought of what my legal risks might be if I'm playing around with this and whether a site could come after me.
- ddebernardy 11y ago> Are there legal implications to scraping a site that actively tries to prevent bots from scraping it? I mean, if the data is publicly accessible on the web, could they go after you? When we do projects, the baseline is if Google can see it we can too. So from a legal standpoint if Google is covered so are we. From a legal standpoint firms do go after web scrapers. And lose more often than not. The exception is when you're logged in when you crawl. In that case you've implicitly accepted the terms of use. Some companies aggressively sue when you're logged in while scraping, so it's best to stay on the safe side. Further reading on the topic: https://www.quora.com/What-is-the-legality-of-web-scraping https://www.quora.com/What-is-the-legality-of-web-scraping
- dante9999 11y ago> code did not cover some weird edge case on the scraped resource and that all data extracted was now basically untrustworthy and worthless. Your data should not be worthless just because you dont catch some edge cases early. Sure there are always some edge cases but best way to handle them is to have proper validation logic in scrapy pipelines - if something is missing some required fields for example or you get some invalid values (e.g. prices as sequence of characters without digits) you should detect that immediately and not after 50k urls. Rule of thumb is: "never trust data from internet" and always validate it carefully. If you have validation and encounter edge cases you will be sure that they are actual weird outliers that you can either choose to ignore or somehow try to force into your model of content.
- kami8845 11y agoHmm, I'll have to investigate that, any tips for libraries to use for validation that tie well into scrapy? What do you do if you discover that your parsing logic needs to be changed after you've scraped a few thousand items? Re-run your spiders on the URLs that raised errors?
- stummjr 11y agoSpider Contracts can help you: http://doc.scrapy.org/en/latest/topics/contracts.html http://doc.scrapy.org/en/latest/topics/contracts.html
- phunge 11y agoPreach! :) That's my preferred design too. The first job of any external data capture process is to capture the full fidelity source data. Everything else belongs in a followon job.
- rcfox 11y agoThe very first thing I do with every scraping project is enable the HttpCacheMiddleware[0]. After downloading a page once, subsequent runs will automatically pull it from the local cache. This makes it way faster to experiment, and doesn't increase the burden of the website. [0] http://doc.scrapy.org/en/latest/topics/downloader-middleware.html#scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware http://doc.scrapy.org/en/latest/topics/downloader-middleware...
- binarysolo 11y agoSeconding the practice of collect, then parse. Storage is cheap relative to hammering a server (or working around the bounds of rate-limitation), so always save your collected things, then crawl it as many time as you desire as you play around with the right scrapes.