5 ms·
I maintain ~30 different crawlers. Most of them are using Scrapy. Some are using PhantomJS/CasperJS but they are called from Scrapy via a simple web service. A
by samtc 9y ago
I maintain ~30 different crawlers. Most of them are using Scrapy. Some are using PhantomJS/CasperJS but they are called from Scrapy via a simple web service.
All data (zip files, pdf, html, xml, json) we collect are stored as-is (/path/to/<dataset name>/<unique key>/<timestamp>) and processed later using a Spark pipeline. lxml.html is WAY faster than beautifulsoup and less prone to exception.
We have cronjob (cron + jenkins) that trigger dataset update and discovery. For example, we scrape corporate registry, so everyday we update the 20k oldest companies version. We also implement "discovery" logic in all of our crawlers so they can find new data (ex.: newly registered company). We use Redis to send task (update / discovery) to our crawlers.
- frik 9y ago> We use Redis to send task (update / discovery) to our crawlers. Some kind of queue implemented with Redis? How does it work?
- CGamesPlay 9y agoProbably not what the GP uses, but Resque does this in Ruby land.
- bdcravens 9y agoSidekiq has emerged as a better option to Resque
- thibaut_barrere 9y agoSee https://sidekiq.org https://sidekiq.org for instance.
- samtc 9y agoIt's a simple redis list containing JSON task. We have a custom Scrapy Spider hooked to next_request and item_scraped [1]. It check (lpop) for update/discovery tasks in the list and build a Request [2]. We only crawl max ~1 request per second, so performance is not an issue. For every website we crawl we implement a custom discovery/update logic. Discovery can be, for example, crawl a specific date range, seq number, postal code.... We usually seed discovery based on the actual data we have, like highest_company_number + 1000, so we get the newly registered companies. Update is to update a single document. Like crawl document for company number 1234. We generate a Request [2] to crawl only that document. [1] https://doc.scrapy.org/en/latest/topics/signals.html https://doc.scrapy.org/en/latest/topics/signals.html [2] https://doc.scrapy.org/en/latest/topics/request-response.html https://doc.scrapy.org/en/latest/topics/request-response.htm...
- CGamesPlay 9y agoI have a similar set up! How do you monitor for failures and deal with the scrape target changing?
- samtc 9y agoWe monitor exceptions with Sentry. We store raw data so we don't have to hurry to fix the ETL, we only have to fix navigation logic and we keep crawling.
- Launchr 9y agoSorry if it's a stupid question/example/comparison, just trying to understand better: You're storing the full html data instead of reaching into the specific div's for the data you might need? This way, separating the fetching from the parsing? I'm a scraping rookie, and I usually fetch + parse in the same call, this might resolve some issues for me :) thanks!
- jimsmart 9y agoWhen I've done scraping, I've always taken this approach also: I decouple my process into paired fetch-to-local-cache-folder and process-cached-files stages. I find this useful for several reasons, but particularly if you want to recrawl the same site for new/updated content, or if you decide to grab extra data from the pages (or, indeed, if your original parsing goes wrong or meets pages it wasn't designed for). Related: As well as any pages I cache, I generally also have each stage output a CSV (requested url, local file name, status, any other relevant data or metadata), which can be used to drive later stages, or may contain the final output data. Requesting all of the pages is the biggest time sink when scraping — it's good to avoid having to do any portion of that again, if possible.
- mapster 9y agoMind if I ask what info/data you are scraping and for what ends?