4 ms·
This will only be a good approach if you are going to scrape a small amount of pages. The problem is using synchronous requests, as this blocks the crawler unti
by niels 14y ago
This will only be a good approach if you are going to scrape a small amount of pages. The problem is using synchronous requests, as this blocks the crawler until a request has finished. Using asynchronous requests such as supported by twisted (and scrapy) will allow you to crawl a lot faster using the same resources.
- boyter 14y agoThis can actually sometimes be a feature. It makes it far less likely to have your IP banned. Its also a far more polite way to crawl someones site.
- niels 14y agoI agree, and for a 101 web scraping tutorial keeping it simple is nice.
- darkarmani 14y agoYou could crawl a lot of different sites one page at a time. When I wrote a large distributed download system, I would use pycurl's bandwidth throttle and also store a 5 minute average of bandwidth per domain that would prevent other downloaders from saturating a domain.
- zevyoura 14y agoI would argue that the proper implementation provides real rate limiting, both in terms of max requests per second and also max concurrent requests. Limiting to one concurrent request is likely to be extremely slow for any significant amount of data, and a couple concurrent requests is not impolite. Obviously I'm not saying you should effectively DoS the site you're scraping, but there's a balance and 1 concurrent request is almost definitely the wrong place to set it.
- gjreda 14y agoI've heard this from countless people who have read the post. It's definitely made me want to look into Scrapy.
- deleted 14y ago[deleted]