4 ms·
The crawling aspect always seems overlooked to me. It's really easy to get a single page and pull the required data from it. However, what strategies do you u
by theworst 12y ago
The crawling aspect always seems overlooked to me. It's really easy to get a single page and pull the required data from it. However, what strategies do you use to crawl? How do you get around IP blocks? Continuous crawling? Etc.
In the end, it's the infrastructure that powers the extraction that requires all my attention. I've got a bunch of techniques I use, but I'd love to compare with how other people do it.
Love that more scraping resources are coming online. I see scraping as the important link between the web as it is now, and the web as it will be in 20 years. (web2 -> web3 for the jargon geeks.) The whole semantic web isn't going to be useful for most non-academics without considerable structuring effort put into existing data.