3 ms·
I do web crawling for a living, the method mentioned in the article does not work for most sites
by ospider 8y ago
I do web crawling for a living, the method mentioned in the article does not work for most sites
- alehul 8y agoI've been working on a project that requires scraping from a large number of sites. Do you have any recommendations on better resources?
- Bromskloss 8y agoDo most sites have scraping detection at all? Are they even opposed to scraping?
- codefined 8y agoOn a website I'd written we had pseudo-randomly generated URLs to show dynamic content (it was a game, the URL contained parameters). On each page we had this little widget that included five random configurations people might like to try. A few times our website went down due to the load going >30, eventually I discovered Google was doing something funky, adding the dynamic domains to the "robot.txt" file fixed the issue. Then some other search engines / scrapers seemed to run into the same issue and started requesting hundreds of thousands of URLs per day (these pages were dynamically generated and took a moderate amount of compute power). We eventually did have to implement basic anti-scraper rules because it was degrading the user experience.
- kapauldo 8y agoWhat does work?