4 ms·
The author makes some great suggestions, namely to cache heavily and throttle requests. However, they lost a lot of credibility for me with "screen scraper traf
by storborg 15y ago
The author makes some great suggestions, namely to cache heavily and throttle requests. However, they lost a lot of credibility for me with "screen scraper traffic should be indistinguishable from human traffic". Sorry, but that's BS--socially responsible scraping leaves control with the publisher. If the publisher doesn't want you scraping their content, you shouldn't try to fake a human in order to be able to.
- dotBen 15y agoI see your point - however I read from it that the author was more referring to the load/level of activity on the server that your requests make should be indistinguishable from human traffic. IE if the server's log files have 100's of requests from the same IP address in successive lines then that doesn't look like human behavior. What would have been nice for a 'best practice' document would be to show how to set the HTTP AGENT string for the crawler so that it had an identifier, version number and some contact method.
- helwr 15y agoThis was already asked: http://groups.google.com/group/comp.lang.java/browse_thread/thread/6923c024ed392c85/88fa10845061c8ba?pli=1 http://groups.google.com/group/comp.lang.java/browse_thread/...