8 ms·
I actually cover this (very briefly) here: http://jakeaustwick.me/python-web-scraping-resource/#thesiteisshowingdifferentcontenttomyscraper http://jakeaustwick.
by Jake232 12y ago
I actually cover this (very briefly) here: http://jakeaustwick.me/python-web-scraping-resource/#thesiteisshowingdifferentcontenttomyscraper http://jakeaustwick.me/python-web-scraping-resource/#thesite...
I should dedicate a section to it though, will stick that on my to-do list.
- tomarr 12y agoI've always opted for Ghost.py rather than selenium as I've found it uses less memory and pretty capable. Admittedly my scrapes are normally pretty targeted and not over 10k+ sites. One query I had from your piece: "When I've found myself in the unfortunate place of getting my proxies banned before on certain sites, they have been more the happy to switch them out for new IP's for me." If you're getting banned from sites, is it not time to leave them alone? If they don't want your traffic (which the admin/system has judged as too much), should you really be circumventing it? With that and the robots.txt bit (which admittedly you justified), you've got to be careful not to slip into a bit of a grey area with scraping, which people regard suspiciously in the first place.