10 ms·
Avoiding Webscraping Throttling Using Python and Tor as a Proxy
- paulryanrogers 7y agoNot sure how long this will last since there are a limited number of exit nodes. Many networks throw up CAPTCHA for such nodes.
- wybiral 7y agoAh, yes, the captcha arms race. I've encountered bots that can solve simple text captchas (especially bots routed through Tor) and as ML gets cheaper and more accurate I do wonder what the end result will be...
- greglindahl 7y agoWell, for Tor in particular, the most likely end result is Tor being banned from an increasing number of websites.
- mirimir 7y agoThat already happened.
- Topgamer7 7y agoYeah honestly surprised scraping through tor would work. I imagine most commercial solutions routinely block exit nodes.
- viraptor 7y agoDepending on how you read the embargo laws for your country, you may be required to block them. If you identify the source of the traffic as a non-country, dropping it may be what you want to do.
- deleted 7y ago[deleted]
- CWuestefeld 7y agoWhile this is intellectually interesting, I'm troubled by the fact that the author seems not to have given the slightest thought that he's breaking the site's T&C, or of how much this abuse costs the service. I'm particularly sensitive to this because I'm constantly dealing with scraper bots from competitors that are trying to monitor our pricing. Without our ongoing policing, and a fair amount of developer time going into it, the traffic coming from these bots - and hence the amount it costs us to operate the site - is significantly larger than that of actual customers. Let me say that again: scraper bots account for more traffic on our sites than do legitimate customers.
- vbezhenar 7y agoAnother problem is that blocking tor nodes is very easy. And with abusing tor nodes as free proxies, it makes a huge disservice to the Tor network, because more websites will ban it and ordinary users won't be able to use it.
- beardog 7y agoThis problem is being partly addressed with Privacy Pass. The addon lets you bypass recaptcha including cloudflare recaptcha minus one initial solve per session. Of course this doesn't account for all blocks and of course some malicious bots bypass recaptcha anyway using AI/speech recognition or human labor farms, but it will make the ecosystem better for everyone. (note: it is not a captcha solving service, its a collaborative project by cloudflare, google and others) https://privacypass.github.io/ https://privacypass.github.io/
- bored_hacker 7y agoI'm actually not so sure that there terms of services prohibits scraping for personal use (or scraping at all) but maybe my understanding is incorrect. I have read through them here https://www.zillow.com/corp/Terms.htm https://www.zillow.com/corp/Terms.htm and looked through there robots.txt here https://streeteasy.com/robots.txt https://streeteasy.com/robots.txt. This is most likely because they use Distil Networks and rely on them to block web scrapers. But you are right this is something of concern and I can add a section talking about the dangers of breaking a sites terms and conditions and how scraping must be done correctly even for personal use
- mirimir 7y agoIt's overkill to use Tor for this. And I'm a little surprised that it works at all well, because Tor exits so commonly trigger CAPTCHAs. Better, I think, would be using HTTPS proxies. But not free ones, which tend to get burned down pretty quickly. There are sites that lease private proxies, and guarantee that they work.
- judge2020 7y agoHTTP/S proxies probably wouldn't want you burning their IPs on web scraping though - it might be against the license/lease terms of the site.
- mirimir 7y agoOK, top DDG result (an ad) for "https proxies web scraping": > https://smartproxy.com/scraping/proxies https://smartproxy.com/scraping/proxies > Global Residential Proxy Network for Web Scraping. Never Get Blocked Again. High Reliability, Easy to Use. Choose Your Plan Now!
- 333c 7y agoWhat are the odds that this is essentially a botnet made of infected computers and IoT devices with unsuspecting owners?
- goatsi 7y agoInfecting computers and IoT devices takes too much effort these days. The latest method is to have the users brought to you - just publish a SDK and pay developers to stick it into their apps. Here's an example: https://luminati.io/sdk https://luminati.io/sdk
- FDSGSG 7y agoLuminati has not killed off the other providers in the market.
- 7y ago
- lammalamma25 7y agoCool article but couldn't you just spoof your IP using something like socket?
- lammalamma25 7y agoReplying to myself here. This will not work because the packet still has to go back to the spoofed address. What a poorly thought out comment.
- foobar_ 7y agoIf only websites have their data dumps for free instead of html reverse engineering.
- Thorrez 7y ago<span class="pull-right" id="ipv4">2a0b:f4c1::7</span> How is that an IPv4?
- unnouinceput 7y agoxxx.xxx.xxx.xxx and the mask. in his example: 2a = 42; 0b = 11; f4 = 244; c1 = 193; mask = 7 (weird i know) so the IP is actually 42.11.244.193. A short trip to any free IP locator site says that IP is from South Korea
- nurettin 7y agoThis method will not work for many websites that require you to stay logged in.
- deleted 7y ago[deleted]
- nomilk 7y agoCould scraping through TOR be risky if site has different content depending on the visitor's IP? Or does TOR allow you to control for that (e.g. get only IPs from a specific location)?
- jjjbokma 7y agoYou can use ExitNodes in your torrc to set a country code, e.g. {us}.
- 256cats 7y agoWell, Tor is usually blocked. If you want something more reliable, use private proxy services, for example https://gimmeproxy.com https://gimmeproxy.com
- bydl0coder 7y agoIt's higly likely that a web site will greet any connections from Tor with captcha.