4 ms·
great site , what im missing on the web is information on how to build or use proxy server for use with the wev scrapers
by umen 5y ago
great site , what im missing on the web is information
on how to build or use proxy server for use with the wev scrapers
- wraptile 5y agoThanks, that's a great suggestion! I actually have an article on proxies in the works however I'm a bit stuck on making it a bit more accessible since proxy access is either really expensive (for a single dev) or extremely unstable (like free proxies). I've always had the luxury of paid quality proxies in my web-scrapers however for article purposes I'd like to have an example of cheap or free proxies for casual usage/education. Are you using something in particular? One idea I've explored was using VPN service as a proxy but seems like most big VPN providers are not providing proxies anymore. Anyway, I can PM you once I figure out how to put this together properly! :)
- spangry 5y agoI once used VPNs as proxies for web scraping, using 2-3 VPN providers with multiple servers in different countries. It was a while ago, so my memory is a little hazy but I'll try to describe how I did it as best I remember. I configured an Alpine Linux docker container with openvpn and a proxy server (I think I settled on squid for stability), and a bash script to to start up the openvpn connection and proxy server with config for both passed into the container. Then just generated a long, line by line list of every possible vpn connection config line by line, shuffled and duplicated. Then in my outer scraping function: grab a line from the config file, start up a vpn-proxy container (passing in the config), do one page download through the container proxy and then stop and delete the docker container. This allowed me to download pages in parallel, with all connections originating from different IP addresses (as long as I made sure not to exceed VPN simultaneous connection limits). Kinda messy, and I spent ages fiddling to get the container config just right, but it worked.
- wraptile 5y agoThat's a very cool hack! Thanks for sharing. This does seem like a very big overhead just for one request: build up/tear down would be quite expensive. I actually noticed that there are personal VPNs that do not have a concurrent devices limit. I guess your hack could be modified to startup persistent images for every VPN server and have them run forever as proxy servers! I'll tinker around with this more but this would definitely make it easier for beginners to onboard on proxy based scraping as many people have VPNs ready for netflix and such already.