4 ms·
I once used VPNs as proxies for web scraping, using 2-3 VPN providers with multiple servers in different countries. It was a while ago, so my memory is a little
by spangry 5y ago
I once used VPNs as proxies for web scraping, using 2-3 VPN providers with multiple servers in different countries. It was a while ago, so my memory is a little hazy but I'll try to describe how I did it as best I remember.
I configured an Alpine Linux docker container with openvpn and a proxy server (I think I settled on squid for stability), and a bash script to to start up the openvpn connection and proxy server with config for both passed into the container. Then just generated a long, line by line list of every possible vpn connection config line by line, shuffled and duplicated.
Then in my outer scraping function: grab a line from the config file, start up a vpn-proxy container (passing in the config), do one page download through the container proxy and then stop and delete the docker container. This allowed me to download pages in parallel, with all connections originating from different IP addresses (as long as I made sure not to exceed VPN simultaneous connection limits).
Kinda messy, and I spent ages fiddling to get the container config just right, but it worked.
- wraptile 5y agoThat's a very cool hack! Thanks for sharing. This does seem like a very big overhead just for one request: build up/tear down would be quite expensive. I actually noticed that there are personal VPNs that do not have a concurrent devices limit. I guess your hack could be modified to startup persistent images for every VPN server and have them run forever as proxy servers! I'll tinker around with this more but this would definitely make it easier for beginners to onboard on proxy based scraping as many people have VPNs ready for netflix and such already.