5 ms·
Update: After a week of doing nothing - they finally noticed their thing is blocked and sprang into action. Apparently, they expanded their pool of available I
by santah 6y ago
Update: After a week of doing nothing - they finally noticed their thing is blocked and sprang into action.
Apparently, they expanded their pool of available IPs they pull data from and now they seem to be endless (so some of the scraping domains actually work now).
I'm investigating what I can do about it. I'd appreciate any advice!
- speedgoose 6y agoDo they use a web browser for the scraping or simply a http library? You could look at the http request headers and perhaps identify the scrapper script. You could also put a javascript challenge that is required to solve before pulling more data, and disable it for Google and Bing ips, so it's more work for them to pull data for some time. Instead of simply blocking, you could detect them and do some kind of http slowloris response.
- santah 6y agoThis is a good question I don't have an answer to. I'll try and find out and also I'll have to learn exactly what "slowloris" is. It may be helpful indeed!
- speedgoose 6y agoslowloris is an old attack on HTTP. The idea is to send garbage HTTP headers as slow as possible while keeping the TCP connection open, you send one letter every 5 seconds for example. The HTTP stack on the other side stays busy waiting for HTTP headers and can't do anything else meanwhile. It's usually targeted to webservers, before most HTTP servers got fixed you could DDOS a server with a tiny connection, but some HTTP clients can also be vulnerable. But you may find better usage of your time than implementing this.
- gnyman 6y agoI had not heard about slowloris, thanks for the tip. I am always on the lookout for ways to make life harder for scrappers and scanners. In this case though, to defeat scrapers maybe create some link which only the scraper sees and leave a "gzip-bomb"' like described here https://blog.haschek.at/2017/how-to-defend-your-website-with-zip-bombs.html https://blog.haschek.at/2017/how-to-defend-your-website-with... and see how their scraper handle that :-) Personally I just used a html-fuzzer to generate 5 MiB of junk html and named it wp-login.php :-) And a ssh-tarpit
- markhowe 6y agoSetup a honeypot page to log the ‘users’ IP. Keep hitting it via their domain and you’ll build up a list of IP’s to block? As an aside, I’ve fought credential stuffers by returning real looking but actually false data, and initiating password resets... start serving different data on each hit, you may need to be annoying enough that they give up.
- santah 6y agoA honeypot is exactly how I caught the IPs the first time around. Problem is - right now I'm over 250 (new) IPs and they keep piling up (their domains now rarely use an IP more than once). I may have to block entire ranges of IPs or whole ASNs.
- cmeacham98 6y agoHow about automatically honeypotting them? Add some code to your site that will IP ban a user that searches for some random string (and when I say random, I mean literally generate a random string - something no legit user would search for). Then, setup a script on your laptop or whatever to search this string on their domains every half hour or so.
- santah 6y agoIt's basically what I've done, though have not automated it yet. It even prepares the expression snippet for me to paste directly into a CloudFlare firewall rule. That's how I got to quickly identify and ban almost 2000 different IPs. If they continue to expand the IP pool I may need to automate it though.
- eps 6y agoThey must've seen this post.
- santah 6y agoUpdate 2: After banning close to 2k individual IPs, it looks like I got it under control, for now. I wonder how banning so many IPs affects CloudFlare performance and if I should optimize it to block whole IP ranges instead ...
- aclelland 6y agoAre the IPs from the same ARN block? Cloudflare should show you that in the Firewall events page. If they're using IPs from a single VPN/Server company then you might be able to just block the ASN. You might also be able to find common user agent headers through the CF firewall page and block them based on their UA. That'd not work if their scraper tool was randomizing the UA string but quite a few of them don't
- santah 6y agoIncluding all IPs I blocked today, they're spread between 5 different ASNs. I may resort to blocking them eventually, but for now - individually blocking the IPs (even in the thousands as it is) - seems to be working well enough. As for user agent - they're using a very common, real browser user agent that's impossible to distinct from legit users.
- Gys 6y agoDo you have many users in Russia? You could block the whole country ;-) Russia has no GDPR or something. So you could put (special key in) a cookie? They probably do not process it so subsequent requests without a cookie are to be discarded?
- lewiscollard 6y ago> Russia has no GDPR or something. So you could put (special key in) a cookie? It is entirely permissible under the GDPR to use cookies for security purposes.
- eythian 6y ago