4 ms·
I did something like this a few weeks ago on my photography site: https://robertmay.photography/journal/meta-has-tried-to-scrape-this-site-1-million-times-in-2-
by robotmay 19d ago
I did something like this a few weeks ago on my photography site: https://robertmay.photography/journal/meta-has-tried-to-scrape-this-site-1-million-times-in-2-weeks-ive-given-them-toasters-instead https://robertmay.photography/journal/meta-has-tried-to-scra...
Meta not only hasn't noticed, but is currently sending about 11 requests per second to my site. I've also seemingly trapped one of those TV proxy scraper nets as I'm getting absolutely hammered by requests from all over the place now. I get maybe 10 legit visitors per day, and I'm currently blocking 406,787 IPs from things that have fallen into my honeypot.
I've tweaked my site to return empty status responses a configurable amount of time but the traffic has been so intense that Traefik is now struggling, so I'm going to have to figure out something else. I was returning over-capacity errors and I think that was a mistake, I've swapped to 400 range status codes now. I don't want to use Cloudflare so I'm not sure what to do after this.
The people at these companies are either incompetent or malicious.
- reaperducer 19d agoSince it's a photography site, route 'em over to goatse. That might get someone's attention.
- robotmay 19d agoHaha I did debate going much worse with the junk images but wanted to err on the side of caution in case I subject possible clients to something like goatse.
- FabCH 19d agoReturn a HTTP 301 pointing to https://facebook.com https://facebook.com? Might make them scan themselves instead.
- nubinetwork 19d agoMost distributed bots won't follow a 301. If 404/400 don't work, just 444 them, they're not worth giving back a response, especially on a personal website.