3 ms·
#OpenStreetMap hammered by scrapers hiding behind residential proxy/embedded-SDK networks.
by molly_radstowe 8mo ago
#OpenStreetMap hammered by scrapers hiding behind residential proxy/embedded-SDK networks.
- direwolf20 8mo agoMore like hammered by Google and Apple so you'll use their apps instead.
- wiredpancake 8mo ago[dead]
- petre 8mo agoUnlikely. The data is freely available for download from geofabrik and other sources.
- direwolf20 8mo agoThe data is, the app isn't. OSM provides a giant data dump, not a way to view maps
- Bender 8mo agoLooks like it is hosted in Equinix in NL? Or just part of it maybe? Is it behind a load balancer, maybe something like HAProxy? If so were stick tables set up to limit rates by cookie and require people be logged in on unique accounts and limit anonymous access after so many requests? I know limiting anonymous access is not great but that is something that could be enabled when under a high load so that instead of the site going offline for everyone it would just be limited for the anonymous users. Degradation vs critical outage On a separate note have tcpdump captures been done on these excessive connections? Minus the IP, what do their SYN packets look like? Minus the IP what do the corresponding log entries look like in the web server? Are they using HTTP/1.1 or HTTP/2.0? Are they missing any expected headers for a real person such as cors, no-cors, navigate, accept_language? tcpdump -p --dont-verify-checksums -i any -NNnnvvv -B32768 -c32 -s0 port 443 and 'tcp[13] == 2' Is there someone at OpenStreetMap that can answer these questions?
- KomoD 8mo agoI think it could be worth trying to block them with TLS fingerprinting, or since they think it's residential proxies they are being hammered by, https://spur.us https://spur.us could be worth a try.
- Bender 8mo agoMy personal preference is to first make a small amount of effort finding something unique to the bots that can more often than not be dropped with a simple firewall rule or load balancer ACL. The botters almost always miss something.
- Firefishy 8mo agoDisclosure: I am part of the mostly volunteer run OpenStreetMap ops team. Technically we able to block and restrict the scrapers after the initial request from an IP. We've seen 400,000 IPs in the last 24 hours. Each IP only does a few requests. Most are not very good at faking browsers, but they are getting better. (HTTP/1.1 vs HTTP/2, obviously faked headers etc) The problem has been going on for over a year now. It isn't going away. We need journalists and others to help us push back.
- Bender 8mo agoI hear ya. This is just my opinion but I don't think journalists are going to be much help. The bots would have to be hurting something belonging to the government or the government is paying for to really get them on it. e.g. some big orgs in the government embed your maps on their site. They would have to create legislation and then someone would have to trace the bots back to their operator for attribution and then someone would have to file lawsuits against them once it is illegal. Or you could try using a ToS/AuP to go after them assuming attribution. I am not a lawyer. I think your only hope would be to either find subtle differences between them and real legit users or change how your site works so that bots have to be authenticated unless they have a whitelisted IP/CIDR or put your site behind something else that spots the bots. Beyond that all anyone can do is beef up their infrastructure to handle much more than the bots could dish out. Have you tried silly simple things like hidden javascript puzzles the browser has to solve?