4 ms·
One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records fo
by DougWebb 7y ago
One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records for any property without a login. However, the data behind this website it the _source_ of the public records, and not the public records themselves (which would be local government databases).
For years now we've been in an arms race with someone using a botnet to scrape all of the account information for a particular county. My client doesn't care so much about the data; it's the server load that's a problem. Normal activity for this site is a few dozen account searches per minute, but when the botnet gets through our blockade it sends hundreds of search requests per second, overwhelmimg the site. The operator of the botnet has NEVER tried to contact my client to ask for an efficient api to access the data, which they'd probably provide for a minimal fee.
Data hosting isn't free, even if the data is.
- gbasin 7y agoCan you just put it behind a CDN that protects you?
- vorpalhex 7y agoWouldn't the solution be to offer a streamlined download (maybe even as a torrent if you're worried about bandwidth) of all the data then?
- TremendousJudge 7y agoprobably, but as GP said, "The operator of the botnet has NEVER tried to contact my client to ask for an efficient api to access the data" Some people just don't care for the commons
- Symbiote 7y agoI work on a fully open data repository. The website has the API linked in 3 places, so when I find inappropriate scraping I block it with "HTTP 420 ... see <API link> or contact <email>". Some people probably switch to using the API, but no-one has ever contacted us. They either give up, or run their scraper on a different computer -- I've seen the same scraper move between university computers, departments, then (in the evening) to a consumer broadband IP.
- magduf 7y agoI really don't understand why anyone would bother writing and using a web scraper when an API exists. Does the API not provide all the same data/functions as the website? Scrapers are a big PITA compared to just using an API: they're much harder to write to be reliable, and they can break at any time, whenever the site makes even the smallest change. APIs avoid all that mess, and make performance far better too (on both sides), since you're only downloading the data you want, not a ton of Javascript and HTML that you don't.
- barrkel 7y agoAPIs are often not as complete as the web interface, since the customer sees the web interface and normally the customer is what drives the revenue model of the company. If pages are driven via an API, then the API is preferable, but publicly facing websites are often a mix of server-side HTML generation and API enrichment, for caching if nothing else.
- magduf 7y agoIn that case it seems that the webmasters complaining about scraping need to make sure their APIs actually provide access to all the same data, if they want people to use the APIs instead of scraping.
- DougWebb 7y agoIf the scraper contacted the client, said what they need the data for, and (probably) paid for api access, then my client would probably go for it. My client is under no obligation to make access to this data easier. It's not really their data either; the information is property addesses, owner names and addresses, and tax assessments and payments. My client wouldn't want to make it easier for scammers to get that data. So they're not going to do anything unless they know the scraper is legit. If that's the case, the api would require authentication, and any fees would be for the server load, not the data.
- briandear 7y agoFor what purpose? That’s like suggesting that if people keep jumping your fence and trampling your roses because it’s a shortcut to a public park (in this case, the county records office) that already has public access roads, that you should be obliged to build a sidewalk through your garden, at your own expense, when the real answer should be that the public road should be improved.
- vorpalhex 7y agoIf your goal is to get your roses to stop being trampled, it's probably easier to install a few pavers than to spend years petitioning to get a road built. The ideal answer and the efficient answer are not usually the same.
- scotty79 7y agoMaybe you could contact the scrapper? Just post magnet links on the site that allows them to get nicely formatted dump of what they want.
- DougWebb 7y agoWe did figure out who the scraper probably is, but only after several years. For a long time they used an untracable botnet, but after blocking that they eventually switched to a corporate network we traced to a data aggregation company. But we don't know for sure who's doing the scraping; it could be the company, a rogue employee, or a botnet that got loose on their network.
- learc83 7y agoAt work we have all of our data available publicly as easy to parse XML files, but no matter what we do the bot owner's refuse to use it. They'd rather hammer our search engine with sequential searches instead.
- moosey 7y agoI'm in a similar job. We block people from scraping if they break a threshold, but we also refer them to the reporting system, which can get all of the information that they are collecting in a variety of formats. I wonder if something like this would be allowed: if all the public information was available in a well-collated format, then can scrapers be blocked? I imagine that will eventually be fought in court as well.
- wgerard 7y agoYeah, the real problem with scraping is that it's often done very haphazardly and bluntly. Sometimes it's very difficult to tell the difference between a scraper and someone trying to DOS your site.
- danmg 7y agojust put an option to download the raw csvs buried somewhere there. someone who is putting in the effort to bot scrapers will find that link, and save your server the load.
- qihqi 7y agoWould rate limiting be a viable solution?
- DougWebb 7y agoWe've done that, but it's tough to rate-limit a botnet because of the ip address spread. Also, their crappy scraper software doesn't even bother to check if requests are successful; it spews them just as fast no matter how our site responds.
- CWuestefeld 7y agoNo. They botnets works through multiple regions on multiple cloud providers - that's how they achieve such high throughput. For any single IP address, the load is reasonable, but for the whole botnet it's absurd. Currently bot traffic accounts for 2/3 of my load, meaning that the cost of providing my service is 3x what it would be without these persistent bots.