3 ms·
You can't copyright facts. As journalists, we scrape things to collect information used toward transformative analysis. Not straight-up mirroring. Facts, as
by pp19dd 7y ago
You can't copyright facts.
As journalists, we scrape things to collect information used toward transformative analysis. Not straight-up mirroring. Facts, as stated by an entity. So we've never run into a legal issue doing this as long as we used the scrapes to synthesize results into data. For example, map of restaurant closures by the health department, with statistics and graphs of violation frequencies. Or analysis of lawyer performance by cross-referencing a state judiciary database search with their team member lists for success rates and other stuff.
Most of the sites we scraped were county, state or federal government sites and they contained information available in the interest of the general public. However, we crawled tons of private sites as well and as long as we wore white hats we considered it fair game.
We typically tried to scrape things fairly without causing technical issues but to be honest, we ignored robots.txt directives all the time but timed it do happen during off hours, with backoff mechanisms in case we contributed negatively to computational loads. The typical issues we ran into were overeager system administrators who squashed or interfered our scraping attempts under their personal interpretation of appropriateness. Sometimes they sicced misguided lawyers after us. Most of the lawyers couldn't tell you what you were doing wrong, let alone how a site is registered, what a glue record is, how DNS differs from IP registry ownership, how collocated servers work and who owns them. They couldn't prove what we did with any of the data to even imply we violated any copyrights.
So we relied on our legal departments to clear the way in case of issues, but in 15 years of doing this, I've never once had a legal issue come up and put a stop to what we were doing under that operating premise of transforming the information into data. Our legal team never got involved for that sort of thing. There were issues, but they got resolved through communication or by reconfiguring our scrapers. Even when we've also made the raw data available to the public or other researchers, it hasn't come up as a problem.
In one case, a police department figure blocked us because they disliked our coverage. Their pretense was that our geocoding wasn't accurate enough from the information they provided, and rather than circumvent their blocking we had face to face meetings to address those concerns and mollify their concerns, on the record. They ended up providing us with additional information to meet that accuracy. In another case, the CEO of a large private company personally threatened us legally claiming we violated their terms of services for their API endpoint. However, their terms of services mentioned nothing about data retention once something became a data point and we felt we were in the clear so we kept doing it for years and nothing came of it.