7 ms·
Web Scraping to Create Open Data
- harperlee 11y agoSo what is the legality of this? Apart from the risk of having someone pull the plug on the way one takes the information out, when is something without a proper license able to be used?
- dsp1234 11y agoIn the US, there is no copyright protection for "facts" on their own. However, a compilation/database of facts can have copyright protections based on a 3 part test[0]. 1. the collection and assembly of pre-existing material, facts, or data; 2. the selection, coordination, or arrangement of those materials; and 3. the creation, by virtue of the particular selection, coordination, or arrangement of an original work of authorship. But specifically there is no protection for the underlying facts themselves, and there is no "sweat of the brow" doctrine. So scraping the data, and rearranging the underlying facts into your own arrangement/organization is almost always not copyright infringement. However, if that data is categorized in some non-trivial way, and you keep that organization, then that is likely to be copyright infringement. However, if what you're scraping are not "facts", but some creative works, such as blog posts, product descriptions, etc, then it is likely to be copyright infringement. Then on top of that, even if there is copyright infringement, other defenses such as a license to use the data, or fair use may apply. [0] - http://www.pddoc.com/copyright/compilation.htm http://www.pddoc.com/copyright/compilation.htm
- toomuchtodo 11y ago> So scraping the data, and rearranging the underlying facts into your own arrangement/organization is almost always not copyright infringement. I'm not so sure. It would definitely be illegal in the US for me to cherry pick data out of Google Maps and add it to OpenStreetMap (and OSM has policies addressing exactly this).
- iolothebard 11y agoYet companies like LexisNexis get most their data they resell this way.
- toomuchtodo 11y agoAre they scraping copyrighted data? Or public records? Big difference.
- iolothebard 11y agoFacts aren't copyrightable. They scrape everything in the world they can get their hands on.
- toomuchtodo 11y agoCollections of facts are: https://www.unc.edu/courses/2006spring/law/357c/001/projects/dougf/node5.html https://www.unc.edu/courses/2006spring/law/357c/001/projects...
- iheartmemcache 11y agoNo one in the US can hold copyrights to the pure 'facts', especially if one demonstrates they invested enough energy to 'creatively reinterpret' it. Scraping hasn't quite seen a Supreme Court ruling yet (@grellas correct me, please), but I'm sure one could make a reasonable argument that the energy invested in re-collating the data is sufficient enough to pass any barrier. See Feist Publications, Inc., v. Rural Telephone Service Co, 1991. and O'Connors opinion.
- ap22213 11y agoWhat part of the law does this fall under? Do people get arrested for this? (i.e. criminal) What's the worst that can happen?
- toomuchtodo 11y ago
- lazyjones 11y agoIANAL but in the EU at least, even databases comprised of simple "facts" are protected. It's a sad state of affairs when i'm not even allowed to scrape data generated using taxpayers' money, like the (required by EU laws) noise maps for cities, which I'd like to use to augment real estate offers, for example.
- Symbiote 11y ago"Europe" would like to partially fund that noise database with income from businesses that use it. The result is less taxpayer money us needed. I think it's only the UK that has copyrightable fact databases
- techdragon 11y agoExcept that doesn't happen because the last thing a new business idea needs is more red tape, paperwork and expenditure.
- rakoo 11y ago"Web scraping to create Open Data" is the exact reason why weboob (http://weboob.org/ http://weboob.org/) was created and still thrives today. CityBikes already seems to be doing a big part of the job, and in Python nonetheless, so it should be easy to integrate its data and use it with Boobsize (http://weboob.org/applications/boobsize.html http://weboob.org/applications/boobsize.html)
- maxaf 11y agoThat naming scheme definitely needs a long, hard rethink.
- snurk 11y ago> ... long, hard ... Heh. Ok, but seriously, I look forward to seeing their appearance on r/drama when twitter discovers this.
- danvoell 11y agoagreed
- rakoo 11y agoIt's funny, everytime Weboob is presented somewhere, and everytime there is a post about the latest version of Weboob, the first comment is a variation of "it's sexist/boobs are unprofessional/grow up", and very very little time is spent talking about the actual thing, what it does and why its only goal is to become irrelevant. Sad thing. Here's what they have to say about it, and why there's very little chance they will change anything: http://laurent.bachelier.name/2013/12/weboob-the-asshole-detector/ http://laurent.bachelier.name/2013/12/weboob-the-asshole-det... (This comment is not directed at you directly)
- maxaf 11y agoIt's a fun read. You know, I'm not normally one to jump on the "offended" bandwagon. In this case I felt compelled to speak up because the name makes it pretty much impossible to reference or recommend the project in a business setting. Wearing a t-shirt to work makes me feel like a true rebel; mentioning a project with components such as "boobsize" and "wetboobs" is going to simply be too weird for many people.
- minimaxir 11y agoI'm not fond of the implication at the end that scraping is justifiable because old websites are dinosaurs without APIs, and those websites are jerks for not doing so, and therefore scraping is the moral thing to do. I've scraped my share of BuzzFeed data and Foursquare data to make data visualizations (with the latter explicitly saying "don't scrape" in their Terms). But if either one told me to stop and take down my results, I would not contest, since data is what drives the Internet ecosystem. (For the record, neither service did; in fact, both tried to recruit me as a result of the visualizations. The difference is that I am not using the data to create a direct competitor that could cause them to lose business.)
- kh_hk 11y agoDisclaimer, I wrote the article. > I'm not fond of the implication at the end that scraping is justifiable because old websites are dinosaurs without APIs, and those websites are jerks for not doing so, and therefore scraping is the moral thing to do. It was not my intention to give that implication. The main implication behind CityBikes is that public services should already provide this information since, well, it is a public service. On the same line, a private company providing a public service should already do so. See motives [1]. > I've scraped my share of BuzzFeed data and Foursquare data to make data visualizations (with the latter explicitly saying "don't scrape" in their Terms). But if either one told me to stop and take down my results, I would not contest, since data is what drives the Internet ecosystem. Same as CityBikes is doing. If we receive a cease and desist, we remove their service from our API. As for Foursquare, I do not see Foursquare as a public service. Your taxdollars at work, and all that. I tried to keep the article balanced but maybe it wasn't clear. There are many transportation companies willing and happy to be scraped, or looking forward to provide their information for people to reuse [2]. [1]: https://blog.scrapinghub.com/2016/03/30/web-scraping-to-create-open-data/#benefits https://blog.scrapinghub.com/2016/03/30/web-scraping-to-crea... [2]: http://nabsa.net/current-members/ http://nabsa.net/current-members/
- seanp2k2 11y agoWhy does your blog intentionally crash browsers that it thinks are Safari?
- l1n 11y agoHeh. I do this with my Student Government data [1]. [1] https://umbc.lin.anticlack.com/finance/ https://umbc.lin.anticlack.com/finance/
- yeukhon 11y agoYou need to fix the certificate before showing that to the public.
- PlzSnow 11y agoCan anyone tell me which cloud provider they are using? I want to make sure that scrapinghub are on the list. I block the IP addresses of all the major cloud providers to prevent parasites such as this.
- jnotarstefano 11y agoI had to use a similar approach when creating a cluster analysis of the amendments in the Italian Senate [0]. The Italian Senate offers a SPARQL endpoint [1], which unfortunately doesn't offer access to the texts of the amendments. So I had to roll my own and create a small spider for them using Scrapy [2]. [0]: https://github.com/jacquerie/senato.py/blob/master/analysis.ipynb https://github.com/jacquerie/senato.py/blob/master/analysis.... [1]: http://dati.senato.it/23 http://dati.senato.it/23 [2]: https://github.com/jacquerie/senato.py/blob/master/senato/spiders/senato_spider.py https://github.com/jacquerie/senato.py/blob/master/senato/sp...