6 ms·
This is the first I've heard of this. Where/what laws prohibit web scraping?
by brogrammernot 10y ago
This is the first I've heard of this. Where/what laws prohibit web scraping?
- jacquesm 10y agohttps://en.wikipedia.org/wiki/Web_scraping#Legal_issues https://en.wikipedia.org/wiki/Web_scraping#Legal_issues https://www.wired.com/2010/10/hacking-captcha/ https://www.wired.com/2010/10/hacking-captcha/ https://www.wired.com/2010/03/wiseguys-indicted/ https://www.wired.com/2010/03/wiseguys-indicted/ http://www.nj.com/news/index.ssf/2011/06/wiseguy_ticket_operators_get_p.html http://www.nj.com/news/index.ssf/2011/06/wiseguy_ticket_oper...
- cookiecaper 10y agoIt's almost always illegal in the United States. It's prohibited by a combination of the CFAA, copyright law, and contractual obligations imposed by Terms of Use, which are usually considered applicable if you load more than one page ("browsewrap"). The CFAA makes it a crime to access any computer network without authorization or in excess of granted authorization. The Terms of Use will usually prohibit "any automated or mechanical access" or use similar boilerplate that can be construed as a restriction on automated access. The implied license to make a copy of the page in RAM is no longer applicable and the scraper is thus infringing copyright. Relevant cases are Craigslist v 3Taps, Facebook Inc. v Power Ventures Inc., and several others. This is at the point where it's basically well-established. The exception is Perfect 10 v. Amazon, where judges ruled that since it was Google and they don't want to break Google, it's OK. Copyright law allows such evaluations because each judge must decide whether a use was "fair" or not.
- TeMPOraL 10y agoThat's a very, very sad turn of events, and I have to wonder, how did we get there? I'm increasingly feeling that the law is giving way too much control over content published on the Internet to the publishers.
- cookiecaper 10y agoI agree. There is a lot more fairness in physical space that doesn't translate to cyberspace primarily due to the implementation details of computers and networks. Whereas products and machines built in the real world are primarily protected by things like patents and trade secrets, practically everything in the digital world falls under uber-restrictive copyright protections, since the "creative" work of code and its compiled/interpreted derivatives is the language by which everything is implemented. Similarly, concepts like the "first sale doctrine" are becoming less applicable with digital delivery, as it's impossible to identify a "hard copy" of something that may be eligible for resell. That completely obliterates the secondary market for many products that are accessed through computers, including software, games, movies, and books. The CFAA essentially allows network operators to arbitrarily make someone a felon overnight. Reddit co-founder Aaron Swartz is the most prominent example of this; his criminal prosecution under the CFAA (for scraping publicly-funded research papers out of a database) was pending when he committed suicide. We badly need digital rights reforms, but since major companies have been allowed to profit handsomely off these shifts and since they find it rather convenient to bully small innovators with serious legal threats, which are easy to craft in this climate, it doesn't seem that anyone is making this a priority.
- tyingq 10y agoThere's even a very highly specific (online ticket sales) bill that passed congress: https://www.congress.gov/bill/114th-congress/senate-bill/3183 https://www.congress.gov/bill/114th-congress/senate-bill/318...
- codedokode 10y agoDoesn't publishing information on a public web server equals to granting authorization to download it? Why publish it otherwise? And copyright laws are supposed to protect only creative works, not every page on the web.
- cookiecaper 10y ago>Doesn't publishing information on a public web server equals to granting authorization to download it? Why publish it otherwise? This is the argument that there is an implied license. The counterargument is that the user agreed to the Terms of Use which explicitly defined automated access as violative. In addition, these cases generally begin with a cease and desist demand letter, which explicitly informs the allegedly-infringing party that the publisher believes their rights are being violated and that they must cease and desist immediately. If the TOS argument doesn't hold, the C&D will surely qualify as a revocation of any implied license to access the content for copyright purposes. It generally also serves as explicit notice that the publisher considers the accessor to be "exceeding authorized use" of their computer systems, which is a crime under the CFAA. In Craigslist v 3Taps, the judge also commented that needing to circumvent IP bans should've made it obvious that 3Taps was "exceeding authorized access" under the CFAA. >And copyright laws are supposed to protect only creative works, not every page on the web. Copyright law protects all works of sufficient originality. Pretty much the only thing it doesn't protect is a plain list of facts (and in the European Union, it even protects that, known as "database rights"). The minimum standard of originality for copyright protection applies to practically every page on the web, yes. In effect, this means that you can copy a list of names and addresses from a phone book, but you can't copy the layout. Since you can't access a web page without making a copy in RAM, if the publisher has revoked your license to access the content, even accessing and extracting the raw factual information within the body of the page is an infringement (because your RAM copy is an infringing copy). IANAL and this is based on my layman's understanding.
- codedokode 10y agoI think that is wrong too. I think that copyright laws should make a distinction between making and distriubuting a copy by a person (for example uploading copyrighted file to a website) and technical processes that happen inside a computer. Copying something from NIC buffers to memory should not be "copying" under copyright law.