4 ms·
ToS is not a legally binding mutual agreement between the website and the scraper. Only when it's materialized in the form of C&D and it caused damage to the we
by brilliantcode 10y ago
ToS is not a legally binding mutual agreement between the website and the scraper. Only when it's materialized in the form of C&D and it caused damage to the website. In the QVC vs Resultly, they caused a website crashed. In the Craigslist vs 3Taps, the damages were non-existant but they still won because of the C&D letters.
So if a website sends you a letter to stop, do not continue scraping them. Until that happens, it's fair game. ToS is not legally binding. Those two cases were the result of damage caused to the website.
STOP. SPREADING. FUD. cookiecaper. All of your comments are the same unsound legal advice yet Mozenda, Import.io, and a whole bunch of tools & service providers are humming along just fine.
Disclaimer: I'm not a lawyer, this is not a legal advice, consult a real lawyer and not random HN comments.
- cookiecaper 10y ago>In the Craigslist vs 3Taps, the damages were non-existant but they still won because of the C&D letters. (Technically they settled.) Read up on that case and you'll see that even absent a C&D, the judge reasoned that Craiglist's IP ban against 3Taps was a separate incident of affirmatively communicating its intention that 3Taps refrain from accessing the site: > The calculus is different where a user is altogether banned from accessing a website. > The banned user has to follow only one, clear rule: do not access the website. The notice > issue becomes limited to how clearly the website owner communicates the banning. Here, > Craigslist affirmatively communicated its decision to revoke 3Taps’ access through its ceaseand-desist > letter and IP blocking efforts. 3Taps never suggests that those measures did not > put 3Taps on notice that Craigslist had banned 3Taps; indeed, 3Taps had to circumvent > Craigslist’s IP blocking measures to continue scraping, so it indisputably knew that Craigslist > did not want it accessing the website at all. The judge continues to suggest that using proxies at all is atypical and may demonstrate an intention to violate the CFAA. Ruling at http://www.volokh.com/wp-content/uploads/2013/08/Order-Denying-Renewed-Motion-to-Dismiss.pdf http://www.volokh.com/wp-content/uploads/2013/08/Order-Denyi... . >So if a website sends you a letter to stop, do not continue scraping them. Until that happens, it's fair game. ToS is not legally binding. Those two cases were the result of damage caused to the website. No, that's incorrect. You can believe this all you want, and I truly hope that it all goes well for you. It's totally possible that you will never piss off someone who has the resources to file a lawsuit over it. But you should know that you can be bound by a browsewrap agreement (and that it's quite easy to be bound by a clickwrap agreement, which, in practice, may not be much different and which scrapers may automatically follow (e.g., a link that says "click here to enter and agree")). It's really going to come down to the judge's belief that the notification is adequately prominent and that a "reasonably prudent user" would be aware of the stipulations. Quoth from Nguyen v. Barnes and Noble (citations removed): > where, as here, there is no evidence that the website > user had actual knowledge of the agreement, the validity of > the browsewrap agreement turns on whether the website puts > a reasonably prudent user on inquiry notice of the terms of > the contract. [...] Whether a user has inquiry notice of a > browsewrap agreement, in turn, depends > on the design and content of the website and the agreement’s > webpage. Where the link > to a website’s terms of use is buried at the bottom of the page > or tucked away in obscure corners of the website where users > are unlikely to see it, courts have refused to enforce the > browsewrap agreement. [...] On the other hand, where the website contains an > explicit textual notice that continued use will act as a > manifestation of the user’s intent to be bound, courts have > been more amenable to enforcing browsewrap agreements. > [...] In short, the conspicuousness and > placement of the “Terms of Use” hyperlink, other notices > given to users of the terms of use, and the website’s general > design all contribute to whether a reasonably prudent user > would have inquiry notice of a browsewrap agreement. Full decision at https://d3bsvxk93brmko.cloudfront.net/datastore/opinions/2014/08/18/12-56628.pdf https://d3bsvxk93brmko.cloudfront.net/datastore/opinions/201... . >STOP. SPREADING. FUD. cookiecaper. All of your comments are the same unsound legal advice yet Mozenda, Import.io, and a whole bunch of tools & service providers are humming along just fine. I'm not giving any legal advice as I'm not a lawyer. For the third or fourth time here, this is all according to my layman's understanding. It's based on things I learned that time I had to close my business or face a lawsuit from a massive company over just such issues. It's crucial for companies that provide scraping services to be aware of these issues and I know of at least one such company who is aware of them and who takes several precautions to provide some distance from potential legal liability, though they are still not 100% out of the woods. As is usually required in entrepreneurship, they're taking a calculated risk. Should someone wage a legal challenge against their activity, they have millions of dollars in the bank from investors who presumably have researched this and are willing to accept the cost of the potential legal liability. I'm not saying that people shouldn't make businesses that depend on scraping data. I just think they should know what they're getting into before they do so. You are correct that some businesses have been able to engage in such activities without being sued out of existence up to this point. Unfortunately, that doesn't mean that others will be as lucky. I fully agree that anyone who is seriously interested/concerned about this should ask a lawyer. I certainly did. Their answers were not good news for me. Maybe they will be for you.
- brilliantcode 10y agoSo the answer is don't get IP banned. That's easy to solve.
- cookiecaper 10y agoThat's potentially an answer if the judge decides that the browsewrap notice was not sufficiently conspicuous to constitute a binding agreement, etc. I would guess that most judges would not be charitable to someone pretending that they've circumvented this by rotating through proxies pre-emptively. In fact, this would likely work against the defendant as it'd be evidence of willful infringement, which is typically 3x damages. If you can convince the judge that you were just incidentally rotating IPs to protect privacy or something, you might get away with it, but it's definitely not a simple answer, and that still only gets you up to the point of receiving a C&D. I'm not sure why you're trying so aggressively to mislead people about the legal precariousness of data scraping, but at this point it should be clear that this is not a simple matter and it's not something to approach lightly or dismissively.
- brilliantcode 10y agoWell, looks like all the HN user's of Scrapy better lawyer up because Scrapy Cloud offers exactly that, as do rest of the web scraping vendors like Mozenda out on the market. They've all been around for 10+ years, doesn't seem like this is an issue for them.
- cookiecaper 10y agoScrapinghub has several proactive/preventative restrictions on the sites they'll allow users to access because they're trying to avoid such liability. They've been successful up to this point and that's great. That doesn't mean that what they're doing is not a legal grey area. For scraping-related activities, Scrapinghub would probably be the party sued, as was the case in 3Taps, though the clients could probably also be legitimately sued for various things, most obviously copyright infringement. Again, I'm really not sure what you're getting at here. Yes, it's a great idea to check with a lawyer and assess your potential legal exposure. That's why lawyers exist! You can then ask them questions, as Scrapinghub surely has, about how to minimize that potential legal exposure. You definitely SHOULD do that, especially since scraping is more or less illegal in the United States. Courts frequently use an analogy to private physical property to address the matter of accessing a web site. Running a business based on scraping until someone sends you a C&D is roughly the same as running a business based on trespassing on private property until someone serves you with a no-trespass order. Maybe it will work out fine, and most of the time, as long as you leave the property promptly upon request, you probably won't have an issue just because there's no benefit in dragging the matter out further. But that doesn't mean there isn't legal risk involved in running such a business, nor does it mean that you won't be liable for damages incurred whilst trespassing. In such a case, questions about whether the borders of the property were clearly delineated, whether "No Trespassing" signs were posted, whether a reasonable person would've understood they weren't allowed to be there or not, etc., would be asked to determine the existence and/or extent of the trespasser's liability. In the same manner, there is substantial risk involved in running a business whose primary function is to scrape websites, and the same types of questions would be (are) asked in a court case related to network access. People deserve to be informed of that. That's not FUD, it's just the law. If you don't like it, well, most people who know what they're talking about don't either, but that doesn't change the law. Saying "$Party_X hasn't been sued over it!" also doesn't change the law or make the process any less legally risky. If you find this arrangement unsettling or absurd, as you obviously do, I would suggest that you direct your energies/attention to your local representatives, the EFF, and other types of political activism that may help rectify the situation rather than accusing HN commenters of spreading FUD. When you do something illegal, you probably won't get sued for it, because it costs a ton of money to sue someone and it's not likely that you're annoying anyone enough to justify that. This is especially the case if you back off at the first sign of annoyance. That's as much as we can say for your angle. If you're comfortable basing a business on that, be my guest.