14 ms·
How can TOS have legal power for the case scraping? A website is a public property. If I'm visiting it without logging in, I don't have a chance to accept TOS.
by twa927 10y ago
How can TOS have legal power for the case scraping? A website is a public property. If I'm visiting it without logging in, I don't have a chance to accept TOS.
Imagine a hotel that makes guests sign a document saying they will not make photographs of the building. If I'm not a guest, I can take photographs of it and I can't even know that would be illegal.
- minimaxir 10y agoThat analogy is not equitable. If you take photographs of a building while on the building's property, they have the right to tell you to stop, or call the police to escort you off if you refuse to do so.
- mindslight 10y agoSure, but they do not have the right to retroactively declare you as having been trespassing, nor even to preemptively put up a "no photography" sign and have you arrested for trespassing if you disobey it. The entire point of protocols is to precisely define the terms of communication. The status code is '200 OK', not '200 OK/Asterisk'. But of course if lawlers didn't force themselves into the situation, they'd be out of jobs. As an aside, I'd really like to see a browser plugin that would scrape sites in the normal course of access, storing the proceeds in a distributed public database.
- cookiecaper 10y ago>As an aside, I'd really like to see a browser plugin that would scrape sites in the normal course of access, storing the proceeds in a distributed public database. This would be copyright infringement, since the content of the page is a substantive unique work that is automatically copyrighted by its author. A site that doesn't want you scraping its content is not going to want you posting dumps of its pages. Much like BitTorrent, they'd get into the protocol and send subpoenas to the ISPs behind the IPs that serve their pages, and use that info to sue the customer. When my company was shut down by a legal threat related to scraping, I did suggest to my lawyer that we create something like a browser extension that would grab the data we needed out of normal client-side browsing sessions. This wouldn't be as nice as controlling the flow of information ourselves but it would've worked OK. My lawyer strongly suggested avoiding that as it could've been construed as conspiratorial conduct that would've made criminal prosecution more likely.
- niftich 10y agoNot the discount the validity of your experience, but the usual counterpoint to this is Google, who (like mentioned elsewhere in the thread) has been continuously scraping since the very beginning and in fact built their entire business model on doing so. They are also responsible for advancing the state-of-the-art of scraping (albeit mostly internally), through the development of V8 and headless Chromium so that they can inspect dynamic pages too. Perhaps this illustrates the fungibility of the legal system: it's an inherently human construct that pits a plaintiff against a defendant, and given a big enough warchest and persuasive-enough arguments, catastrophe can be avoided -- by Google; perhaps not by you, me, or someone else.
- cookiecaper 10y agoYeah, Google violates the CFAA and infringes on copyright as a matter of course. Their service would be impossible if they weren't doing so. The main difference when Google was small was that Google was not dependent on any data source in particular, so even if someone denied their robot or sued them, they could cease and desist without affecting the overall value of their offering. This is different if you are getting data that is only available from one or two sources. Now, the main difference is that Google is one of the biggest companies in the world, and they'll sick an army of $1,000/hr lawyers on you if you even think about taking legal action against them. The only people who can afford to fight are other big companies, but that's not going to happen because they all depend on breaking the CFAA for their own purposes and then using their position as a huge company to bully small innovators.
- wtracy 10y ago> they'll sick an army of $1,000/hr lawyers on you They don't even need to do that. They just cheerfully agree to not scrape you, and wait for you to come back and beg to be re-instated when your search traffic plummets.
- mthoms 10y agoGoogle's crawling and caching has been largely found to be fair use and thus is not considered to be infringing copyrights. https://en.wikipedia.org/wiki/Field_v._Google,_Inc https://en.wikipedia.org/wiki/Field_v._Google,_Inc. There are similar rulings for thumbnail images: https://en.wikipedia.org/wiki/Perfect_10,_Inc._v._Amazon.com,_Inc https://en.wikipedia.org/wiki/Perfect_10,_Inc._v._Amazon.com.... And of course books: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
- Asooka 10y agoWell, it's just a technical response code. 200 OK - everything went as normal, here's your data. By the same margin, the door on a shop doesn't stop you walking out without paying and the road markings don't stop you from driving in the wrong lane. I think imbuing technical protocols with legal implications would be even worse than the current situation since then changing anything on a protocol would require changing the law and getting a protocol implementation slightly wrong would carry real-world legal repercussions on the order of licensing your work in the public domain rather than retaining copyright. Let the lawyers make the law and check the human terms of service before using the data. Trying to out-lawyer the lawyers is like challenging a hedgehog to a butt-kicking brawl.
- Doctor_Fegg 10y agoThe protocol is also that you send a valid, non-faked User-Agent: "The User-Agent request-header field contains information about the user agent originating the request. This is for [...] the tracing of protocol violations [...]. User agents SHOULD include this field with requests" Many scrapers disregard this part of the protocol. Of course, whether a headless browser should send a different UA is an interesting question. https://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html https://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html
- TheCoelacanth 10y agoUser-Agent is a SHOULD, not a MUST. There are also practically no browsers that send a non-fake User-Agent, since they almost all claim to be Mozilla/5.0.
- jessaustin 10y agoOne rarely visits corporate property in order to access corporate websites. The analogy may be flawed, but this objection to it is as well.
- ysavir 10y agoLet's go with a more apt analogy: If you're entering a country, do its laws not apply to you until you've seen a copy of them? "Oh, sorry, no one told me theft is illegal here. Where does it say that? Oh, I see. Okay. I'll stop now. Thanks for letting me know." If you cross the border without necessary documents, does that country have no right to detain you, simply because you haven't checked the laws? Just because a website is visible and public doesn't mean its content is public domain. It just means that your first order of business as a user should be to check the terms of service. Sure, most people using a website probably don't need to--same as not needing to check a country's stance on murder--and so can just use the website as intended without violating the terms. But when you plan on using it in a way that might not be intended, and you don't check the terms of service, well, that's on you.
- slrz 10y agoBut when you plan on using it in a way that might not be intended, and you don't check the terms of service, well, that's on you. I don't need to check your terms of services if I'm doing something that I'm allowed to do by law anyway; the TOS cannot deny me those rights (they might, of course, grant me additional rights provided that I follow certain conditions).
- baddox 10y agoCountry's laws are a bit different, simply because a country has virtually absolute legal power over its territory. Countries can and do punish people for breaking laws that one cannot feasibly know they were breaking. Does any human know all the laws in the United States? Would that even be physically possible?
- mentat 10y agoThere's some interesting science fiction opportunities here. When you open a connect to a site then all traffic over that connection is subject to the jurisdiction of the ToS for that site regardless of disclosure. Also we don't even know how many laws there are in the United States for I'd say knowing the content is impossible.
- 10y ago
- baddox 10y agoRegardless of whether that would be reasonable, is it actually true? I know that the United States has specific rules for "public accommodations," which are private properties that are generally accessible to the public, like retail businesses. Property owners in this case don't have complete control over who enters their property. The obvious example is refusal of service due to membership of a protected class like race or religion. So I'm not so sure that police will escort you out of a Walmart because they caught you taking a picture of the parking lot with your smartphone.
- cookiecaper 10y ago>How can TOS have legal power for the case scraping? A website is a public property. If I'm visiting it without logging in, I don't have a chance to accept TOS. This is called "clickwrap". There is usually a notice in the footer of each page that says something like "By using this site, you agree to our Terms of Service." Typically, this kind of notice has been held enforceable. More recently, judges have been demanding that such notices be placed more prominently before they're held enforceable (e.g., somewhere above the fold), but that's it. >Imagine a hotel that makes guests sign a document saying they will not make photographs of the building. If I'm not a guest, I can take photographs of it and I can't even know that would be illegal. The reasonable laws that exist in meatspace are not applicable online, because once you hit someone else's server, you're considered to be on their property and they have the right to control what you do there. There is no "public property" from which to safely stand and take photographs in the internet. Also, photographs of structures may not be free to use. Architectural copyrights went into effect in the early 90s and have a term of either 90 or 120 years. Thus, if you take a photograph of a building built in 1991 and the year is not yet 2111, there is a chance that the architect can claim infringement.
- wang_li 10y agoI have a custom X-TOS header in all of my http/https requests that states that the company who owns rights to the website my request is sent to and replies with data owes me: 1. Total privacy, they will not track me activity on their website, including any logs. 2. They will send me a cashier's check for $1,000 for each byte that they send to me. 3. They will provide me with Mana Sakura's cell phone number. I'm still waiting for checks and a phone number.
- cookiecaper 10y agoIf you can convince a judge that this represents an enforceable contract, as has been done and established with clickwrap, then you should be able to get what you're owed. :) It is ridiculous. Something like "pagewrap" can't trump the consumer protections that apply to a physical good like a book, it would be laughed off. But the law doesn't contemplate network access so reasonably.
- dragonwriter 10y ago> A website is a public property. No, its not. It may be in public view, but that's a different issue.
- makmanalp 10y agoThat's an interesting analogy - though you're allowed to take photographs of whatever is in public view in many jurisdictions. Now if you wanted you could take this argument to the extreme, but surely there's some parallel between sending and receiving photons across the border of someone else's property (perfectly agreeable) and sending and receiving requests?
- gsnedders 10y agoBoth the LinkedIn and OKC cases involved the scrapers using logged in accounts.
- hooph00p 10y ago> A website is public property. This is a gross misunderstanding of how the internet works.
- sseveran 10y agoOne of the original court cases covering this was eBay vs Bidders Edge. https://en.wikipedia.org/wiki/EBay_v._Bidder%27s_Edge https://en.wikipedia.org/wiki/EBay_v._Bidder%27s_Edge The courts have generally disagreed with that interpretation.
- duaneb 10y ago> A website is a public property. This isn't even true metaphorically. It's like a shop front: there may be public access, but it is NOT public property.
- taftster 10y ago"public property" may not be the correct metaphor. But neither is "shop front" correct. Taking the store metaphor further, it would be more like you knocking on the front door of a clothing store and the store owners open the door and throw every possible piece of clothing at you, shirts, shorts, underwear, including coupons to "partner" stores, when all you wanted was a pair of pants. Upon knocking, if the store owner hands you instructions on how to enter their store and interact with their products in a personalized shopping experience, that would be one thing. But when the clothing owner throws everything at you at once, what they flung at you is for all practical purposes public property.
- pyre 10y agoA better metaphor for this would be the "Sunday Flyers" that come in the newspaper (e.g. for big box stores like Best Buy). They sent that information to you, they can not then attempt to restrict how you use that information (though they have tried to claim copyright over pricing against sites that aggregate the flyers).
- buro9 10y agoThe UK has a database law: https://en.wikibooks.org/wiki/UK_Database_Law#Database_Right https://en.wikibooks.org/wiki/UK_Database_Law#Database_Right If you scrape, and effectively reconstitute a database, then so long as the database originally had a "substantial investment" in it's "obtaining, verifying or presenting the contents" then yup... you have breached the database right, which is a modified form of copyright. You may access said database (via the web), but as soon as you start reconstituting the database from scraping... you're in breach. It's a law, it is illegal in the UK, I'm sure most countries have some equivalent law on their books, all of the EU does. The law looks recent, but UK copyright and patent used to cover it, the 1997 date is just a separate statute to clarify the position.
- flukus 10y agoWhat is the definition of "reconstituting a database"? Aren't googles indexes doing that?
- cookiecaper 10y agoYes. Google is violating practically every law of this type. They're allowed to do it because they have a lot of money.
- greglindahl 10y agoActually, such database laws are rare. The US and Canada don't have one. See Feist v. Rural Telephone for an example of databases getting scraped & the scraper winning in court.
- devbug 10y agoExactly. You can't copyright facts.
- cookiecaper 10y agoYou can't copyright facts in the US. You effectively can in the EU, as the grandparent discussed, as long as you demonstrate that it took significant investment to arrange the compendium of facts from which they were drawn.