5 ms·
Every request from a different place ultimately is fulfilled by the original host barring some extreme caching (which you can also use http headers to instruct
by falien 17y ago
Every request from a different place ultimately is fulfilled by the original host barring some extreme caching (which you can also use http headers to instruct against). This is completely different than the case of it being taken from the original host and put up on scribd, which could easily be a copyright violation. An HTTP 200 response is implied consent for an end-user (or even a robot) to view it, not to redistribute it. AFAIK scribd does not crawl for content (with them then hosting that would be blatantly illegal in the US) so robots.txt is not really applicable.
- Zak 17y agoAFAIK scribd does not crawl for content (with them then hosting that would be blatantly illegal in the US) As far as I know, they don't do that, but if they did, how would it be different from Google cache?
- mos1 17y agoGoogle cache doesn't beat your main site in the rankings, and it clearly cites you as the original source of the data. That said, it's irrelevant because they don't crawl for content. The HN admins are way out of line with their practices on this one. They're making illegal copies to help their friends at Scribd, and doing so without the requisite consent.
- natrius 17y agoSearch rankings have nothing to do with copyright law, which seems to be the basis of your argument. You have yet to give a decent argument that materially distinguishes Scribd from Google's cache, caching proxies, or even the basic routing required for any request on the internet, which copies content by definition. See also: http://docs.google.com/viewer http://docs.google.com/viewer
- deleted 17y ago[deleted]
- deleted 17y ago[deleted]
- mos1 17y agoI was just responding to Zak's comment about differences, giving a few. It wasn't meant as a complete argument of anything, or any sort. That said, two points: 1) doing so is not just immoral, it's against Scribd's TOS. 2) your claim that re-hosting content without permission is indistinguishable from transport makes clear that your beliefs are so different than mine that I cannot possibly find a way to communicate with you.
- natrius 17y agoThe basic concept of copyright law is that an author is the only one who is allowed to make copies of her work, and only she can give others permission to do so as well. At a technical level, sending a response to a request involves telling another machine to pass a message along for you, so there is at least implied consent to copy it and send it to another user. However, what are the bounds of this consent? Can a router store the data it has been passed? For how long? Can it serve it to others besides the IP address the response was intended for? There may be case law and/or actual laws that clarify these points, but I am not a lawyer. I presume you aren't either. If you think the concept of IP law has a straightforward and indisputable application to the internet that clears all questions about what routers can and can't do with the data they are passed, feel free to explain. It seems like you're operating within the "lots of people do it, so it must be legal somehow" school of thought when it comes to routing. This isn't necessarily a problem as long as it's applied consistently. Several other commenters and I effectively made the same argument in saying that Google does almost the same thing Scribd does in terms of copying and redistributing content. It isn't possible for you to call that argument invalid then rely on that argument as proof of routing's obvious legality. Scribd is legally indistinguishable from a caching proxy. Feel free to let Opera and all the other caching proxy operators know. http://www.opera.com/business/solutions/turbo/ http://www.opera.com/business/solutions/turbo/
- mos1 17y ago... Scribd is legally indistinguishable from a caching proxy. You already said you weren't a lawyer. No need to prove it dramatically.
- Zak 17y agoYou have a valid point about ethics, but I was talking about copyright law. The two are not closely related.
- natrius 17y agoThis interpretation makes caching illegal and puts routing in a gray area. If not instructing against caching is implied consent to redistribute the content, then you're essentially agreeing with me. robots.txt is indeed intended for crawling, but if it's there and you redistribute someone's content anyway, I'd consider it less defensible.