7 ms·
Sorry if i don't understand this totally, but why is scraping for pdf and ppt/pptx files and mirroring them illegal? If you can reach that files just scraping
by EMRZ 8y ago
Sorry if i don't understand this totally, but why is scraping for pdf and ppt/pptx files and mirroring them illegal?
If you can reach that files just scraping it means they are somehow open to public access.
No joke, i am genuinely asking.
- consp 8y agoScraping, mostly no. Using them to make profit: yes. Since you then violate the copyright of the author as the original action was solely to make it public (assuming no profit was intended). edit: Rehosting is not allowed as far as I know in Dutch law if you are attempting to make profit of it by not requesting it from the original owner. Rehosting it and not taking advantage and linking to the original article is allowed as far as I know. But maybe ask your lawyer for info if you really want to know.
- jerrre 8y agoNAL but I'd say the profit is not the problem. Rehosting it (also without credit) is.
- ForHackernews 8y agoHuh? IANAL, but that doesn't correspond to my understanding of copyright laws.
- prepend 8y agoIf a file is available through non-auth http it’s unclear what the copyright is. I think it depends on the particular item whether copyright was violated. Making a file public without restriction doesn’t mean another can’t make profit (my ISP makes a profit by transmitting the file to me; gmail makes profit when I email the file to me; google makes a profit when they cache the file; archive makes a profit when they archive a file; etc etc). I think if this were non-public docs then the case is clearer. But by releasing a document publicly with unlimited access via URL the author explicitly allows unlimited distribution (and due to the nature of tcp/ip redistribution). If an author wants to restrict distribution then they should restrict distribution using available protocols.
- consp 8y agoSince the article contained some private Dutch governmental documents, I think you are correct to state the last case (some non-public). Though if the are public, reselling them (or making profit as a side business) without referring to the original document is problematic at best, illegal at worst. There are plenty of sites who reference or even quote published documents from the overheid.nl (goverment) websites and make profit from them, but not by copying them without reference. Though they all serve some legitimate purpose and you can actually download them instead of having to click though quizzes and other annoying things in the hope you will click one advertisement.
- JdeBP 8y agoThis is a common copyright fallacy. So common that Brad Templeton was combatting it 20 years ago. Publication does not waive copyright. * https://www.templetons.com/brad/copymyths.html https://www.templetons.com/brad/copymyths.html
- dragonwriter 8y ago> If a file is available through non-auth http it’s unclear what the copyright is. Protocol has nothing to do with copyright. “Non-auth http” doesn't suddenly make copyright murky. > But by releasing a document publicly with unlimited access via URL the author explicitly allows unlimited distribution (and due to the nature of tcp/ip redistribution). There is an implicit license to exactly the redistribution necessary to effect access by URL, sure, but this is not “unlimited distribution”. Republishing is, particularly, not what is licensed.
- JdeBP 8y agoThere is no implicit licence. This is an oft-circulated but fallacious rationalization made by computer people that is not in line with the law. The law does not recognize any such thing, and until the turn of the 21st century this was a recognized hole. Technically, the World Wide Web (and indeed other systems from FidoNet to SMTP) was a violation of many countries' copyright laws. Legislators fixed the hole, but not by introducing the idea of implicit licencing. The EU introduced provisions in 2001, by way of a Directive, that the making of temporary transient copies for the likes of HTTP and other network transmission mechanisms to work was explicitly not copyright violation in the first place. This became Netherlands law in 2004, and U.K. law in 2003. Neither the Directive nor the implementing Acts and Instruments talk about licensing. It is simply not a violation of copyright by definition. * https://eur-lex.europa.eu/legal-content/en/ALL/?uri=CELEX:32001L0029 https://eur-lex.europa.eu/legal-content/en/ALL/?uri=CELEX:32... * https://www.ivir.nl/publicaties/download/RIDA2005_206.pdf https://www.ivir.nl/publicaties/download/RIDA2005_206.pdf * http://wetten.overheid.nl/jci1.3:c:BWBR0001886&hoofdstuk=I¶graaf=5&artikel=13a&z=2018-10-11&g=2018-10-11 http://wetten.overheid.nl/jci1.3:c:BWBR0001886&hoofdstuk=I&p... * https://www.legislation.gov.uk/uksi/2003/2498/regulation/8/made https://www.legislation.gov.uk/uksi/2003/2498/regulation/8/m... * https://www.legislation.gov.uk/uksi/2003/2498/note/made https://www.legislation.gov.uk/uksi/2003/2498/note/made
- pjc50 8y agoJust because they're published on the internet doesn't mean you can rehost them - that's copyright infringement. (The Internet Archive is very careful to skirt the edge of what's permissible and what they can get away with in this regard)
- EGreg 8y agoSo copyright infringement of this form may or may not be immoral depending on who you ask. And may even be legal depending on where you live
- Cthulhu_ 8y agoI get the impression ATM that copyright infringement is legal as long as nobody complains. When one person complains or files a DMCA takedown notice, no problem, it gets taken down. It needs to be a big company's documents - like Elsevier or (in the Bittorrent era) music publishers - before a major legal action is undertaken to take a site offline. TL;DR it's not a crime if nobody complains. Now that someone complained though, things might end up badly for the owner.
- heinrichhartman 8y ago
- stevespang 8y agoIf the site owner is actually a Russian living in Russia - - law enforcement and courts of democratic countries cannot get access to him, Russia does not extradite it's citizens, only Google has the power to stop this (but they obviously have a conflict of interest because like credit card co's and real estate agents and so many others they only make $$$ if the deal goes through).
- aasasd 8y agoIf you don't have the author's or publisher's permission to republish, you're in breach. That's the entirety of practical crux of copyright law. In practice, in some cases the authors' intent is for others to republish, especially on the web. So they deliberately don't pursue damages―but they still have the right to do so, and users are still not entitled to anything unless there's a specific broad license. Republishing files en masse stretches this unspoken agreement. People keep grossly misinterpreting copyright law, because they got used to the share culture on the web. "Fair use" is especially invoked left and right where it doesn't apply. "Public access" is also not a thing in the law (afaik).