8 ms·
For one Sam scraped under the veil of a corporation which helps reduce or remove personal liability. Second, if the crime was the act of scraping then it’s dir
by menzoic 2y ago
For one Sam scraped under the veil of a corporation which helps reduce or remove personal liability.
Second, if the crime was the act of scraping then it’s directly comparable. But if the crime is publishing the data for free, that’s quite different from training AI to learn from the data while not being able to reproduce the exact content.
“Probabilistic plagiarism” is not what’s happening or even aligned with the definition of plagiarism (which matters if we’re talking about legal consequences). What’s happening is that it’s learning patterns from the content that it can apply to future tasks.
If a human reads all that content then gets asked a question about a paper, they too would imperfectly recant what they learned.
- jazzyjackson 2y agoAsk any LLM to recite lyrics and see that it's not so probabilistic after all, it's perfectly capable of publishing protected content, and the filter to prevent it from doing so is such a bolt-on its embarrassing.
- menzoic 2y agoWe have to understand what plagiarism is if making claims of it. Claiming that you authored content and reciting content are different things. Reciting content isn’t plagiarism. Claiming you are the author of content that you didn’t author is plagiarism. > it's perfectly capable of publishing protected content At most it can produce partial excerpts. LLMs don’t store the data that it’s trained on. That would be infeasible, the models would be too large. Instead, it stores semantic representations which often uses entirely different words and sentence structures than the source content. And of course most of the data is lost entirely during this lossy compression.
- Earw0rm 2y agoThat's a little like saying downloading mp3s isn't music piracy, because it's not encoding the actual music, just some lossy compressed wavelets that sound like it.
- ben_w 2y agoYour username represents a thing that can happen in a human brain which reproduces the perceptual content of a song. Are earworms copyright infringement? If I ask you what the lyrics were, and you answer, is that infringement, or fair use? The legal and moral aspects are a lot more complex than simply the mechanical "what it's done" or "is it like a brain".
- Earw0rm 2y agoAre earworms infringement? No, they stay inside your head, so they exist entirely outside the scope of copyright law. If you ask me the lyrics, fair use acknowledges that there's a copyright in effect, and carves out an exemption. It's a matter-of-degree argument, is this a casual conversation or a written interview to be published in print, did you ask me to play an acoustic cover of the song and post it on YouTube? Either way, we acknowledge that the copyright is there, but whether or not money needs to change hands in some direction or other is a function of what happens next.
- CaptainFever 2y agoNo, the difference is that MP3s can almost completely recreate the original music, while LLMs can't do that with specific pieces of authored works.
- croemer 2y agoYou've got it completely backwards. MP3s just trick you into thinking it's the same thing. It's actually totally different if you analyze it properly, in a non-human sense. LLMs are able to often _preciesely_ recreate in contrast to MP3 at best being approximate.
- CaptainFever 2y ago> LLMs are able to often _preciesely_ recreate Is there actual proof of this? Especially the "often" part?
- theyinwhy 2y agoLLMs (OpenAi models included) are happy to reproduce books word by word, page by page. Just try it out yourself. And even if some words were reproduced wrong, it still would be copyright violation.
- throw5959 2y agoIt's not really some words, it's more like you won't be able to get more than a page out of it and even that is going to be so wrong it's basically a parody and thus allowed.
- angoragoats 2y agoI’d love to see you try to defend this notion in court. Parody requires deliberate intent to be humorous. And courts have repeatedly held that changing the words of a copyrighted work while keeping the same general meaning can still be copyright infringement.
- throw5959 2y agoIt's not just "changing some words". The majority of words will be different, sentences will be different. The general meaning might be generally the same, but I don't think that's enough to claim copyright protection.
- angoragoats 2y agoI didn’t use the word “some.” Please don’t misquote me. As I said in the comment you’re replying to, there’s case law proving you wrong.
- throw5959 2y agoGood thing that's not global, right? I'm not in the US, our courts work differently.
- dkjaudyeqooe 2y agoYour argument might make sense to you, but it doesn't make sense legally. The fact is that “Probabilistic plagiarism” is a mechanical process, so as much as you might like to anthropomorphize it for the sake of your argument ('just like a human learning') it's still a mechanical reproduction of sorts which is an important point under fair use, as it the fact that it denies the original artists the fruits of their labor and is a direct substitute for their work. These issues are the ones that will eventually sink (or not) the legality of AI training, but they are seldom addressed in these sorts of discussions.
- menzoic 2y ago> The fact is that “Probabilistic plagiarism” is a mechanical process, so as much as you might like to anthropomorphize it for the sake of your argument I did not anthropomorphize anything. “Learning” is the proper term. It takes input and applies it intelligently to future tasks. Machines can learn, machine learning has been around for decades. Learning doesn’t require biology. My statement is that it is not plagiarism in any form. There is no claim that the content was originally authored by the LLM. An LLM can learn from a textbook and teach the content, and it will do so without plagiarism. Just as a human can learn from a textbook and teach. Making an analogy to a human doesn’t require anthropomorphism.
- neuroelectron 2y ago> I did not anthropomorphize anything. Machines don't learn. They encode, compress and record.
- Earw0rm 2y agoUntil or unless the law decides otherwise. The 2020s ethic of "copying any work is fair game as long as you call the copying process AI" is the polar and equally absurd opposite to the 1990s ethic of "measurement and usage of any point or dimension of a work, no matter how trivial, constitutes a copyright infringement".
- wiseowise 2y ago
- qwertox 2y agoDo you really think that OpenAI has deleted the data it has scraped? Don't you think OpenAI is storing all this scraped data at this moment on some fileservers in order to re-scan this data in the future to create better models? Models which may even contain verbatim copies of that data internally but prevent access to it through self-censorship? In any case it's a "Yes, we have all this copyrighted data and we're constantly (re)using it to produce derived works (in order to get wealthy)". How can this be legal? If that were legal, then I should be able to copy all the books in a library and keep them on a self-hosted, private server for my or my companies use, as long as I don't quote too much of that information. But I should be able to have all that data and do close to whatever I want with it. And if this were legal, why shouldn't it be legal to request a copy of all the data from a library and obtain access to it via a download link?
- edanm 2y agoWhat? This makes no sense. Of course you're allowed to own copyrighted material. That's the whole point. I have bookshelves worth of copyrighted material at home at this very minute. If you're implying that the scraping and storing of the things itself breaks copyright, then maybe, but I don't think so? If you're saying that training on copyrighted material breaks copyright, then yes, that's the whole argument. But just having copyrighted material on a server somewhere, if obtained legally, is not by itself illegal.
- Terr_ 2y agoNot the other poster, but chiming in here that: > If you're implying that the scraping and storing of the things itself breaks copyright, then maybe, but I don't think so? Suppose I "scrape and store" every book I ever borrow or temporarily-owned, using the copies to fill several shelves in my own personal library-room. Yes, that's still copyright infringement, even if I'm the only one reading them. > But just having copyrighted material on a server somewhere, if obtained legally, is not by itself illegal. I see two points of confusion here: 1. The difference between having copies and making copies. 2. The difference between "infringing" versus "actually illegal." Copyright is about the right to make copies. Simply having an unauthorized copy is no big deal, it's making unauthorized copies where you can get in trouble. Also, it is generally held that the "copies" of bytes in the network etc. do not count, but if you start doing "Save As" on everything to create your own archives of the news site, then that's another story.
- ajb 2y agoIt's weird that corporations get to remove liability from people acting on their behalf. The law is only that they remove liability for debts. If you act on behalf of a corporation to commit either a criminal offence or civil wrong, as far as I can tell (I'm not a lawyer) you are guilty of at least conspiracy. And yet, it's rare for individuals to be prosecuted for such offences, even for criminal offences. We treat the liability shield as absolute. It may seem unfair to prosecute the little guy for "just following orders" but the fact that we don't do it is what allows corporations to offend with impunity.
- nonrandomstring 2y agoIf you join a gang you get protection from that gang. There are some gangs you can join where you get to kill people, with total impunity, if you're into that sort of thing. Corporations are just another kind of gang. Don't let the legal sugar frosting fool you.
- ajb 2y agoYes, but I don't think that's the full explanation. It seems to apply even to companies of modest size, which should not be able to intimidate.
- nonrandomstring 2y agoDefinitely not the full explanation. I made a strident, blunt statement to move the discussion along. If you follow the Hobbsian thing the biggest gang is the state. And it's supposed to be "our" gang cos we vote for it and give tacit assent to a violent monopoly we call "the law". What's happened is the law has become very weak, and so has the democracy that keeps the biggest baddest gang on the side of 'the people'. I think US society, and small companies as you say, sense that the West has been "re-wilded". In the new frontier anything goes so long as you have the money. Aaron was a principled, smart and courageous dude, but he was acting without support from a strong enough base. The gang he thought he was in, MIT, betrayed him. (Chomsky has plenty of wisdom on what terrible cowards universities and academics really are.) The law that should have protected someone like Aaron was weak and compromised. It remains so. The same (lack of) state and laws are now protecting others doing the same things as Aaron, at an even bigger scale, and for money. At least that's how I read TFA.
- sgt101 2y ago>Second, if the crime was the act of scraping then it’s directly comparable. But if the crime is publishing the data for free, that’s quite different from training AI to learn from the data while not being able to reproduce the exact content. They often do reproduce the exact context; in fact it's quite a probable output >“Probabilistic plagiarism” is not what’s happening or even aligned with the definition of plagiarism (which matters if we’re talking about legal consequences). What’s happening is that it’s learning patterns from the content that it can apply to future tasks. That's what I think people wish would happen. Sometime they have been shown to learn procedural knowledge from the training data but mostly it's approximate retrieval.
- CaptainFever 2y agoDo you have proof of this?
- sgt101 2y agoProof? :) where has proof ever been found in truth? It's not proof, but there is some good evidence out there. Here are two interesting papers I found informative. https://arxiv.org/html/2403.04121v2 https://arxiv.org/html/2403.04121v2 https://arxiv.org/html/2411.12580v1 https://arxiv.org/html/2411.12580v1
- zx10rse 2y agoAn algorithm that predicts the best next possible outcome, with trillions of parameters can hardly be called intelligence in my book, it is an artificial I will give you that. I am old enough to remember a bot called SeNuke which was widely used 10-15 years ago, in the so called black hat SEO community, the purpose of the bot was to feed it with 500 words article, so the words can be scrambled in a way to pass the Google algorithm for duplicated content. It was plagiarism 101, now I don't recall anyone talking about AI back then or how all the jobs of copy writers will extinct, and how we are all doomed. What I remember is that every serious agency would not use such tool so that they can't be associated with plagiarism and duplicate content bans. Maybe it is just me but I cannot fathom the craziness, and hype of a first person output. What we get now with LLM models it not simply an output of link and description of lets say a question like What is an algorithm? We get an output that starts with "Let me explain" ... how is this learning and intelligence? We are just witnessing the next dot com boom, the industry as whole haven't seen such craziness despite all the efforts in the last 25 years. So I imagine that everyone wants to ride the wave to become the next PayPal mafia, tech moguls, philanthropist, inventors, billionaires... Chomsky summed it best. RIP Aaron
- sambeau 2y agoIf I make an MP3 of a song or a JPEG of an artwork I cannot use them to reproduce the exact content, but I will still have violated the artist’s copyright.
- nialv7 2y agoSo, let me get this straight. Scraping data and publishing it so everyone get equal and free access to information and knowledge, arguably making the world better, is a crime. But scraping it to benefit and enrich just yourself, is A-OK?? This legal system is truly fucked.
- CaptainFever 2y agoNo, if you scrape and create a open-weight model, that is OK too. The difference is if you re-publish the original dataset, or just the weights trained from it. Please stop mis-interpreting posts to match doomer narratives. It's not healthy for this forum.
- tsimionescu 2y agoTalking about plagiarism here is a complete red herring. Plagiarism is essentially a form of fraud: you are taking work that someone else did and presenting it as your own. You can plagiarize work that is in the public domain, you can even plagiarize your own work that you own the copyright to. Avoiding a charge of plagiarism is easy: just explicitly quote the work and attribute to the proper author (possibly yourself). You can copy the entirety of the works of Disney, as long as you are attributing them properly, you are not guilty of plagiarism. The Pirate Bay has never been accused of plagiarism. And plagiarism is not a problem that corporations care about, except insofar as they may pay a plagiarist more money than they deserve.' The thing that really matters is copyright infringement. Copyright infringement doesn't care about attribution - my example above with the entire works of Disney, while not plagiarism, is very much copyright infringement, and would cost dearly. Both Aaron Swartz and The Pirate Bay have been accused and prosecuted for copyright infringement, not plagiarism.