7 ms·
Why training AI can't be IP theft
- re-thc 1y agoThe argument in the article breaks down by taking marketing by definition and try to apply it to a technical argument. You might as well start by saying that the "cloud" as in some computers really float in the sky. Does AWS rain? This "AI" or rather program is not "training" or "learning" - at least not the way these laws conceived by humans were anticipated or created for. It doesn't fit the usual dictionary term of training or learning. If it did we'd have real AI, i.e. the current term AGI.
- Latty 1y agoI agree you can't just say it's learning and be done with it, but I think there is a discussion to be had about what training a model is. When they made the MP3 format, for example, they took a lot of music and used that to create algorithms that are effective for reproducing real-world music using less data. Is that a copyright violation? I think the answer is obviously no, so there is a way to use copyrighted material to produce something new based on it, that isn't reproduction. The obvious answer is that MP3 doesn't replace the music itself commercially, it doesn't damage the market, while the things produced by an AI model can, but by that logic, is it a copyright violation for an instrument manufacturer to go and use a bunch of music to tailor an instrument to be better, if that instrument could be used to create music that competes with it? Again, no, but clearly there is a difference in how much that instrument would have drawn from the works. AI Models have the potential to spit out very similar works which makes them much more harmful to the original works' value. I think looking at it through the lens of copyright just isn't useful: it's not exactly the same thing, and the rules around copyright aren't good for managing it. Rather, we should be asking what we want from models and what they provide to society. As I see it, we should be asking how we can address the artists having their work fed into something that may reduce the value of their work, it's clearly a problem, and I don't think pushing the onus onto the person using the model not to create anything that infringes is a strategy that will actually work. I do think the author correctly calls out gatekeeping as a huge potential issue. I think a reasonable route is that models shouldn't be copyrightable/patentable themselves, companies should not be allowed to rent-seek on something largely based on other people's work, they should be inherently in the public domain like recipes. Of course, legislating something like that is hard at the best of times, and the current environment is hostile to passing anything, let alone something pro-consumer.
- re-thc 1y ago> they took a lot of music and used that to create algorithms that are effective for reproducing real-world music using less data That's redefining history. MP3 didn't evolve like this. There was a series of studies, experiments etc and it took many steps to get there. MP3 was not created by dumping music somewhere to get back an algorithm. > I think a reasonable route is that models shouldn't be copyrightable/patentable themselves Why? Why can't they just pay. AI companies have the highest valuation and so have the most $$$ and yet they can't pay? This is the equivalent of the rich claiming they are poor and then stealing from the poor.
- deleted 1y ago[deleted]
- Latty 1y ago> MP3 was not created by dumping music somewhere to get back an algorithm. This wasn't what I was trying to suggest, clearly I wasn't clear enough given the context, but my point was to give a very distant example where humans are using copyrighted works to test their algorithms as a starting point, as I later go on to say, I think the two cases are fundamentally different, but the point was to make the case that there are different types of "using copyrighted works to create tools", which is distinct from "learning". > Why? Why can't they just pay. AI companies have the highest valuation and so have the most $$$ and yet they can't pay? This is the equivalent of the rich claiming they are poor and then stealing from the poor. I don't think them paying solves the problem. 1) These are trained on such enormous amounts of data that is sourced unreliably, how are these companies going to negotiate with all of the rights holders? 2) How do you deal with the fact the original artists who previously sold rights to companies will now have their future work replaced in the market by these tools when they sold a specific work, not expecting that? Sure, the rights owners might make some money, but the artists end up getting nothing and suffering the impact of having their work devalued. 3) You then create a world where only giant megacorps who can afford to get the training rights can make models, they can then demand all work made with them (potentially necessary to compete in future markets) give them back the rights, creating a viscous cycle of rent-seeking where a few companies control the tools necessary to be a commercial artist. Paying might, at best, help satisfy current rights holders, which is a fraction of the problems at hand, in my opinion. I think making models inherently public domain solves far more of them.
- basch 1y ago"I think the unambiguous answer to this question is that the act of training is viewing and analysis, not copying. There is no particular copy of the work (or any copyrightable elements) stored in the model. While some models are capable of producing work similar to their inputs, this isn’t their intended function, and that ability is instead an effect of their general utility. Models use input work as the subject of analysis, but they only “keep” the understanding created, not the original work." The author just seems to have decided the answer and worked backwards. When in reality this is very much a ship of theseus type problem. At what point does a compressed jpeg not become the original image but a transformation? The same thing applies. If i ask a model to recite frankenstein and it largely does, is that not a lossy compression of the original. Would the author argue an mp3 isnt a copy of a song because all the information isnt there? Calling it "training" instead of compression lets the author play semantic games.
- Retr0id 1y agoThere clearly is a point when a compressed jpeg becomes a transformation, even if the precise point is ambiguous. Take 'The Bee Movie at 3000% speed except when they say "bee"', for example - https://www.youtube.com/watch?v=7apltfVJBwU https://www.youtube.com/watch?v=7apltfVJBwU. It hasn't been taken down for over 5 years so I'm going to assume it's considered acceptable/transformational use. Personally, I'd say what matters is whether you'd plausibly use the transformed/compressed version as a drop-in substitute for the original. ChatGPT can probably reproduce the complete works of shakespeare verbatim if prompted appropriately, but is anyone seriously going to read it that way?
- basch 1y agoAgreed. So not EVERY LLM is automatically copying in principle, but most of the current implementations probably retain TOO MUCH of the original sources to NOT be copies.
- kemotep 1y agoDo you know that YouTube hasn’t reassigned the video’s ad revenue to the IP holder and the IP holder hasn’t requested it be taken down because they now receive compensation for it from YouTube?
- blagie 1y agoI asked AI to complete an AGPL code file I wrote a decade ago. It did a pretty good job. What came out wasn't 100% identical, but clearly a paraphrased copy of my original. Even if we accept the house-of-cards of shaky arguments this essay is built on, even just for the sake of argument, where Open AI breaks my copyright is by having a computer "memorize" my work. That's a form of copy. If I've "learned" Harry Potter to the level where I can reproduce it verbatim, the reproduction would be a copyright violation. If I can paraphrase it, ditto. If I encode it in a different format (e.g. bits on magnetic media, or weights in a model), it still includes a duplicate. On the face of it, OpenAI, Hugging Face, Anthropic, Google, and all other companies are breaking copyright law as written. Usually, when reality and law diverge, law eventually shifts; not reality. Personally, I'm not a big fan of copyright law as written. We should have a discussion of what it should look like. That's a big discussion. I'll make a few claims: - We no longer need to encourage technological progress; it's moving fast enough. If anything, slowing it down makes sense. - "Fair use" is increasingly vague in an era where I can use AI to take your picture, tweak it, and reproduce an altered version in seconds - Transparency is increasingly important as technology defines the world around us. If the TikTok algorithm controls elections, and Google analyzes my data, it's important I know what those are. That's the bigger discussion to have.
- naming_the_user 1y agoCleanroom implementation comes to mind. If I just remember the source code of a 100 line program and then reproduce it verbatim a week later that doesn’t suddenly make it a new work.
- franktankbank 1y agoThis is why I don't believe in restrictive licensing of open work.
- HPsquared 1y agoMaybe the infringement occurs when a user uses the model to produce the facsimile output.
- EdwardDiego 1y agoThat's a lot of words to justify what I presume to be the author's pre-existing viewpoint. Given that "training" on someone else's IP will lead to a regurgitation of some slight permutation of that IP (e.g., all the Studio Ghibli style AI images), I think the author is pushing shit up hill with the word "can't".
- seanhunter 1y agoYup. Nothing quite like someone who clearly has no legal background trying to use first principles reasoning + bullshit to make a quasi-legal argument that justifies their own prior opinion.
- mdp2021 1y ago> Given that "training" on someone else's [would] lead to a regurgitation of some slight permutation That is not necessary. It may happen with "bad" NNs.
- TimorousBestie 1y agoThe assumption that human learning and “machine learning” are somehow equivalent (in a physical, ethical, or legal sense—the domain shifts throughout the essay) is not supported with evidence here. They spend a long time describing how machine learning is different from human learning on a computational level, but that doesn’t seem to impact the rest of the argument. I wish AI proponents would use the plain meaning of words in their persuasive arguments, instead of muddying the waters with anthropomorphic metaphors that smuggle in the conclusion.
- armoredkitten 1y agoExactly. In particular, when I train a model, I have a defined process for training, and I can flip the switch between "learning" and "not learning" to define exactly when the model adjusts its weights as a result of inputs. Humans can't do that with their brains. Thus, for humans, learning can't be decoupled from viewing, but it absolutely can be for AI.
- djoldman 1y agoThere are a few stages involved in delivering the output of a LLM or text-to-image model: 1. acquire training data 2. train on training data 3. run inference on trained model 4. deliver outputs of inference One can subdivide the above however one likes. My understanding is that most lawsuits are targeting 4. deliver outputs of inference. This is presumably because it has the best chance of resulting in a verdict favorable to the plaintiff. The issue of whether or not it's legal to train on training data to which one does not hold copyright is probably moot - businesses don't care too much about what you do unless you're making money off it.
- mdp2021 1y ago> businesses don't care too much Not really so, since the deranged application of an idea of "loss of revenue" decades ago.
- techpineapple 1y ago“If humans were somehow required to have an explicit license to learn from work, it would be the end of individual creativity as we know it“ What about text books, in order to train on a textbook, I have to pay a licensing fee.
- Ylpertnodi 1y ago>What about text books, in order to train on a textbook, I have to pay a licensing fee. Would that also apply if you bought the text books second-hand (or were given it)?
- Ekaros 1y agoFirst sale doctrine applies in those cases. That is original buyer can transfer the license that is sell the book. With AIs I think it is clear it would fall under some other limitations like you cannot broadcast CD over radio. And with movies not start movie theatre and play movies without paying creators...
- pitaj 1y agoIf you pirate a text book, learn from it, and then apply that knowledge to write your own textbook: your textbook would not be a copyright violation of the original, even though you "stole" the original.
- mdp2021 1y ago> I have to pay Fortunately, others have libraries. There is no need to pay for the examination of material stored in libraries (and similar).
- realharo 1y agoEven if you accept the premise, what does it matter? AI are not humans. Laws were made up by people at a specific time for a specific purpose. Obviously our existing laws are "designed" around human limitations. As for future laws, it's just a matter of who is powerful and persuasive enough to push through their vision of the future.
- yniopper 1y ago[dead]
- gavinhoward 1y agoCopyright reserves most rights to the author by default. And copyright laws thought about future changes. Copyright laws (in the US) added fair use, which has four tests. Not all of the tests need to fail for fair use to disappear. Usually two are enough. The one courts love the most is if the copy is used to create something commercial that competes with the original work. From near the top of the article: > I agree that the dynamic of corporations making for-profit tools using previously published material to directly compete with the original authors, especially when that work was published freely, is “bad.” So essentially, the author admits that AI fails this test. Thus, if authors can show the AI fails another test (and AI usually fails the substantive difference test), AI is copyright infringement. Period. The fact that the article gives up that point so early makes me feel I would be wasting time reading more, but I will still do it. Edit: still reading, but the author talks about enumerated rights. Most lawsuits target the distribution of model outputs because that is reproduction, an enumerated right. Edit 2: the author talks about sunstantive differences, admits they happen aboit 2% of the time, but then seems to argue that means they are not infringing at all. No, they are infringing in those instances. Edit 3: the author claims that model users are the infringing ones, but at least one AI company (Microsoft?) had agreed to indemnify users, so plaintiffs have full right to go after the company instead.
- light_hue_1 1y agoThis it totally the wrong analysis. Think of AI tools like any other tools. If I include code I'm not allowed to use, like reading a book I pirated, that's copyright infringement. If I include an image as an example in my image editor, that's ok if I am allowed to copy it. If someone decides to use my image editor to create an image that's copyrighted or trademarked, that's not the fault of the software. Even if my software says "hey look, here are some cool logos that you might want to draw inspiration from". People are getting too hung up on the AI part. That's irrelevant. This is just software. You need a license for the inputs and if the output is copyrighted that's on the user of the software. It's a significant risk of just using these models carelessly.
- Calwestjobs 1y agolook, quickest example if it IS or it IS NOT ip theft is - go to any image generation ML wizardry prompt machine and ask it this : "generate image of jack ryan investigating nuclear bomb. he has to look like morgan freeman." (and do it quickly before someone in FAANGM manually plays with something altering result of that prompt) problem is opposite, is "original" work IP a original in itself or is it just remix or someone just gave lawyer some generic text and make it arbitrarily protected for adding 0.000000001% to previous work.
- EPWN3D 1y agoI couldn't get through it, did he actually make an argument eventually?
- prophesi 1y agoI think it can be IP theft, and also require labor negotiations. And global technical infrastructure for people to opt-in to having their data trained on. And a method for creators to be compensated if they do opt-in and their work is ingested. And ways for their datasets to be audited by third parties. It sounds like a pipedream, but ethical enforcement of AI training across the globe will require multifaceted solutions that still won't stamp out all bad actors.
- tete 1y agoLike any kind of copyright law?
- fithisux 1y ago[flagged]
- mdp2021 1y agoThat only teaches us your opinion: too little information and too much.
- fithisux 1y agoIf you want to protect interests of the rich feel free, it is good if you are rich. I will not play your game.
- mdp2021 1y agoAgain: anybody can come and say "A is B". It is gratuitous. The argument is all that counts. You are speaking out of your inner representation and interpretaton of "A". We are those who are outside your mind.
- hyperman1 1y agoIn the EUCD, a copy in RAM falls under copyright, but there is an exception defined (art 5) if the copy is transitory and the target use is legal under copyright. Neither is true for AI, so this article is probably wrong in the EU. Apart from that, I wonder uf an AI is learning in the legal sense of the word. I'd suspect removing copyright trough learning is something only humans can do, seen trough legal glasses. An AI would be a mechanical device creating a mashup of multiple works, and be a derived work of all of them. Main problem with this rebuttal is how you prove the AI copied your work specifically, and finding out which of the zillions of creative works in that mashup are owned by who.
- deleted 1y ago[deleted]
- alganet 1y agoThat's a lot of text. Where is AI disruptive? If it is disruptive in some area, should we apply old precedents to a thing so radically new? (rethorical). Good fresh training data _will end_. The entire world can't feed this machine as fast as it "learns". To make a farming comparison, it's eating the seeds. Any new content gets devoured before it has a chance to grow and give fruit. Furthermore, people are starting to manipulate the model instead of just creating good content. What exactly will we learn then? No one fucking knows. It's a power grab free for all waiting to happen. Whoever is poor in compute resources will lose (people! the majority of us). If I am right, we will start seeing anemic LLMs soon. They will get worse with more training, not better. Of course they will still be useful, but not as a liberating learning tool. Let's hope I am not right.
- bionhoward 1y agoDid the article mention the part about how these companies turn around and say you’re not allowed to use the output to develop competitive models? I couldn’t find mention of this
- hulitu 1y ago> Why training AI can't be IP theft Because Microsoft is part of BSA. /s If you steal our software, it is theft. If we still your software, it is fair use. Can we train AI on leaked Windows source code ?
- ConspiracyFact 1y agoThe problem is that model outputs are wholly derivative. This is easy to see if you start with a dataset of one artistic work and add additional works one at a time. Clearly, at the start the outputs are derivative. As more inputs are added, there’s no magical transformation from derivative to non-derivative at any particular point. The output is always a deterministic function of the inputs, or a deterministic output papered over with randomness. “But,” you say, “human art is derivative too in that case!” No. A human artist is influenced by other artists, yes, but he is also influenced by the totality of his life experience, which amounts to much more in terms of “inputs”.