4 ms·
>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will e
by DowsingSpoon 3y ago
>you can train an AI on copyrighted material just like people can learn from copyrighted material.
No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide.
A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It therefore violates the copyrights of a large number of rights holders. The outputs of the model are derivative works which also violate copyright.
And anyone using or training a model trained on works for which they do not have the rights? Completely fucked. Or at least, they must accept this as a real risk.
- deleted 3y ago[deleted]
- kelnos 3y agoI think a reasonable interpretation is also that what you are saying is correct, that doing all that does indeed infringe others' copyright, but that a fair use defense is valid. I won't be particularly thrilled if that turns out to be the case, but I wouldn't be surprised if it does. But as you say, we won't know until it's tested in court. And even then, often court cases around a complex topic like this will end up with a ruling that only clarifies a narrow aspect of it. So it might take many related court cases before we have a pretty good understanding of where the law stands. And then, of course, the law could change.
- AnthonyMouse 3y agoDerivative works are typically things like language translations or film adaptions of a novel. A large language model is something like a probability breakdown of the order of word fragments in a body of text. It's a collection of statistics and math. It's different. Now, can you get it to output a derivative work? Maybe. Is every output a derivative work? Maybe not.
- DowsingSpoon 3y agoThings like sports statistics and directories of phone numbers are not copyrightable. Maybe models fall into a similar category. Could be. It’ll be interesting to see how this shakes out over the next few years.
- buildbot 3y ago> A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It therefore violates the copyrights of a large number of rights holders. The outputs of the model are derivative works which also violate copyright. What is the blackbox “limit” here? Is the mean value of all images in imagenet (which contains many copyrighted images) violating copyright? Is the character count of sarah silverman’s books? What about a prime number representing them - https://en.wikipedia.org/wiki/Illegal_number?wprov=sfti1 https://en.wikipedia.org/wiki/Illegal_number?wprov=sfti1 Training is much more similar to a character count than an illegal prime in my view, and thus, is almost certainly going to be okay/found to be okay. If not, something like, 90% of all models used today had some component trained on copyrighted data of some form.
- gaganyaan 3y ago[flagged]
- lotsoweiners 3y agoYou sound like my coworker 5-10 years ago that told me how I wouldn’t be driving today because of the proliferation of self driving vehicles. I told him he’s a 28 year old dum dum who didn’t understand how things operate in the real world when the constraints aren’t based on technology but on government regulations, the economy, and other factors. I’d say I won that argument for now. I live in the Waymo pilot city and still haven’t taken one mostly due to the limited area they drive in. Just this past week we learned that a traffic can disable a self driving vehicle. I’m interested to see the traffic cone era of AI.
- gaganyaan 3y agoI'm basing this on what US courts have decided. You're free to disagree with them all you'd like, but AI-generated art is not copyrightable. We're seeing an explosion in non-copyrightable art, and when we get down to some small fraction of art being copyrightable, nobody will give a shit about copyright anymore. You also talk about the economy, and guess where the economoc incentives are aligned towards? Hint: it's not towards having expensive humans generate art. I guess we'll just have to feel that the other is wrong as we wait to see what happens, but I'm betting on existing US regulation and basic human behavior. That's a tall order to bet against.
- jazzyjackson 3y agojeez, way to make it personal I agree that copyright is done for in an age of generative models, and as a pirate I'm kind of rooting for it, but i'm not so sure it's unequivocally a Good Thing. I'm interested to understand history better, how art and science was produced and distributed before the legal fiction of intellectual property. the point of allowing someone a monopoly on their work is to share it with the public, same as patents. without the legal framework, the way to protect your work may be to not publish it at all, which is where I see the internet going from here, private enclaves that go to great lengths to prevent LLMs from drinking their milkshake.
- dragonwriter 3y ago> A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. Its unmistakably not a derivative work of the inputs individually or collectively, since a derivative work must be itself an distinct work of authorship (the same as the work of authorship requirement for copyright), and the output of a purely mechanical process is not. The collection of inputs itself might be a derivative work of the individual inputs, before considering Fair Use.
- dang 3y ago> No amount of whining and hand wringing from engineers will ever make this true Please omit flamebait and swipes, as the site guidelines ask: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html. Your comment would have been fine without that bit.