5 ms·
it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work
by metalspot 3y ago
it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them.
the process of training requires reproduction and distribution of the works internally as part of the data processing pipeline so why wouldn't you need a license for that?
- subw00f 3y agoIf I publish an article on a subject I extensively read about on books and add no new information, I’m just reproducing them. Should that be considered a violation of copyright?
- bathtub365 3y agoThe analogy would hold if a single human could read all books and then be copied an infinite number of times with little effort and respond to an infinite number of prompts simultaneously. This is fundamentally different than the example of a single person reading books and being inspired by them to produce something. The scale is important and I see this lost in a lot of discussions.
- akasakahakada 3y agoSimply no. Logic must be consistent from 1 to infinity. If someone read and speak too fast is illegal, then let's do it. Put everyone has IQ > 130 into jail. Oh Rainman cannot forget things, obviously that is illegal, shot it down.
- metalspot 3y agofalse analogy. if openai is downloading and copying and distributing copyrighted material internally and then compressing that information into a model that can reproduce it later and selling access to that model that is a very different thing.
- akasakahakada 3y agoWhat if I just make a robot to go to public library and book store and then OCR everything, would it make a difference? Again there is no "compression of information" in deep learning.
- metalspot 3y ago> there is no "compression of information" in deep learning i would strongly disagree. when you are training a model you are taking the information from a document and extracting the relationships between tokens and storing that information conglomerated with the same information from a massive amount of other documents. the model that results is a compressed form of all of the information from all of the documents where you have extracted and stored a synthesis of the relationships between the tokens in all of them. this is a lossy compression, but it does reproduce exact sequences of source documents in some cases, so the original information is stored there. you can very plausibly argue that an LLM model trained on copyrighted material violates the copyright on every single copyrighted document that was fed to it.
- rockemsockem 3y agoDistributing data to servers only accessible by a machine isn't what folks have in mind when they talk about the distribution of copyrighted content.
- metalspot 3y ago> isn't what folks have in mind you would have to ask a judge about that because its a novel legal question. obviously nobody anticipated this technology at the time it was written so ultimately it will have to be a court that decides how to apply existing laws.
- melagonster 3y agoI trust this just because they trust they can escape from this.