4 ms·
The big question is: if copyrighted material was used in the training material, is the LLM's output copyright infringement when it resembles the training materi
by peacebeard 7mo ago
The big question is: if copyrighted material was used in the training material, is the LLM's output copyright infringement when it resembles the training material? In your example, you are taking the copyrighted material and giving it to the LLM as input and instructing the LLM to process it. Regardless of where the legal cards fall, this is a much less ambiguous scenario.
- LPisGood 7mo agoI think Disney ran into this with people generating Marvel characters etc
- shagie 7mo agoFictional characters can have their own copyright. https://www.nolo.com/legal-encyclopedia/protecting-fictional-characters-under-copyright-law.html https://www.nolo.com/legal-encyclopedia/protecting-fictional... https://en.wikipedia.org/wiki/Copyright_protection_for_fictional_characters https://en.wikipedia.org/wiki/Copyright_protection_for_ficti...
- wvenable 7mo agoThere's a couple of different issues here that all get mangled together. If you're producing effectively the same expression that's infringement. You draw Captain America from memory, it's still Captain America, and therefore infringement. If you draw Captain Canada by tracing around Captain America that's also infringement but of a different type. When it comes to software, again it's the expression that matters -- literally the actual source code. Software that does the same thing but uses entirely different code to do it is not the same expression. Like with the tracing example above, if you read the original source code then it's harder to claim that it isn't the same expression. This is why clean room implementations are necessary.
- alex1sa 7mo ago[flagged]
- wvenable 7mo agoClean room is merely a defence in case you get sued by someone saying that you copied the work. It's not legally necessary. If the presumption is that LLM training, despite reading all the source code of everything everywhere, ultimately doesn't actually contain that source code (in a compressed form) then that is the significant bit. If training is truly doing something transformative, maybe even a machine analogy to human learning, then anything produced directly by that LLM without another work in it's context is an entirely new work. That's all that is important. > I think the practical answer is that clean room as a legal concept was designed for a world where reimplementation was expensive and intentional. Whether or not it's expensive or intentional is immaterial. It always was and it's still true now. All that matters is that the actual expression, the real source code, is not copied. Clean room is just one way to have evidence that you didn't copy.