4 ms·
Model weights, if they can reproduce something like the original, are just a form of lossy compression (or even lossless for text), where the LLM answering the
by devit 2y ago
Model weights, if they can reproduce something like the original, are just a form of lossy compression (or even lossless for text), where the LLM answering the prompt is a more powerful version of asking software to retrieve a specific file from a Zip archive (or a webserver answering an HTTP query) of such lossy compressed data.
So if model weights don't infringe, that would also imply that saving an image as a JPG or a video using AV-1 doesn't infringe, which would obviously effectively implies that copyright doesn't apply to images or videos on the web, which is not current law/policy, so I think that reasoning cannot possibly work.
- fenomas 2y agoThat comparison would only make sense if compressed images were considered derivative works. They're not - copyright doesn't protect bytes on a disk, it protects creative expressions. Lossy compression doesn't affect the creative expression, so in copyright terms a compressed JPG is just a copy, and is covered exactly like the original image. In contrast a derivative work is one creative expression that contains elements of another - like when you take an image and add commentary, or draw your own addition onto it, etc. And I'm pointing out that a trained model is not that - it's not itself a copyrightable expressive work. (We could think of it as a kind of algorithm for generating works, but algorithms aren't copyrightable.)
- devit 2y agoWell then the model weights would be a compilation of copies of the original works, which has the same effect as it being a derivative work unless the copyright holder chose to allow copies but not derivative works.