4 ms·
> 100 GB of weights don't resemble any work I don't think that argument has a lot of weight. By the same logic a zip file of a book doesn't resemble the origi
by ajnin 3y ago
> 100 GB of weights don't resemble any work
I don't think that argument has a lot of weight. By the same logic a zip file of a book doesn't resemble the original work if you look at the raw bytes but you can extract it and get it back. It would be hard to argue that the copyright violation happens only when you extract the zip, and not, say, if you distribute the archive.
- panta 3y agoLLM could even be seen as a form of lossy compression after all
- datadrivenangel 3y agoA sufficiently overfit model is indistinguishable from bad compression.
- Barrin92 3y agoThe same logic doesn't apply because you can't get the data back out of the model by unpacking it. It's theoretically not possible because the model is magnitudes smaller than the totality of the data. To use, and invert the example from the other commenter, you can even get an existing poem out of a generative model that genuinely was not in the training data at all. This is because it's not an archive, which does correspond directly to one particular work, but a generative technology.
- danaris 3y agoIf I create an original file in Photoshop that's 8K, then produce a JPG of it, the JPG is both orders of magnitude smaller, and clearly a rendition of the same work. There's no way to get back to the original 587MB 8K PSD from the 87KB 1024x768 JPG, but that's irrelevant to whether the latter is a version of the former. Just because the models are not perceptually the same as their training data does not mean that they are not effectively a lossy compression of that data.
- Marazan 3y agoBut people absolutely have extracted individual source images out of the models by not even particularly careful prompting.
- artifabrian 3y agoThat's absolutely not true. You might get replicas of incredibly popular works like Mona Lisa due to overfitting but that's it. If I am wrong please do provide an example as this is very relevant and interesting (and impossible from an information theory pov).
- Marazan 3y agohttps://www.technologyreview.com/2023/02/03/1067786/ai-models-spit-out-photos-of-real-people-and-copyrighted-images/ https://www.technologyreview.com/2023/02/03/1067786/ai-model... The approach I've seen is to prompt for people with unusual names, they're often be only a single source image in the input data set that gets reproduced by the AI. I've seen examples with the AI "generated" images and the source image side by side - I'll try and find them. EDIT: Link to the paper https://arxiv.org/abs/2301.13188 https://arxiv.org/abs/2301.13188 and this articles has the example images: https://www.theregister.com/2023/02/06/uh_oh_attackers_can_extract/ https://www.theregister.com/2023/02/06/uh_oh_attackers_can_e... Look at the Ann Graham Lotz image and tell me that isn't the source image being reproduced in a lossy manner.
- SilasX 3y agoRight, the LLM is more like the zip algorithm than the zip file. Yes, if you feed certain bits (the zip file) to the zip algorithm, you can get a copyrighted work out, but that means the zip file is violating, not the zip software.
- otabdeveloper4 3y agoSo if I lossily compress a novel so random typos and spelling errors are introduced then it isn't copyright infringement? Wow, that changes everything. BRB, I got books to pirate.