5 ms·
If I compress an artist's painting into a jpeg and rehost part of it for individual t-shirt designs I am committing a crime. If I compress an artist's painting
by greentext 3y ago
If I compress an artist's painting into a jpeg and rehost part of it for individual t-shirt designs I am committing a crime.
If I compress an artist's painting into a model & rehost what's essentially a highly flexible complete version of their painting for infinite, perpetual use of any kind ... I'm not committing a crime?
- hexage1814 3y ago>compress an artist's painting into a model That's not how image models work.
- greentext 3y agoPainting features => back propagation => weights. Yes it is.
- Loocid 3y agoIt has been shown that image models can produce originals, or at least extremely close to the originals. If the outcome is the same, what is the difference between compression/decompression vs training/generation regarding copyright?
- sdiupIGPWEfh 3y ago> It has been shown that image models can produce originals Not in the general case, no. For the study done against Stable Diffusion [1], researchers were only able to reproduce about 0.03 percent of the images tested. Those were also believed to be cases of overfitting on images which were over-represented in the training data and they're not something you'd hit upon by accident. Generative text models seen to be more problematic, depending on the subject. Code seems especially prone to overfitting, probably due to insufficient amounts of it compared to other text sources as well as lots of copying going on between the repos the models were trained on. [1](https://arstechnica.com/information-technology/2023/02/researchers-extract-training-images-from-stable-diffusion-but-its-difficult/ https://arstechnica.com/information-technology/2023/02/resea...)
- icehawk 3y agoThey got 94 direct matches, which is 94 instances where copyright infringement could be argued.
- sdiupIGPWEfh 3y agoCould be argued, sure. If you have to already have access to the copyrighted images to find them in the model, the argument seems weak. A sufficiently advanced model could, in theory, generate any image. You could then, again in theory, find an embedding for any image. Does said model then infringe on all copyrighted images? A program that creates Fourier epicycle drawings could be given input that causes trademarked output. An evolutionary algorithm iterating on noise could, given metrics for an image and the right fitness function, generate infringing images. Hypothetically, and admittedly absurdly, you could extract any image in the binary expansion of Pi and share it by "just" providing an index and length. If you have to know exactly what you're looking for and have to perform a substantial amount of computation to get it, it could be argued that the act of infringement is in the effort made by the person seeking infringing content (and distributing the results) rather than whatever it is they're attempting to extract the content from. But hey, courts don't always make sensible rulings, so who knows.
- a13o 3y agoLLM training is a compression algorithm, LLM weights are a compressed dataset of the training materials, and executing LLMs is accessing the compressed source material. That the compression technology relies on parameterizing the copyrighted material, and as a result can produce hallucinations remixing the copyrighted material, is super cool but doesn't change that this is compression at rest. All the hullabaloo about artificial intelligence is science fiction laundering a (cool new) compression algorithm. Copyright law shouldn't apply any differently to a LLM as it does to gzip.
- epups 3y agoCan I talk to gzip in natural language and have it produce novel output not contained within any of its source files? If not, I think your comparison is deeply flawed. LLM's are not simply compression algorithms.
- a13o 3y agoMaybe if someone were to build it. You can't talk to LLMs in natural language either. They have a very precise query language. The natural language component is an additional feature bolted onto the front. Also the output of LLMs is not novel, it's a derivative work of the training dataset. LLMs can't produce anything not present in the training set. They are operating on parameterized classifications of artwork instead of the artwork itself, which is the new technology; but they can't do anything except combine those existing Legos in new ways. We also talk to databases in what was considered 'natural language' for that era - SQL. The LLMs are no different, which is why we have prompt engineering same as we have database engineering. Just because the database is processing the dataset and storing it in a proprietary way, approachable via query language, does not change the fact that the original data is still in there, and everything that comes out is derivative.
- m463 3y agoWhat I wonder is: Load the source code to all versions of unix, with all licenses. "write me a version of unix" Since there is no model copyright and the result was written by AI the software is now in the public domain.
- Zuiii 3y ago> Since there is no model copyright and the result was written by AI the software is now in the public domain. No that's not how copyright works. If you have photographic memory and reproduce a work exactly you still commit copyright infringement.