3 ms·
The OP was asking a different question with a different answer -- OP wanted to know if novel ChatGPT output is copyrightable, not whether existing copyrights mi
by thwayunion 4y ago
The OP was asking a different question with a different answer -- OP wanted to know if novel ChatGPT output is copyrightable, not whether existing copyrights might cover ChatGPT outputs. That said...
> The compression ratio of the source corpus versus the model demonstrates it doesn't have all the source material, ipso facto, results are not infringement.
What?! This is not at all how copyright law works and that's also not at all how LLMs or compression work! You're wrong about both the law and the technology!
Just because the model can't losslessly reproduce the entire training dataset doesn't mean that the model can't losslessly reproduce copyrightable fragments of the training dataset.
And the law truly doesn't give a damn that a verbatim copy of a piece of code or sequence of paragraphs came out of a model instead of a copy/paste from github. If it's the same basket of bits, or close enough in certain cases, it's covered by copyright, full stop.
> As an easy to understand analog, if you grab a Shutterstock photo, change it from 100% quality to 5% quality, the file size drops to 100th of the original and on display, none of the pixels are the same. While it may be the same general concept depicted badly, it is certainly not the original and clearly isn't copyrightable.
As an easier to understand analog, if you grab a BITMAP of a Shutterstock photo and convert the file type to a reasonable JPEG (lossy compression!!!), tht's very clearly still covered by the copyright on the .bmp.
- Terretta 4y agoYou seem to have missed "Note: That last paragraph is firmly tongue in cheek /s" Firmly tongue in cheek sarcasm, as in, this is obviously wrong.