4 ms·
> You basically don't know what was there and you cannot prove they didn't achieve those results with perfectly clean data. I mean you totally do though, right
by Arelius 4y ago
> You basically don't know what was there and you cannot prove they didn't achieve those results with perfectly clean data.
I mean you totally do though, right? You just need one instance of the LLM reproducing information that would only have been able to by violating copyright.
I mean, it's theoretically possible that it could have reproduced it from scratch, infinite monkies on typewriters sort of thing, but statistically we can rule that out on pretty short notice.
Adding on to this, I don't think the argument that OpenAI, Google and others are ultimately making will be that they don't violate copyright, but instead will ultimately be that their violation is sufficiently transformative such that it constitutes fair-use.
- muyuu 4y agonot only it's theoretically possible, it happens and it can already be observed on clean lab experiments with normally used parameters the probability that LLMs produce copyrighted information is no proof that it was trained with it exactly, esp. when parameters are set so they don't repeat outputs
- Arelius 4y agoI think you'll find that it is in fact proof by all practical standards we use outside of formal mathematics.
- muyuu 3y agothe moment you cannot in any practical way tell if the data set was corrupted with copyrighted material, nobody will convict you for any accidental violations that may occur, even in the astronomically low probability that they do with standard parameters
- Arelius 3y agoMy point is being able to reliably reproduce copyright works will function as a very practical way to tell if the dataset was corrupted with copyrighted material. In that way it’ll be a lot easier to prove that a dataset was corrupted, then proving the negative.