3 ms·
> reproduce books word by word, page by page This statement is a figment of the commenters imagination with no basis in reality. All they would have to do is t
by menzoic 2y ago
> reproduce books word by word, page by page
This statement is a figment of the commenters imagination with no basis in reality. All they would have to do is try it to realize they just spouted a lie.
At most LLMs can produce partial excerpts.
LLMs don’t store the data that it’s trained on. That would be infeasible, the models would be too large. Instead, it stores semantic representations which often uses entirely different words and sentence structures than the source content. And of course most of the data is lost entirely during this lossy compression.
- rhubarbtree 2y agoThe NYT has extracted long articles from ChatGPT and submitted the evidence in court.
- ben_w 2y agoCrucial size difference between an article and a book. Size difference meaning that people often share complete copies of articles to get around pay walls — including here. As I understand it, this is already copyright infringement. I suspect that those copies are how and why it's possible in cases such as NYT.
- Lerc 2y agoGiven that it has been submitted in court, does that mean you can say what the longest verbatim extract was? It seems like that would be a fact that couldn't be argued with.
- angoragoats 2y ago> At most LLMs can produce partial excerpts. Glad you agree that LLMs infringe copyrights.