4 ms·
> If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringe
by jsmith45 3y ago
> If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringement.
Sure. Slightly more interesting is if that same human with those same breifcases was taking money to answer questions and referenced those papers, but did not just provide the article or headlines, and might not even be paraphrasing the article at all. Is that okay?
To the extent it is merely paraphrasing articles, or outputting headlines that it just looked up, I agree that could well be infringement. If it more transformative processes those articles into something distinct, then it is not nearly as clear cut. The latter is arguably the intent of openAI, even if the current results might be closer top the former.
- __e 3y agoWhile LLMs are designed to generalize from their training data, rather than simply memorizing it, overfitting occurs in niche areas. There's be plenty of niche areas and LLMs of today merely repeat training data, like with the NYT case. Larger datasets and better algorithms will help to an extent, but you'll always have niche topics where overfitting happens. The intent has always been to generalize well, however, it may not be feasible to do so in the long tail of the internet. How should copyright law address this?