4 ms·
This battle will be fascinating.
by Biologist123 3y ago
This battle will be fascinating.
- thecleaner 3y agoNo it wont lead to anywhere. They will just get rid of the NYT bits of the corpus, then settle or just flat out not care if they lose. Transformers are relatively easy to train they arent GANs, the bottleneck is hardware and OpenAI has those bits sorted out.
- theonemind 3y agoPresumably, the outcome will start the series of legal precedents regarding copyright and GPT (and other AI stuff) training potentially going far beyond OpenAI and NYT.
- nicklecompte 3y agoIn the lawsuit, NYT pointed out that their articles are the third-most used source in GPT-3, behind Wikipedia and the US patent office. We don't know what they used for GPT-4 but I imagine NYT is still very prominent. High-quality newspaper articles are valuable for training LLMs because they contain good prose and an enormous variety of factual information. They can't be replaced by throwing in a bunch of Reddit comments and hoping for the best. "NYT bits of the corpus" is understating things - I imagine getting rid of NYT articles would significantly impact the overall quality of GPT-4's responses in many unexpected domains.
- vidarh 3y agoI very much doubt it'd be a problem. OpenAI's valuation is such a large multiple of NYT that they can afford to buy out a dozen or so publishers of NYT's valuation to get content to replace (or "just" license content from far more) if NYT itself isn't willing to play ball cheaply enough. The NYT might be better than average, but it's not so special that you can't replace it. The long-term effect if NYT wins this case will be that the "established" players and companies with deep pockets will get a massive moat protecting them against open models and new entrants - in the long run it might well be well worth whatever a loss will cost them.
- Biologist123 3y agoThe US is placing a big bet on AI to break the economic secular stagnation of recent decades. If IP is an obstacle to that, it would be great if IP laws are loosened.
- NemoNobody 3y agoThis is what I expect/hope will happen. The NYT could prolly win under current law - due to the interpretation of how GPT is reproducing the articles. The judge that may ultimately decide the future of how AI will fall for now in the US, will likely know very little about AI/LLMs. I don't think reproducing a current article is that big deal tbh - it just retrieved the info from itself rather than the NYT bc it had already saved the info before and can't forget things. That's humanely possible by anyone with a photographic memory - it's more similar to recollection than infringement. The NYT is pissed the LLMs learned from their data - you can't stop something from learning - especially after the fact. Obviously, the user intent not machine intent is key also. No infringement can happen without a prompt that suggests copyrighted material - easy to do or not, requires a person nonetheless. The technical aspects of this is where they have a chance, taking a step back, it all just looks kind of silly.
- NemoNobody 3y agoSeriously, just consider this: I memorize a textbook, verbatim, about WWII and you have a term paper on WWII so you ask me to write some notes about the highlights, so I write notes that I take write from the verbatim text in my head and give them to you. You use the notes I gave you verbatim in your paper and are found guilty of plagiarism. You obviously didn't realize the notes I gave you on the spot were from a textbook - that wasn't what you were looking for. I didn't realize you were just going to use whatever I wrote in your paper. Did I commit plagiarism or did you?
- LargeTomato 3y agoIf the NYT doesn't settle this could set precedent that will have massive repercussions on how all AI companies collect data. This could end up being a hugely important case.
- BobaFloutist 3y agoEven if they do, settlements often establish de facto precedents that corporate lawyers (and insurance companies) will strongly advise you steer away from.