4 ms·
I feel that all of these deals are an implicit acknowledgement that what OpenAI and all of these companies did training on everyone's' content and data was ille
by BadHumans 2y ago
I feel that all of these deals are an implicit acknowledgement that what OpenAI and all of these companies did training on everyone's' content and data was illegal. If you're so certain that what you did was fine and fair use then why go through the trouble of spending all this money on licensing deals? Unless they just think that dealing with all of the lawsuits would cost more than just paying people to go away.
- drooby 2y agoPerhaps because that is still being disputed in court? The outcome is unknown and they are making a defensive play? If it turns out they shouldn't have done that, then when the dust settles they will be ahead of the competitors that didn't sign deals.
- KoolKat23 2y agoNo it's a simple net present cost calculation. Its cheaper than continued litigation and ensures they maintain access in future (rather than the publisher in question starts blocking their web crawlers). Once they've got deals with the few really big players, the rest of the industry falls into line (the smaller guys don't have the financial means to out-litigate or block OpenAI).
- jstummbillig 2y agoThe difference between "ack it was illegal" and "certain it was fine" and how easily you skip over it. Trying to make things better is good and requires no admittance of guilt.
- dylan604 2y agoI've thought of it as more of a they didn't have money to license at first, then the models got decent enough for them to charge for use, now they are trying to avoid getting the models shut down for violations by making deals. Too bad only the large players that got scraped will have these deals made while the smaller players are left without
- HeavenFox 2y agoThe content directly obtained from the source would also likely be cleaner, and even if it's settled that scraping and training is legal, the cost savings in data cleaning could still warrant a paid partnership.
- stranded22 2y agoSurely it is about not allowing others to have access to the data…
- dylan604 2y agoAre you suggesting that OpenAI made a deal that will mean Googs/MS/others will now not be able to license the same content?
- tourmalinetaco 2y agoMicrosoft already has deals with OpenAI and owns a sizable stock, they’re not looking to make ML models.
- smegger001 2y agoalso Microsoft owns the number 2 search engine no one is blocking Microsoft from scraping their sites else they loose all that traffic from Bing searches. I suspect the same with google thus googles deals are strictly speaking an anti competitive measure and insurance against the law deciding against them
- MangoCoffee 2y agoThese publishers are lucky to get a deal from Western AI companies. I doubt Chinese AI companies will give a f'ck. In the age of China-US AI competition, this will be one advantage for Chinese AI companies.
- c1sc0 2y agoArguably what they win in input training cost they may lose in output censorship cost? But then again the line between Western “Guardrails” and Eastern “Censorship” isn’t all that clear.
- nickthegreek 2y agoIt’s just the next TikTok/Douyin. One version for their country, another for the rest of the world.
- vineyardmike 2y agoI don’t know I think it’s pretty clear… There was a massacre in Tiananmen Square. America has also committed massacres, like over throwing governments of foreign nations. No one in my family is at risk from either of those statements. And there is no automated system to stop them. That’s the difference. Just because people decided that massive platforms should limit hate-speech doesn’t mean the west is performing censorship of a comparable level. Not even close.
- lancesells 2y agoThese publishers are really being stupid IMO. Just speeding up their demise on whatever very short term profits they have by making these deals.
- bloppe 2y agoNobody that's familiar with copyright precedent seriously thinks training on copyrighted data is illegal. What's illegal is simply reproducing the copyrighted content for others. See Google Books for an interesting precedent. Google still scans every book ever written without paying licensing fees. They just can't legally make all those scans freely available to all. Clearly, some of the responses GPT gives to users have infringed training data copyright, but the majority of their responses do not. They basically have 3 options: 1. Figure out how to engineer an LLM so that it reliably avoids "unfair" use of copyrighted content. "Fair use" is a legal doctrine with no rigorous definition, so this would be very difficult even if they had a clue where to start. I wouldn't hold my breath for this. 2. They can continue without any licensing, and field copyright lawsuits on a case-by-case basis for each individual prompt and response. That would be a logistical nightmare for the courts, plaintiffs, and defendant alike. It would certainly stress test the whole system, possibly result in knee jerk legislation that OAI may not like, or simply bury them in legal fees if infringement is common enough (which is not entirely clear yet). 3. They can strike deals, eliminating legal uncertainty and allowing them to plow ahead with reproducing copyrighted content without worries, while also getting other goodies like exclusivity deals at the same time. Seems like a no brainer to me
- fredgrott 2y agoIts a little more nuance that as otherwise archive or library system would still be in business..
- bloppe 2y agoLibraries and archives rely on first sale doctrine, which doesn't apply to digital copies, only physical ones.
- hiatus 2y agoTo add a bit of color, first sale doctrine does not apply to licensing, rentals, etc, which is why it won't apply to most digital sales (though it does not apply to _all_ as you can still outright buy some digital media beyond mere licensure).
- ben_w 2y agoHypothetically, because the deal is cheaper than fighting the case. (In practice, being not a lawyer, I have no idea how even Google's search indexing is legal, nor where the boundaries are between the legal bit vs. the times they got in trouble for indexing newspapers and at least one separate case about images).
- deleted 2y ago[deleted]
- 2OEH8eoCRo0 2y agoConde Nast is massive with a massive back catalog of material that is not online or is behind paywall.
- smegger001 2y agobecasue after they did it everyone put roadblocks in place to prevent anyone from doing it again so even if you believe its fair use (I happen to think it is clearly a transformative work thus clear on copyright grounds) you still have to acknowledge the practical issue that its simply easier to pay for access to newly data. Also having new data of known provenance to prevent feeding ai generated data to the ai which is known to cause reduction in quality of output, as well as preventing lawsuits which even if your are in the right are expensive and time consuming at give you a bad look in the public mind, paying right now just makes since on to many levels not to.