4 ms·
Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation
by zaptrem 3mo ago
Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).
- andriy_koval 3mo agoUs models didnt pay for licenses too
- joe_mamba 3mo agoWe're still in the early days of the AI industry timeline(relative to traditional industries). Not everything has yet been litigated. Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature. If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pay extra taxes on any and all storage mediums and on devices with built-in storage (tapes, CDs, DVDs, HDDs, SSDs, tablets, phones, etc) simply because they can be used to store pirated content, decisions based on laws from 50-100 years ago, and the money goes to the national unions and associations of music and arts IP holders. It's basically a lobby pushed and government legalized extortion racket that no voter agrees with or can change but has no choice but to conform either way. So I guarantee you in the future, it will be the same for AI subscriptions and hardware capable of running LLMs locally. Every time you purchase a Claude or ChatGPT subscription, an Nvidia GPU, Intel/AMD SoC PC or an Apple/Qualcomm powered smartphone, you'll pay a government enforced tax to the likes of Sony, Axel Springer, etc. for licensing their IP, whether you want to or not. In the EU at least. US maybe not.
- andriy_koval 3mo agoI think we are going to direction where AI corps will have stronger lobby compared to IP holders.
- tristanj 3mo agoThat is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt the Chinese models operate under similar licensing agreements.
- maximus_01 3mo agoLet’s not forget that Anthropic only paid that to settle a class action lawsuit.
- andriy_koval 3mo agothey didn't pay yet, because court challenged settlement as inadequate. > I doubt the Chinese models operate under similar licensing agreements. US corps likely pay licenses when afraid to be sued, or have troubles getting that data, otherwise they just take data, which was demonstrated many times. The same apply to Chinese corps, alibaba totally can be sued in US.
- tristanj 3mo agoChina is infamous for weakly enforcing copyright law. Even when it is completely obvious that Chinese labs are training models on pirated data, US copyright holders face a virtually impossible task of proving it in court. Those lawsuits won't go anywhere.
- andriy_koval 3mo agoThere are tons of lawsuites which resulted in banning Chinese companies from doing business in US, those lawsuits totally have consequences.
- no-name-here 3mo agoWhat are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?
- deleted 3mo ago[deleted]