4 ms·
Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every o
by boondongle 2mo ago
Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape."
Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether it's complete loss of access, or the amount of control you'll have to give up to access them will be ridiculous.
Sort of like the "stealing music is fine" but "lets freak out now that it's producing visual art", in the end the entire thing is a social construct. Whether this is treated as theft or "business as usual" is entirely societal.
Eventually the gap will close, unless there's a major breakthrough that hasn't been made yet.
- zaptrem 2mo agoGiven these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).
- andriy_koval 2mo agoUs models didnt pay for licenses too
- joe_mamba 2mo agoWe're still in the early days of the AI industry timeline(relative to traditional industries). Not everything has yet been litigated. Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature. If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pay extra taxes on any and all storage mediums and on devices with built-in storage (tapes, CDs, DVDs, HDDs, SSDs, tablets, phones, etc) simply because they can be used to store pirated content, decisions based on laws from 50-100 years ago, and the money goes to the national unions and associations of music and arts IP holders. It's basically a lobby pushed and government legalized extortion racket that no voter agrees with or can change but has no choice but to conform either way. So I guarantee you in the future, it will be the same for AI subscriptions and hardware capable of running LLMs locally. Every time you purchase a Claude or ChatGPT subscription, an Nvidia GPU, Intel/AMD SoC PC or an Apple/Qualcomm powered smartphone, you'll pay a government enforced tax to the likes of Sony, Axel Springer, etc. for licensing their IP, whether you want to or not. In the EU at least. US maybe not.
- andriy_koval 2mo agoI think we are going to direction where AI corps will have stronger lobby compared to IP holders.
- tristanj 2mo agoThat is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt the Chinese models operate under similar licensing agreements.
- maximus_01 2mo agoLet’s not forget that Anthropic only paid that to settle a class action lawsuit.
- andriy_koval 2mo agothey didn't pay yet, because court challenged settlement as inadequate. > I doubt the Chinese models operate under similar licensing agreements. US corps likely pay licenses when afraid to be sued, or have troubles getting that data, otherwise they just take data, which was demonstrated many times. The same apply to Chinese corps, alibaba totally can be sued in US.
- tristanj 2mo agoChina is infamous for weakly enforcing copyright law. Even when it is completely obvious that Chinese labs are training models on pirated data, US copyright holders face a virtually impossible task of proving it in court. Those lawsuits won't go anywhere.
- andriy_koval 2mo agoThere are tons of lawsuites which resulted in banning Chinese companies from doing business in US, those lawsuits totally have consequences.
- no-name-here 2mo agoWhat are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?
- deleted 2mo ago[deleted]
- reinitctxoffset 2mo agoHow is distillation an "attack" but gigascraping the Internet to the point of crashing servers and everyone needs Cloudflare and Anubis now not an "attack"? I'm not aiming for a what about kickflip here: I'm saying we need to either agree on some rules or stop crying foul. Maybe the coherent legal theory is that neural networks and intellectual property don't interact. That would be weird but it would be consistent, a market could price it, I could do coding stuff and know if I was illegaling. But this weird gerrymander that no judge will really rule on in an emphatic way is like, bad for the planet, bad for markets, bad business. There are a lot of reasons to look forward to DeepSeek Huggingface drop kicking the unambiguous frontier weights in like, November, but I think my favorite one will be "who's distilling now bitch?"
- gpm 2mo agoI think you've basically got the legal theory. Training a neural network isn't prohibited by copyright law so if you can legally get your hands on something (e.g. by sending a GET request to someone with rights to serve the contents of their web page, or by buying a book) without signing a contract to not train on it, you can train on it. But the American AI companies only let you query their models if you first sign a contract to not train on the output. It's hypocrisy and unfair, but I think there's a strong legal argument for it. Of course China can simply decline to assist in enforcing that contract... But I would expect US courts to do their best to.
- nickysielicki 2mo agoContract law is never going to prevent this.
- gpm 2mo agoWhy not? Seems like a perfectly normal contact term to me. Or do you just mean that US courts don't have enough teeth to prevent Chinese companies from violating contracts? On that I agree.
- 2mo ago
- Forgeties79 2mo agoThis is nothing like music piracy.
- throw1234567891 2mo agoAmerican labs have ripped everything out of the internet. And now they cry someone else is “stealing” from them. Cry me a river.
- areoform 2mo ago> The problem with distillation attacks I think it's worth stepping back here and pointing out the obvious. Y'all waging war on math. And I'm sorry, but that's the computing equivalent of legislating gravity. Apologies for repeating myself here, but what you call "distillation" is function approximation. I feel for the teams at Anthropic and Open AI, but unlike startups from prior eras; Anthropic and OpenAI have decided to be in the business of selling compute. Not creating a product that uses compute, but a product that's math running on compute. This is different from what Google is (or, rather was. As always, RIP Google 1998-2019). Google's algorithm might be math, but Google search isn't. Google search is a process that's continuously operating in the background. Google crawls pages. Google stores and indexes what it finds. Google then exposes this to retrieval via its algorithm. User uses algorithm. Now, let's compare that to AI models. When Anthropic serves Mythos / Opus etc, they're taking input or x from their user, doing compute, and then serving the result of the Mythos / Opus function, i.e., f(text) -> (text_transform) Where f is a continuous function, https://www.turing.ac.uk/sites/default/files/2025-11/language_models_are_implicitly_continuous.pdf https://www.turing.ac.uk/sites/default/files/2025-11/languag... According to Stone-Weierstrass, given enough values of y for f(x), anyone can approximate this function. The fidelity and sophistication of this approximation definitely requires a lot of cleverness and effort, and it is arguably an imposition on Anthropic and OpenAI. But on a long-enough timeline, they don't even have to poll Anthropic or OpenAI. As the internet is flooded by PRs, content, emails written by Mythos / Claude, and just people otherwise sharing the results of Claude prompts, then there's an ever increasing set of data to approximate the f(x) that's f_Claude. Eventually, in the future, anyone will be able to create a good enough approximation of the f_Mythos. Which is Anthropic's product. Anthropic and OpenAI can now wage war on mathematics and the open-ended compute. Or, they can adapt and build a better product. Choosing Option B was the Silicon Valley option / choice. I think the OG large-scale Valley lobbying effort, the Semiconductor Industry Association, was unique in that it prioritized and chose to do real research. https://en.wikipedia.org/wiki/Semiconductor_Industry_Association https://en.wikipedia.org/wiki/Semiconductor_Industry_Associa... https://en.wikipedia.org/wiki/Semiconductor_Research_Corporation https://en.wikipedia.org/wiki/Semiconductor_Research_Corpora... This helped the industry to survive and outcompete the pressure they were facing (at the time).
- 2mo ago