11 ms·
Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going t
by nickysielicki 3mo ago
Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a cheaper version of it. There was never any plausible explanation for why this wouldn’t happen. There was never any practical mechanism to prevent someone from saving a conversation and using it to train their own model.
Even if it didn’t happen here, it was still the case that it was going to happen going forward. It was always going to end like this. Invest in the hardware companies, not the model companies.
- jstummbillig 3mo ago> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthropics IP and b) what Anthropic did to build their models is legally questionable (or might be ruled illegal, even though I doubt it).
- amazingamazing 3mo agoThe value is simply that it is easier. The same way it is easier to ask someone who has experience for advice than reading hundreds of textbooks.
- nickysielicki 3mo agoRegardless of whether it’s intellectual property or it isn’t intellectual property, it doesn’t actually matter. If AI doesn’t stop seeing diminishing returns in scaling up, and it hasn’t yet in the 10 years since the attention/transformers paper, the advent of AI will be the most important development in the history of humanity. Controlling that machine, or at least having one of your own, is an existential problem for nation states. It’s like a matter of national defense. Do you really think intellectual property laws will prevent this in practice? It’s like as if we said, “hey, USSR, you can’t make a nuke, too! We patented that already.” Asking China to not distill our models down is equally as ridiculous.
- ctippett 3mo agoIt's unlikely the USA would be granted an exclusive patent for the atomic bomb given the well-established existence of prior-art in the form of nuclear fission on the sun. (I actually appreciated your analogy, despite my lark)
- idiotsecant 3mo agoIn fact, China stealing fire from the gods is essential to the future balance of power, so long as they keep making the results freely available.
- tadfisher 3mo agoI do like that you mention diminishing returns, because we are hitting them in building out all the external requirements for competing at the frontier. Even if model performance scales linearly with energy input, the top labs are now competing with other uses for that energy. How far are we willing to go as a nation (and as a species) to prove out the scaling laws? Are we willing to sacrifice our industrial base? Would we rather train models or smelt aluminum?
- athrowaway3z 3mo ago>People infringe on Anthropics IP No. Authors do not infringe on IP when they read another's book, nor should the lumber company be able to dictate how I use planks and if I can resell them if i'm done with them. You're framing it as if the added value of the author or lumber company, awards them consideration when somebody uses the products to create more value. IP law was always a big mess, and these questions cross far into ideology instead of law; but I do not understand people who think we need an ideology where more IP-law is good for society.
- throe74844949 3mo agoThere are some quite interesting legal implications here. If Anthropic has IP over output produced by agents, do they somehow have legal rights to code and documents produced by such agents? This would demolish agent usage by corporations.
- throw1234567891 3mo agoYou say “if”. How did Anthropic obtain this IP, if the model serves ripped internet and all human knowledge?
- p_l 3mo agoGeneral consensus is that neither the model nor its outputs can be protected IP
- jstummbillig 3mo agoIt's more simple: They infringe on the IP by way of violating the ToS. If you violate ToS and the company suffers financial harm, they usually can (usually) sue you in civil court for damages.
- ForHackernews 3mo agoYou can't violate ToS you never agreed to. If I use pirate Claude through a third-party reseller, I have entered no agreement with Anthropic.
- justincormack 3mo agoThe output of Anthropic's models is not Anthropic's IP, as that would destroy their market, if Anthropic owned all the software it generated, and all the content. So distillation, which is just using those outputs is always going to exist.
- Jtarii 3mo agoWhether or not its legal to distil models, it is obviously morally permissible to do so. Anthropic, OpenAI, etc do not deserve legal protection.
- InsideOutSanta 3mo agoI'm pretty sure that LLM output is not intellectual property. Nobody owns it, and it can't even be copyrighted. So using output from Anthropic's LLMs in ways Anthropic does not condone is not IP infringement.
- whateveracct 3mo agoanthropic model output is not their IP that would be existential doom for them because then they have a case to claim ownership of their users' codebases no corporation would sign off on that
- NitpickLawyer 3mo ago> People infringe on Anthropics IP Unless someone literally stole the weights somehow (which is not out of the question, I doubt either oAI/Anthropic have the capabilities to prevent a state-level actor getting those weights), distillation from generations is not infringement on anyone's IP nor is it stealing nor is it an attack. It can't be. As long as you pay for tokens you get to do whatever you want with them. Someone saying you can't doesn't mean it's an attack or their IP or whatever. They either sell the tokens or not. They can decide to not sell them to anyone, but again that's not stealing. And their ToS are a joke. Imagine how people would react if MS had ToS saying that you can't use MS software to develop solutions that compete with MS. They'd be laughed out of the room. Somehow it's ok for token sellers to decide what you do with the tokens? Why? If you pay for something you get to do whatever you want with that output. Train, distill, whatever.
- nonethewiser 3mo agoIts definitely an attack. Thats established from anthropics perspective. No one has a right to use Anthropic’s services in ways that directly violate the ToS and user agreements.
- georgemcbay 3mo ago> Its definitely an attack. Thats established from anthropics perspective. How do things get "established" from someone's perspective, exactly? By that logic it is established from my perspective that Anthropic has no right to train on anything I've written that is publicly available on the internet. Of course, they don't care about my perspective, but then again I don't care about theirs.
- anigbrowl 3mo agoThere's a lot of people diligently shilling for Anthropic in this thread. That's established from my perspective.
- NitpickLawyer 3mo ago> from anthropics perspective I guess I can see that, if you mean the targeted effort of creating many accounts w/ the intent of doing it at scale. Sure, they may see that as an attack. But again, it's only an attack from their perspective if you agree that using generations to distill is "wrong". I just don't see it, in general. You can't both sell tokens and decide that distilling is somehow illegal. Something, something, cake and eat it.
- camillomiller 3mo agoAnthropic’s IP is basically null and void for how they created it. And they might not want to try and challenge this in court, considering how they had to settle for using text books they had no right to use
- sensanaty 3mo agoConsidering they were the original infringers, I don't know how anyone can expect tears to be shed here. The best we can hope for is for all these cancerous - and they really are the definition of a cancer - money burning entities to all fall apart to distillation attacks like these.
- anon373839 3mo ago> People infringe on Anthropics IP Anthropic’s model outputs contain no IP. This is actually a simple legal proposition (rare in this field!) that derives from the fact that only specific classes of IP exist: copyrights, patents, trade secrets, and trademarks. Examining each, it is clear that API outputs do not qualify. Anthropic disclaims copyright in outputs; the outputs are not patented; the outputs are not secret (a prerequisite to having trade secrets); and trademarks are irrelevant in concept.
- inigyou 3mo agoWhat IP is being infringed on? AI outputs aren't copyrightable.
- milkshakes 3mo agoassume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts. what is the end game for this strategy? if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?
- traverseda 3mo agoIn public with budgets that don't risk destroying the American economy presumably. Yes it may be slower.
- milkshakes 3mo ago> with budgets and what will fund these budgets exactly? inference is cheap, distillation is cheap, training is what's expensive.
- Jtarii 3mo agoPresumably the US military / NSA.
- milkshakes 3mo agothe USG/NSA will fund chinese labs? to what end?
- Jtarii 3mo agoI was more thinking they would be funding US labs.
- milkshakes 3mo agothe question was: what is the endgame for the stated "second class labs" strategy of distilling their frontier competitors then undercutting them on price?
- boondongle 3mo agoAlmost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape." Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether it's complete loss of access, or the amount of control you'll have to give up to access them will be ridiculous. Sort of like the "stealing music is fine" but "lets freak out now that it's producing visual art", in the end the entire thing is a social construct. Whether this is treated as theft or "business as usual" is entirely societal. Eventually the gap will close, unless there's a major breakthrough that hasn't been made yet.
- zaptrem 3mo agoGiven these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).
- andriy_koval 3mo agoUs models didnt pay for licenses too
- joe_mamba 3mo agoWe're still in the early days of the AI industry timeline(relative to traditional industries). Not everything has yet been litigated. Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature. If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pay extra taxes on any and all storage mediums and on devices with built-in storage (tapes, CDs, DVDs, HDDs, SSDs, tablets, phones, etc) simply because they can be used to store pirated content, decisions based on laws from 50-100 years ago, and the money goes to the national unions and associations of music and arts IP holders. It's basically a lobby pushed and government legalized extortion racket that no voter agrees with or can change but has no choice but to conform either way. So I guarantee you in the future, it will be the same for AI subscriptions and hardware capable of running LLMs locally. Every time you purchase a Claude or ChatGPT subscription, an Nvidia GPU, Intel/AMD SoC PC or an Apple/Qualcomm powered smartphone, you'll pay a government enforced tax to the likes of Sony, Axel Springer, etc. for licensing their IP, whether you want to or not. In the EU at least. US maybe not.
- solumunus 3mo agoThanks for the models guys, sorry for your losses. Once this reality becomes mainstream and undeniable, surely the bubble pops and then what then. Future model development stops? Becomes private? Becomes a public effort?
- nickysielicki 3mo agoThe existing models are still going to exist. As hardware improves, there will be a day where it might cost a tenth of a penny to churn through 100M tokens a second of Opus 4.8. Established compute providers will invest in improving the models incrementally when margins drive them to look there.
- dzonga 3mo agoor the application layer - which will capture majority of the value. yeah hardware companies make for nice stories or green numbers on Wall Street - but value will be captured by application layer. look at history.
- nickysielicki 3mo agoThat’s true up until the point where you can ask the hardware you made to make its own application layer.
- nonethewiser 3mo ago>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?
- ahofmann 3mo agoIs this some form of rage bait? 2005 we hadn't the GPUs, we have today. There are other factors, but I think this is the big one. The mathematics of building an LLM are really old, we just hadn't the hardware to do the needed calculations.
- nonethewiser 3mo agoRight. Therefor it's not simply a derivative of information. The hardware is required to build the model. Software as well. The model uses information, it is not "distilled" from it. "Distillation" literally means to separate and take some components out of something. You can distill how a model works from a model. You cant distill a model from information because the information does not contain the model. People are happy to conflate distilling with building because they dont like how the information was used. You distill how the model works from the model, and you build a model with information. Both could be morally good or bad but its not the same thing.
- realusername 3mo agoThat argument is moot as distillation also requires a lot of hardware and software, if copying models was as easy as that, we would have hundreds of competing models.
- nonethewiser 3mo agoNo. Building models and distilling models both require the hardware and software. It doesn't mean building models is distillation.
- girvo 3mo ago
- genxy 3mo agoLook how hard Anthropic is to even be able scroll back on your conversation, or look at the thinking tokens or subagents. They want to keep everyone coming back to the watering hole but never to learn how to dig a well.
- drcongo 3mo agoWhy is it hard to scroll?
- arcanemachiner 3mo agoDid you enable the flicker-free TUI mode?
- anon373839 3mo agoI strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “cold start” problem faster. By far, what matters more is the quality and variety of RL environments the model learns from.
- tristanj 3mo agoAPI distillation doesn't have to explain all of K3's capabilities for it to have happened. Kimi K3 reproducibly identifies itself as Claude: https://x.com/denisewu/status/2077984660211269870 https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/writeups/write_up.md https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... This analysis observed K3 identifies itself as Claude approximately 15% of the time. K3 reproduces Claude's correct current model id, which the real Claude models themselves do not emit. This suggests K3 was trained on Claude data labeled with deployment metadata (API logs, tagged synthetic data), rather than Claude's chat outputs. And there's an entire Reddit thread discussing Kimi's similarities with Claude https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kimi_k2_train_on_claudes_generated_code_i/ https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kim... This analysis shows K3 and Opus/Fable have unexpected correlated outputs https://typebulb.com/u/lab/you-re-relatively-right/full https://typebulb.com/u/lab/you-re-relatively-right/full
- deminature 3mo agoIt does not reproducibly identify itself as Claude, there's evidence to the contrary in the very thread you linked: https://x.com/bobbyNewcomb5/status/2078151562828947954 https://x.com/bobbyNewcomb5/status/2078151562828947954
- 3mo ago
- kamranjon 3mo agoThe fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusable. These Chinese labs are producing novel models, publishing their techniques and sharing their open weights and the first topic of conversation is how they stole from U.S. AI labs. Setting aside the fact that it doesn’t make any feasible sense to do API distillation, these models are outperforming frontier models on a number of benchmarks, and often times run more efficiently by several orders of magnitude. We have to stop crying distillation, it’s getting embarrassing and at this point feels even a bit delusional.
- tristanj 3mo agoThere's little doubt that Kimi K3 was distilled off Claude. Anthropic stated in February that Moonshot AI (the creator of Kimi) distilled ~3.4 million exchanges from Claude models, as explained in their press release https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.anthropic.com/news/detecting-and-preventing-dist...
- overfeed 3mo agoWhile it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model? Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.
- tristanj 3mo ago3.4 million is the number of sessions Anthropic detected. The actual number of Claude sessions trained on is likely >100 million. There are tens of thousands of accounts funneling Claude sessions into Chinese labs https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... They are used for post-training, i.e. calibrating the model to understand and use tools/command line more effectively.
- akrymski 3mo agoWell, there is precedence: Google can scrape the web, but you can't scrape Google. Laws around compiled databases exist for a reason: you can't just copy the phone book if effort has gone into compiling it, it is itself copyrightable
- p_l 3mo agoAnd funnily enough, most laws about compiled databases might not apply, for one. And then there's new updates related to AI that fully take out LLMs from protection.
- philipkglass 3mo agoThat varies by jurisdiction. In the United States, copying the phone book (or otherwise copying facts from someone else's collection) has been legal since 1991: https://en.wikipedia.org/wiki/Feist_Publications,_Inc._v._Rural_Telephone_Service_Co https://en.wikipedia.org/wiki/Feist_Publications,_Inc._v._Ru....
- anigbrowl 3mo agoThis is the opposite of legal reality, at least as far as the US is concerned.
- matheusmoreira 3mo ago> Distillation “attacks” are not attacks. Say it louder for the people in the back. All these complaints about "distillation" from frontier labs are bordering on felony contempt of business model at this point. It's great for us. Maybe it's bad for them but nobody other than shareholders really cares. The optimal outcome for humanity is for oligarchs to spend trillions training a godlike AI, only for the precious weights to just leak. No "distillation" required.
- inigyou 3mo agoIt's literally felony contempt of business model, except, unlucky for them, it isn't a felony or any kind of crime, so they just have to seethe harder.
- casualscience 3mo ago> There was never any plausible explanation for why this wouldn’t happen. What a nice post hoc revision of history. Distillation is still an active area of research, that you can distill models as easily as you can it genuinely interesting and absolutely not something that was taken for granted even 12 months ago. Even 6 months ago this idea that 'using model outputs as training examples' was listed as the reason that all models would fail in the near future due to some spooky circular training catastrophe. Don't pretend like this was so obvious.
- stanfordkid 3mo agoI think you’re being overly combative. It’s intuitively quite obvious that it’s incredibly easy to implement and the circular training catastrophe was only ever a conjecture. It’s kind of like releasing a crypto primitive without knowing a proof. Like… maybe it works, but you can’t assume that just because you don’t know how to break it. You have to remember that 100s of billions of enterprise valuation rely on frontier models being moats. The burden of proof is on those raising valuations assuming they will capture the full market.
- chrishare 3mo agoI agree that hindsight is doing work here, but DeepSeek R1 from Jan 2025 seemed to heavily leverage distillation, and 18 months is an eternity in this climate.
- MichaelMoser123 3mo agoI suspect that distillation attacks may be slightly exaggerated. Most of the training data used during fine-tuning is now synthetic data. You can't just repeat the same stuff twice, therefore another LLM is writing a text book that is explaining a topic in detail, ideally without any gaps in the material.
- bentocorp 3mo agoCalling distillation an 'attack' is exactly what I've been describing as "AI Exceptionalism": https://www.magiclasso.co/insights/ai-exceptionalism/ https://www.magiclasso.co/insights/ai-exceptionalism/
- rayiner 3mo agoThe desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patterns, etc. There will be fields where China is a global leader, and Americans and Europeans will have to learn Chinese and move there, or else be stuck in some satellite office of a Chinese company. We’re all in Europe circa 1895 not realizing the behemoth America will become in WWI.
- warshinder 3mo agoSo the efficient market hypothesis is wrong?
- rayiner 3mo agoHow is the efficient market hypothesis applicable here?
- inigyou 3mo agoThe efficient market hypothesis, read loosely, says that capitalism is the best system. Yet here it is being thoroughly pwned by a series of 5-year plans.
- s1artibartfast 3mo agoexactly zero serious people draw that connection, and it suggests that you either dont know what the term means or are constructing a strawman. Most of the economic literature for the last 150 has explored the constraints on markets and their function. The efficient market hypothesis TM is a very narrow theory about price information, it says nothing about Economic Development or Direction. It is a theoretical extreme that can be used to compare real world Systems. Further, China's success relies heavily on market processes
- 3mo ago
- blitzar 3mo agoIt is unfair, they stole the dataset that we stole.
- JSR_FDED 3mo agoThe court decided that LLMs are a transformative fair use of the data they trained on, and therefore aren’t copyright infringement. Maybe Kimi is a derivative work as well
- dannyw 3mo agoLLM outputs do not have copyright protection in the US, there is no copyright element here.
- XorNot 3mo agoThe hand wringing over whether internationally located AI labs are "stealing" output from American ones is the funniest thing in a while. It's international politics with people talking about AI success as a matter of national strategic advantage and survival. So at best "this was built off our work" mostly tells you that apparently you've got months of advantage when a new model drops before it can be cloned. That's certainly some sort of advantage, sure hope it represents a consistent ability to stay ahead and causes people to redouble their efforts. Or...of course none of these companies are worth what they say, but the advantage is also not really that great, and a whole lot of people are just really worried about their stock payouts.
- geokon 3mo agoare there any "open source" efforts to do distillation? Like some place one can submit one's anonymized chat logs? So they can be pooled and used as an open training set (similar to OpenCrawl)
- guesswho_ 3mo ago[dead]
- dkrich 3mo agoPulling on this thread, if the model companies become commoditized and make no money then who is buying the hardware? Seems like it would be the next shoe to drop
- Spacecosmonaut 3mo agoI wonder, how does distillation deal with unprobed spaces in the knowledge landscape? Is a distilled model worse in some niche area that was not probed? Presumably, this is why frontier labs dont distill their own models internally to release them to the public as a servicable frontier model.
- tngranados 3mo agoDo frontier labs not distill their own bigger models into the smaller/cheaper variants? I thought that’s been the case for a while
- iloveoof 3mo agoA “distillation attack” is like a concrete company calling a competitor building a factory with its concrete a “construction attack”
- titaniumrain 3mo agoif distillation is the key, why the fuck all other competitors do not release competitive models? and only Chinese can distill this great?! Am I smoking too much?
- ohhnoodont 3mo ago> There was never any plausible explanation for why this wouldn’t happen. There was never any practical mechanism to prevent someone from saving a conversation and using it to train their own model. At this point it may not even be happening intentionally given the quantity of LLM-generated content that is appearing online and is likely being re-ingested by models.
- jrichard2026 2mo agoI saw your post about the Kimi K3 Moment and how you mentioned the pricing of models from Anthropic, OpenAI, and Kimi, I thought you might be interested in ShipDiff, which could help with tracking competitor pricing strategies, feel free to ignore this if not relevant