5 ms·
The fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusa
by kamranjon 3mo ago
The fact that API based distillation is even a conversation right now makes me feel like the U.S. has their heads so far in the sand that it’s not really excusable.
These Chinese labs are producing novel models, publishing their techniques and sharing their open weights and the first topic of conversation is how they stole from U.S. AI labs.
Setting aside the fact that it doesn’t make any feasible sense to do API distillation, these models are outperforming frontier models on a number of benchmarks, and often times run more efficiently by several orders of magnitude.
We have to stop crying distillation, it’s getting embarrassing and at this point feels even a bit delusional.
- tristanj 3mo agoThere's little doubt that Kimi K3 was distilled off Claude. Anthropic stated in February that Moonshot AI (the creator of Kimi) distilled ~3.4 million exchanges from Claude models, as explained in their press release https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.anthropic.com/news/detecting-and-preventing-dist...
- overfeed 3mo agoWhile it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model? Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.
- tristanj 3mo ago3.4 million is the number of sessions Anthropic detected. The actual number of Claude sessions trained on is likely >100 million. There are tens of thousands of accounts funneling Claude sessions into Chinese labs https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... They are used for post-training, i.e. calibrating the model to understand and use tools/command line more effectively.
- overfeed 3mo ago> 3.4 million is the number of sessions Anthropic detected. The actual number of Claude sessions trained on is likely >100 million. That's an increase of only a single order of magnitude, increasing my estimate of exfiltrated tokens from 0.05 to 0.15 trillion - a far cry from the 15 trillion required. > They are used for post-training Possibly - it may be too much data for post-training, unless further curation was done. However, this is not distillation; you know it, I know it, Dario knows it, but "Distillation Attack" is a short, memorable, sciencey-sounding, political sound-bite with enough malevolence to be deployed on the floors of congress, or by the usual fear-mongering newstainment talking heads.
- tristanj 3mo agoYou're conflating pre-training data volume with post-training data volume. Nobody is suggesting Moonshot used 15 trillion tokens of Claude data to pre-train a base model from scratch. That would be impossible and nonsensical. This is entirely about distillation, which happens during post-training (alignment and SFT). Here, datasets are measured in millions or billions of tokens, not trillions. 50 billion Claude tokens is far, far than enough to copy Claude's reasoning logic, writing style, and tool-use ability to the pre-trained base model. > However, this is not distillation I don't understand how you're so caught up on the term "distillation". Distillation is using a larger model's outputs to train a (weaker) student model. Which is exactly what's happening. It's a standardized term that has been in use for a decade.
- overfeed 3mo agoThere is a lot of supposition going on your part and mine. IMO, Chinese labs are not dependent on OpenAI/Anthropic outputs; they definitely use the outputs, but along other training/post-training data. Now that Anthropic hides the real thinking tokens in a way that precludes future CoT distillation, we'll find out which side is correct based on whether Chinese AI labs close the gap or not. My bet is they'll close the gap; nothing about frontier AI is magic, once something is shown to be possible, experienced practitioners almost always figure out how to accomplish the same feat, though not always on the same way. This is why frontier US labs keep leapfrogging each other every few months.
- ACCount37 3mo agoWhat do you think those tokens are used for? Distillation attacks aren't about replacing the entire pretraining dataset with questionably sourced synthetics. It's all about post-training. Train your own base model - but tune it off Claude output to make it perform more in line with Claude. Yoink the products of Anthropic's expensive SFT, RLHF and RLVR work for yourself by training on the outcomes. The post-training datasets are small, but they are what controls the final model behavior.
- overfeed 3mo agoHow does yoinking outputs from from prior generation Claude model and post raining on them result in a model competitive with the latest generation? That doesn't add up - nevermind Anthropic hasbeen summarizing thinking tokens since January to counter distillation.
- ACCount37 3mo agoDo I really have to explain the shape of AI training pipelines to you? Train a big, wide base model with a lot of potential. Mid-train or post-train that on Claude Opus 4.5 reasoning/agentic traces (i.e. Claude Code data from Chinese API resellers) to make your model approximate a high baseline of chatbot behavior, reasoning, agentic work and tool use. Then run your own expensive SFT, RLHF and RLVR on top of that yoinked baseline to dial it in further. Actually doing RLHF and RLVR is extremely expensive. Distillation gives you a lot of dense, high quality post-training signal for cheap. This can get your model into the basin of "the right way to tackle this kind of problem" without a frontier lab compute budget. It's a big shortcut that gets you closer to the target - you can take it from there and build on top of it with your own work. Also, it's unclear whether "summarizing thinking tokens" actually ruins distillation, or just makes it work worse. I'd bet on the latter, really. Because it's an approximation game, and summarized reasoning is still a better approximation of true reasoning than most of what you get online and in pre-training datasets.
- zmmmmm 3mo ago> Train your own base model - but tune it off Claude output to make it perform more in line with Claude Is that actually genuine distillation though? Distillation suggests the core model is being pre-trained using output from another model. For the above to work, you have to already have all the core intelligence trained into your base model. If distillation just comes down to post-training then it's tantamount to admitting that the Chinese base models are just as good as frontier US lab models. Because you can't post-train frontier intelligence into a model. It has to be there in the base. Then you can change how that intelligence is expressed through post-training.
- kamranjon 3mo agoIt’s so funny to me that Anthropic can make claims like this one with zero evidence provided. DeepSeek and others like Minimax are publishing deep research on Multi-Head Latent Attention and Mixture of Experts, Multi-Token Prediction, novel Sparse Attention approaches, I mean they trained long context models on a fraction of the resources and gave everyone the recipe. Chinese labs might not have the funding of labs like Anthropic, but at least they provide the receipts.
- tristanj 3mo agoThere's reproducible evidence of Kimi K3 spontaneously identifying itself as Claude https://x.com/denisewu/status/2077984660211269870 https://x.com/denisewu/status/2077984660211269870 This behavior is exactly what you'd expect from a model distilled from Claude. Someone even took the time to analyze Kimi's ambiguous identity, in great detail: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/writeups/write_up.md https://github.com/rgreenblatt/which_claude_is_k3/blob/main/... And there's an entire Reddit thread discussing this https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kimi_k2_train_on_claudes_generated_code_i/ https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kim... That doesn’t prove Anthropic’s specific 3.4m-session allegation, but calling it “zero evidence” is no longer credible. Kimi K2.5 was worse in a hilarious way, it identified itself as Claude and referenced Anthropic's Constitutional AI as some of its guiding principles https://huggingface.co/moonshotai/Kimi-K2.5/discussions/38 https://huggingface.co/moonshotai/Kimi-K2.5/discussions/38
- staticman2 3mo ago> This behavior is exactly what you'd expect from a model distilled from Claude. This is not at all what I would expect because it's trivial to change the training data to replace Claude with Kimi. In fact I'd argue it's almost certainly not saying that due to distillation.
- tristanj 3mo agoI encourage you to review the links before committing to a position. The writeup on K3's anomalous trans-model identity is very comprehensive. K3 reproduces Claude's internal model identifier when prompted, something which the real Claude models themselves do not emit. This is highly suggestive that K3 was trained on Claude metadata (API logs, tagged synthetic data), rather than Claude's chat outputs. And it's well documented that Chinese labs are buying large amounts of raw Claude metadata https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
- overfeed 3mo ago> We have to stop crying distillation, it’s getting embarrassing and at this point feels even a bit delusional. It's a PR campaign - when they say its an "attack" they don't mean on Anthropic - but on America itself. What kind of American can let such a brazen attack go unanswered? At the very least, they ought to demand the dangerous, pinko, stolen models be banned in all 50 states, and pay whatever price demanded by the patriotic, freedom-loving, all-American AI labs that can never be accused of stealing.
- spaceman_2020 3mo ago[dead]