4 ms·
I understand people just get off posting stuff like this. But creating LLMs from the entire corpus of human text was a huge achievement. Distilling those models
by slibhb 23d ago
I understand people just get off posting stuff like this. But creating LLMs from the entire corpus of human text was a huge achievement. Distilling those models is much less of an achievement. It means China is further behind than we thought.
- vanviegen 23d agoThey're using competitor output as additional training input. Shrug.. It's not like they have access to the weights.
- ezekiel68 23d agoFirst movers rarely take the prize, though, do they? Les Paul and Mary Ford pioneered overdubbing voices back in the 1940s and 50s. The Beatles stole it from Buddy Holly. And Elton John from the Beatles.
- davyAdewoyin 23d agoHow is China further behind if distillation cannot stop? I think it's a reasonable strategy to follow, even if they could train from scratch.
- undeveloper 23d agodistillation results in a worse product than the actual teacher model iirc
- SXX 23d agoNobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models. Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.
- Zambyte 23d agoDistilling is a massive achievement. I can run Qwen. I can't run GPT (TM). It's not a matter of X is better than Y. It's a matter of Y exists, X does not.