5 ms·
> their main trick for model improvement is distilling the SOTA models Could you elaborate? How is this done and what does this mean?
by _fizz_buzz_ 10mo ago
> their main trick for model improvement is distilling the SOTA models
Could you elaborate? How is this done and what does this mean?
- MobiusHorizons 10mo agoI am by no means an expert, but I think it is a process that allows training LLMs from other LLMs without needing as much compute or nearly as much data as training from scratch. I think this was the thing deepseek pioneered. Don’t quote me on any of that though.
- tickerticker 10mo agoYes. They bounced millions of queries off of ChatGPT to teach/form/train their DeepSeek model. This bot-like querying was the "distillation."
- SirMaster 10mo agoWhy would OpenAI allow someone to do that?
- qcnguy 9mo agoThey don't anymore. They introduced ID verification shortly after, but it's hard to stop completely while also scaling fast.
- MadnessASAP 10mo agoThey didn't, but how do you stop it? Presuming the scale that OpenAI is running at?
- orbital-decay 10mo agoThey definitely didn't. They demonstrated their stuff long before OAI and the models were nothing like each other.
- tensor 10mo agoNo, distillation is far older than deepseek. Deepseek was impressive because of algorithmic improvements that allowed them to train a model of that size with vastly less compute than anyone expected, even using distillation. I also haven’t seen any hard data on how much they do use distillation like techniques. They for sure used a bunch of synthetic generated data to get better at reasoning, something that is now commonplace.
- MobiusHorizons 10mo agoThanks it seems I conflated.