3 ms·
"SWE-2 is post-trained from Kimi K3"
by Tsarp 22d ago
"SWE-2 is post-trained from Kimi K3"
- airstrafer 22d agoYeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model. Maybe still worth it if their "64% cheaper" figure holds.
- Tsarp 22d agoWith the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.
- Bolwin 22d agoI don't think you know what distill means
- airstrafer 22d agoI guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...?
- FergusArgyll 22d agoDistilling you don't have the actual model weights of the teacher. All you have are the teachers answers to a lot of questions. You then teach your own smaller model to answer more similarly to the big teacher model. Fine tuning you have the actual model weights of the original model, you then train that model to answer in a different (or better) way.
- Bolwin 22d agoDistillation requires you to have the actual logits of each token from the teacher model, which in practice means having the model itself. What you're describing is just synthetic data. Note Anthropic misused the term in their post about Chinese model distillation, deliberately I assume.
- FergusArgyll 22d agoThere are multiple kinds of distillation https://arxiv.org/abs/2106.03310 https://arxiv.org/abs/2106.03310 https://arxiv.org/abs/2207.12106 https://arxiv.org/abs/2207.12106
- xlbuttplug2 22d agoI presume post training is significantly easier than the distillation/training the top Chinese labs are doing. I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.