3 ms·
Wait so Qwen trained QWQ 32B from Qwen 32B and then they distill QWQ back into Qwen 32B? What's the point? This is massive marketing scam here. Borderline acad
by rlforllms 2y ago
Wait so Qwen trained QWQ 32B from Qwen 32B and then they distill QWQ back into Qwen 32B? What's the point?
This is massive marketing scam here. Borderline academic dishonesty.
- barrenko 2y agoNot sure if scam, honestly depends on the data sometimes it might work.
- rlforllms 2y agoThe goal is distillation is to distill into smaller models like 7B, 1.5B. They didn't even change the model size, let alone try a different class of models. Getting expert model's trajectories is trivial if you have vLLM to do batched inference.
- jojaja 2y agoSo you are better off just using QwQ
- andy_xor_andrew 2y agoI wouldn't go that far, but I agree, my reaction to reading the details was to go "huh?" From the title, my best guess was they applied some kind of RL/GRPO to an existing model. But... they took an existing model that had already undergone SFT for reasoning... and then used it to generate data to SFT the exact same model... nothing wrong with that, but it doesn't seem to warrant the title they chose.