3 ms·
That's not an issue. DeepSeek R1 is massive but people still managed to distill and quantize it. They'll just do the same for this one.
by rany_ 1y ago
That's not an issue. DeepSeek R1 is massive but people still managed to distill and quantize it. They'll just do the same for this one.
- wkat4242 1y agoDeepseek distilled it themselves using Qwen and Llama, right from the start.
- kelipso 1y agoPeople are reproducing the training for Deepseek R1.
- wkat4242 1y agoOh that's interesting, because they distilled only with 1 method, not both methods that were used on R1 (the one was reinforcement learning, the other supervised finetuning, I think only the former was used on the distillations). So there might be some room for improvement. I hadn't seen any third party distillations yet but I'll have a look.