3 ms·
Aren't they based on qwen? I thought the clever bit with DeepSeek is a way to fine tune with reinforcement learning, not training a huge model from scratch.
by chpatrick 2y ago
Aren't they based on qwen? I thought the clever bit with DeepSeek is a way to fine tune with reinforcement learning, not training a huge model from scratch.
- vineyardmike 2y agoNo… mostly. They developed their “zero” model, which (they claim) is a from-scratch model. Base models are typically not fine-tuned for particular applications (eg chat). They trained their Zero model into a Chat+Reasoning model, which is what is attracting news. They ALSO fine-tuned small Qwen models using their big models as a teacher (distillation technique).
- supermatt 2y agodeepseek-r1 (reasoning) is a post-trained (not simply finetuned) deepseek-r3 (base model) by deepseek. qwen is a completely separate 3rd party model by alibaba. deepseek distilled (finetuned on the outputs of) r1 into qwen. They also did the same with llama.