3 ms·
aren't the smaller param models all just Qwen/Llama trained on R1 600bn?
by john_alan 2y ago
aren't the smaller param models all just Qwen/Llama trained on R1 600bn?
- whimsicalism 2y agoyes, this is all ollamas fault
- john_alan 2y agoYeah I don’t understand why
- yetanotherjosh 2y agoollama is stating there's a difference: https://ollama.com/library/deepseek-r1 https://ollama.com/library/deepseek-r1 "including six dense models distilled from DeepSeek-R1 based on Llama and Qwen. " people just don't read? not sure there's reason to criticize ollama here.
- whimsicalism 2y agoi’ve seen so many people make this misunderstanding, huggingface clearly differentiates the model, and from the cli that isn’t visible