7 ms·
This continues the pattern of all other announcements of running 'Deepseek R1' on raspberry pi - that they are running llama (or qwen), modified by deepseek's d
by tofof 2y ago
This continues the pattern of all other announcements of running 'Deepseek R1' on raspberry pi - that they are running llama (or qwen), modified by deepseek's distillation technique.
- corysama 2y agoYeah. People looking for “Smaller DeepSeek” are looking for the quantized models, which are still quite large. https://unsloth.ai/blog/deepseekr1-dynamic https://unsloth.ai/blog/deepseekr1-dynamic
- zozbot234 2y agoYes this is just a fine-tuned LLaMa with DeepSeek-like "chain of thought" generation. A properly 'distilled' model is supposed to be trained from scratch to completely mimick the larger model it's being derived from - which is not what's going on here.
- kgeist 2y agoI tried the smaller 'Deepseek' models, and to be honest, in my tests, the quality wasn't much different from simply adding a CoT prompt to a vanilla model.
- rcarmo 2y agoYet for some things they work exactly the same way, and with the same issues :)
- whereismyacc 2y agoI really don't like that these models can be branded as Deepseek R1.
- sgt101 2y agoWell, Deepseek trained them?
- yk 2y agoYes, but it would've been nice to call them D1-something, instead of constantly having to switch back and forth between Deepseek R1 (here I mean the 604B model) as distinguished from Deepseek R1 (the reasoning model and it's distillates.)
- mdp2021 2y ago? Alexander is not Aristotle?!
- sgt101 2y agoyou made my day!
- tucnak 2y agoI don't know if they'd changed the submission title or what, but it says quite explicitly "Deepseek R1 Distill 8B Q40" which is a far-cry from "Deepseek R1" which would be misrepresenting the result, indeed. However, if you refer to Distilled Model Evaluation[1] section of the official R1 repository, you will note that DeepSeek-R1-Distill-Llama-8B is not half-bad; it supposedly out-performs both 4o-0513 and Sonnet-1022 on a handful of benchmarks. Remember sampling from formal grammar is a thing! This is relevant, because llama.cpp has GBNF, and lazy grammar[2] setting now, which is making it double not-half-bad for a handful of use-cases, less of all deployments like this. That is to say, the grammar kicks in after </think>. Not to mention, it's always subject to further fine-tuning: multiple vendors are now offering "RFT" services, i.e. enriching your normal SFT dataset with synthetic reasoning data from the big-boy R1 himself. For all intents and purposes, this result could be much more valuable prior than you're giving it credit for! 6 tok/s decoding is not much, but Raspberry Pi people don't care, lol. [1] https://github.com/deepseek-ai/DeepSeek-R1#distilled-model-evaluation https://github.com/deepseek-ai/DeepSeek-R1#distilled-model-e... [2] https://github.com/ggerganov/llama.cpp/pull/9639 https://github.com/ggerganov/llama.cpp/pull/9639
- hangonhn 2y agoCan you explain to a non ML software engineer what these distillation methods mean? What does it mean to have R1 train a Llama model? What is special about DeepSeek’s distillation methods? Thanks!
- littlestymaar 2y agoSubmit a bunch of prompts to Deepseek R1 (a few tens of thousands), and then do a full fine tuning of the target model on the prompt/response pair.
- dcre 2y agoDistilling means fine-tuning an existing model using outputs from the bigger model. The special technique is in the details of what you choose to generate from the bigger model, how long to train for, and a bunch of other nitty gritty stuff I don’t know about because I’m also not an ML engineer. Google it!
- lr1970 2y ago> Distilling means fine-tuning an existing model using outputs from the bigger model. Crucially, the output of the teacher model includes token probabilities so that the fine-tuning is trying to learn the entire output distribution.
- numba888 2y agoThat's possible only if they use the same tokens. Which likely requires they share the same tokenizer. Not sure that's the case here, R1 was built on OpenAI closed model's output.
- anon373839 2y agoThat was an (as far as I can tell) unsubstantiated claim made by OpenAI. It doesn’t even make sense, as o1’s reasoning traces are not provided to the user.
- HPsquared 2y agoAnd DeepSeek itself is (allegedly) a distillation of OpenAI models.
- alexhjones 2y agoNever heard that claim before, only that a certain subset of re-enforced learning may have used ChatGPT to grade responses. Is there more detail about it being allegedly a distilled OpenAI model?
- IAmGraydon 2y agoHe didn't say it's a distilled OpenAI model. He said it's a distillation of an OpenAI model. They are not at all the same thing.
- scubbo 2y agoHow so? (Genuine question, not a challenge - I hadn't heard the terms "distilled/distillation" in an AI context until this thread)
- janalsncm 2y agoThere were only ever vague allegations from Sam Altman, and they’ve been pretty quiet about it since.
- blackeyeblitzar 2y agohttps://www.newsweek.com/openai-warns-deepseek-distilled-ai-models-reports-2022802 https://www.newsweek.com/openai-warns-deepseek-distilled-ai-... There are many sources and discussions on this. Also DeepSeek recently changed their responses to hide references to various OpenAI things after all this came out, which is weird.
- littlestymaar 2y agoMeanwhile on /r/localllama, people are running the full R1 on CPU with NVMe drives in lieu of VRAM.
- numba888 2y agoDid they get the first token out? ;) Just curious, NVidia ported it, and they claim almost 4 tokens/sec on 8xH100 server. At this performance there are much cheaper option.
- littlestymaar 2y ago> Did they get the first token out? ;) Suprisingly it's not *that* bad, with 3t/s for the quantized models: https://www.reddit.com/r/LocalLLaMA/comments/1in9qsg/boosting_unsloth_158_quant_of_deepseek_r1_671b/ https://www.reddit.com/r/LocalLLaMA/comments/1in9qsg/boostin... > NVidia ported it, and they claim almost 4 tokens/sec on 8xH100 server. What? That sounds ridiculously low, someone just got 5.8t/s out of only one 3090 + CPU/RAM using the KTransformers inference library: https://www.reddit.com/r/LocalLLaMA/comments/1iq6ngx/ktransformers_21_and_llamacpp_comparison_with/ https://www.reddit.com/r/LocalLLaMA/comments/1iq6ngx/ktransf...
- btown 2y agoSpecifically, I've seen that a common failure mode of the distilled Deepseek models is that they don't know when they're going in circles. Deepseek incentivizes the distilled LLM to interrupt itself with "Wait." which incentivizes a certain degree of reasoning, but it's far less powerful than the reasoning of the full model, and can get into cycles of saying "Wait." ad infinitum, effectively second-guessing itself on conclusions it's already made rather than finding new nuance.
- pockmarked19 2y agoThe full model also gets into these infinite cycles. I just tried asking the old river crossing boat problem but with two goats and a cabbage and it goes on and on forever.
- avereveard 2y agoThis has been brilliant marketing from deepseek and they're gaining mindshare at a staggering rate.