3 ms·
> Edit: Also, side note, why are so many people running their LLMs in the cloud? All the cutting edge models are open weight licensed, and run locally. You don'
by krapht 2y ago
> Edit: Also, side note, why are so many people running their LLMs in the cloud? All the cutting edge models are open weight licensed, and run locally. You don't need to depend on some corporation that will inevitably rug-pull you.
???
Deepseek R1 doesn't run locally unless you program on a dual socket server with 1 TB of RAM. Or enough cash to have a cabinet of GPUs. The trend for state-of-the-art LLMs is to get bigger over time, not smaller.
Look, I've played with llava and llama locally too, but the benchmarked performance is nowhere near what you can get from the larger cloud providers who can serve hundred-million+ parameter models without quantization.
- DiabloD3 2y agoYou wouldn't use full fledged R1 for coding. There are distilled models using R1 for coding that get you most of the way there. R1 also doesn't take 1TB of RAM, go use read Unsloth's writeup on how to reduce model size without reducing quality (they got it to fit into 131GB): https://unsloth.ai/blog/deepseekr1-dynamic https://unsloth.ai/blog/deepseekr1-dynamic tl;dr parameter count is where the statistical model lives or dies, not weight precision; you can't blindly shrink every weight, and tooling is learning how to not butcher models. Also, performance between cloud-ran models and models I've ran locally with llama.cpp seem to be actually pretty similar. Are you sure your model didn't fit into your VRAM, or something else may have been misconfigured? Not fitting into VRAM slows everything to a halt. All the coder models that are worth looking at fit into 24GB cards in their full sized variants with the right quantization.
- satvikpendem 2y agoDistilled "DeepSeek" models are not actually DeepSeek and should not be referred to as such.
- DiabloD3 2y agoNo one said they were. They're distilled using the original model and the same weights that match the ones in R1. Its ostensibly the original, but better. There are also fused models such as https://huggingface.co/FuseAI/FuseO1-DeepSeekR1-Qwen2.5-Coder-32B-Preview https://huggingface.co/FuseAI/FuseO1-DeepSeekR1-Qwen2.5-Code... that also seem to perform interestingly.
- satvikpendem 2y agoThey are not better, they are strictly worse in every way, and the performance characteristics show such degradation. There is a difference between distilling and quantizing, as while the latter does show some degradation too, it is not to the extent of distilled models and at least it's still the original model.
- DiabloD3 2y agoDepends how you define "worse in every way". With my own personal testing, DeepSeek's distillations have been able to do the task when the original upstream model either couldn't, or was marginally worse yet. You're preaching to the anti-choir on this, though: I do not think LLMs are ready for use yet. Maybe another few years, maybe another few decades, we'll find out I guess, but what we have today sure as hell isn't it.