5 ms·
Everyone is trying to say its better than the biggest closed models. It feels like it has parity, but its not the clear winner. But, its free and open and the
by strangescript 2y ago
Everyone is trying to say its better than the biggest closed models. It feels like it has parity, but its not the clear winner.
But, its free and open and the quant models are insane. My anecdotal test is running models on a 2012 mac book pro using CPU inference and a tiny amount of RAM.
The 1.5B model is still snappy, and answered the strawberry question on the first try with some minor prompt engineering (telling it to count out each letter).
This would have been unthinkable last year. Truly a watershed moment.
- the_real_cher 2y agoyou don't mind me asking how are you running locally? I'd love to be able to tinker with running my own local models especially if it's as good as what you're seeing.
- strangescript 2y agohttps://ollama.com/ https://ollama.com/
- rpastuszak 2y agoHow much memory do you have? I'm trying to figure out which is the best model to run on 48GB (unified memory).
- Metacelsus 2y ago32B works well (I have 48GB Macbook Pro M3)
- whimsicalism 2y agoyou’re not running r1 dude. e: no clue why i’m downvoted for this
- smokel 2y agoYou are probably being downvoted because your comment is not very helpful, and also a bit rude (ending with "dude"). It would be more helpful to provide some information on why you think this person is not using R1. For example: You are not using DeepSeek-R1, but a much smaller LLM that was merely fine-tuned with data taken from R1, in a process called "distillation". DeepSeek-R1 is huge (671B parameters), and is not something one can expect to run on their laptop.
- zubairshaik 2y agoIs this text AI-generated?
- tasuki 2y agoProbably. It's helpful tho, isn't it?
- smokel 2y agoI actually wrote it myself. I set a personal goal in trying to be more helpful, and after two years of effort, this is what comes out naturally. The most helpful thing that I do is probably not posting senseless things. I do sometimes ask ChatGPT to revise my comments though (not for these two).
- tasuki 2y agoYou have reached chatgpt level helpfulness - congrats!
- zubairshaik 2y agoWasn't a value judgement
- john_alan 2y agoaren't the smaller param models all just Qwen/Llama trained on R1 600bn?
- whimsicalism 2y agoyes, this is all ollamas fault
- john_alan 2y agoYeah I don’t understand why
- yetanotherjosh 2y agoollama is stating there's a difference: https://ollama.com/library/deepseek-r1 https://ollama.com/library/deepseek-r1 "including six dense models distilled from DeepSeek-R1 based on Llama and Qwen. " people just don't read? not sure there's reason to criticize ollama here.
- whimsicalism 2y agoi’ve seen so many people make this misunderstanding, huggingface clearly differentiates the model, and from the cli that isn’t visible
- whimsicalism 2y agoyou’re probably running it on ollama. ollama is doing the pretty unethical thing of lying about whether you are running r1, most of the models they have labeled r1 are actually entirely different models
- semicolon_storm 2y agoAre you referring to the distilled models?
- whimsicalism 2y agoyes, they are not r1
- BeefySwain 2y agoCan you explain what you mean by this?
- baobabKoodaa 2y agoFor example, the model named "deepseek-r1:8b" by ollama is not a deepseek r1 model. It is actually a fine tune of Meta's Llama 8b, fine tuned on data generated by deepseek r1.
- ekam 2y agoIf you’re referring to what I think you’re referring to, those distilled models are from deepseek and not ollama https://github.com/deepseek-ai/DeepSeek-R1 https://github.com/deepseek-ai/DeepSeek-R1
- whimsicalism 2y agothe choice on naming convention is ollama's, DS did not upload to huggingface that way
- strangescript 2y ago* Yes I am aware I am not running R1, and I am running a distilled version of it. If you have experience with tiny ~1B param models, its still head and shoulders above anything that has come before. IMO there have not been any other quantized/distilled/etc models as good at this size. It would not exist without the original R1 model work.