9 ms·
Deepseek R1-0528
- dubrado 1y ago[flagged]
- transcriptase 1y agoOut of sheer curiosity: What’s required for the average Joe to use this, even at a glacial pace, in terms of hardware? Or is it even possible without using smart person magic to append enchanted numbers and make it smaller for us masses?
- jacob019 1y agoI'm sure it will be on OpenRouter within the next day or so. Not really practical to run a 685B param model at home.
- behohippy 1y agoAbout 768 gigs of ddr5 RAM in a dual socket server board with 12 channel memory and an extra 16 gig or better GPU for prompt processing. It's a few grand just to run this thing at 8-10 tokens/s
- wongarsu 1y agoAbout $8000 plus the GPU. Let's throw in a 4080 for about $1k, and you have the full setup for the price of 3 RTX5090. Or cheaper than a single A100. That's not a bad deal. For the hobby version you would presumably buy a used server and a used GPU. DDR4 ECC Ram can be had for a little over $1/GB, so you could probably build the whole thing for around $2k
- phonon 1y agoThis is the state of the art for such a setup. Really good performance! https://github.com/kvcache-ai/ktransformers https://github.com/kvcache-ai/ktransformers
- JKCalhoun 1y agoBeen putting together a "mining rig" [1] (or rather I was before the tariffs, ha ha.) Going to try to add a 2nd GPU soon. (And I should try these quantized versions.) Mobo was some kind of mining rig from AliExpress for less than $100. GPU is an inexpensive NVIDIA TESLA card that I 3D printed a shroud for (added fans). Power supply a cheap 2000 Watt Dell server PS off eBay.... [1] https://bsky.app/profile/engineersneedart.com/post/3lmg4kiz4fk2v https://bsky.app/profile/engineersneedart.com/post/3lmg4kiz4...
- terhechte 1y agoYou can run the 4bit quantized version of it on a M3 Ultra 512GB. That's quite expensive though. Another alternative is a fast CPU with 500GB of DDR5 RAM. That of course, is also not cheap and slower than the M3 Ultra. Or, you buy multiple Nvidia cards to reach ~500GB of VRam. That is probably the most expensive option but also the fastest
- lodovic 1y agoIf you use the excess memory for AI only it's cheaper to rent . A single H100 costs less than $2 per hour. (incl power)
- diggan 1y agoVast.ai has a bunch of 1x H100 SXM available, right now the cheapest at $1.554/hr. Not affiliated, just a (mostly) happy user, although don't trust the bandwidth numbers, lots of variance (not surprising though, it is a user-to-user marketplace).
- qingcharles 1y agoEvery time someone asks me what hardware to buy to run these at home I show them how many thousands of hours at vast.ai you could get for the same cost. I don't even know how these Vast servers make money because there is no way you can ever pay off your hardware from the pennies you're getting.
- omneity 1y agoWorth mentioning that a single H100 (80-96GB) is not enough to run R1. You're looking at 6-8 GPUs on the lower end, and factor in the setup and download time. An alternative is to use serverless GPU or LLM providers which abstract some of this for you, albeit at a higher cost and slow starts when you first use your model for some time.
- girvo 1y agoIt is enough to run the dynamically quantised 1.56 bit version I believe, which is fun to play around with.
- deleted 1y ago[deleted]
- hu3 1y agoIt's probably going to be free at OpenRouter. There's already a 685B parameter DeepSeek V3 for free there. https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free
- latchkey 1y agoIt is free to use, but you're feeding OR data and someone is profiting off that.
- ankit219 1y agoThats how a lot of application layer startups are going to make money. There is a bunch of high quality usage data. Either you monetize it yourself (cursor), get acquired (windsurf) or provide that data to others at a fee (lmsys, mercor). This is inevitable and a market for this is just going to increaase. If you want to prevent this as an org, there arent many ways out. Either use open source models you can deploy, or deal directly with model providers where you can sign specific contracts.
- inquirerGeneral 1y ago[dead]
- 85392_school 1y agoYou're actually sending data to random GPUs connected to one of the Bittensor subnets that run LLMs.
- latchkey 1y agoThat can, today, collect that data and sell it. There is work being done to add TEE, but it isn't live yet.
- dist-epoch 1y agoNot every prompt is privacy sensitive. For example you could use it to summarize a public article.
- hadlock 1y agoAs mentioned you can run this on a server board with 768+ gb memory in cpu mode. Average joe is going to be running quantized 30b (not 600b+) models on an $300/$400/$900 8/12/16gb GPU
- rahimnathwani 1y agoI'm not sure that's enough RAM to run it at full precision (FP8). This guy ran a 4-bit quantized version with 768GB RAM: https://news.ycombinator.com/item?id=42897205 https://news.ycombinator.com/item?id=42897205
- SkyPuncher 1y agoPractically, smaller, quantized versions of R1 can be run on a pretty typically Macbook Pro setup. Quantized versions are definitely less performant, but they will absolutely run. Truthfully, it's just not worth it. You either run these things so slowly that you're wasting your time or you have to buy 4- or 5-figures of hardware that's going to sit, mostly unused.
- danielhanchen 1y agoWe made DeepSeek R1 run on a local device via offloading and 1.58bit quantization :) https://unsloth.ai/blog/deepseekr1-dynamic https://unsloth.ai/blog/deepseekr1-dynamic I'm working on the new one!
- screaminghawk 1y agoI use this a lot! Thanks for your work and looking forward to the next one
- danielhanchen 1y agoThank you!! New versions should be much better!
- CamperBob2 1y agoYour 1.58-bit dynamic quant model is a religious experience, even at one or two tokens per second (which is what I get on my 128 MB Raptor Lake+4090). It's like owning your own genie... just ridiculously smart. Thanks for the work you've put into it!
- danielhanchen 1y agoOh thank you! :) Glad they were useful!
- nxobject 1y agoLikewise - for me, it feels how I imagined getting a microcomputer in the 70s was like. (Including the hit to the wallet… an Apple II cost the 2024 equivalent of ~$5k, too.)
- danielhanchen 1y ago:) The good ol days!
- behnamoh 1y ago> 1.58bit quantization of course we can run any model if quantize it enough. but I think the OP was talking about the unquantized version.
- mechagodzilla 1y agoI have a $2k used dual-socket xeon with 768GB of DDR4 - It runs at about 1.5 tokens/sec for the 4-bit quantized version.
- threeducks 1y ago> even at a glacial pace If speed is truly not an issue, you can run Deepseek on pretty much any PC with a large enough swap file, at a speed of about one token every 10 minutes assuming a plain old HDD. Something more reasonable would be a used server CPU with as many memory channels as possible and DDR4 ram for less than $2000. But before spending big, it might be a good idea to rent a server to get a feel for it.
- z2 1y agoHardware: any computer from the last 20 or so years. Software: client of choice to https://openrouter.ai/deepseek/deepseek-r1-0528 https://openrouter.ai/deepseek/deepseek-r1-0528 Sorry I'm being cheeky here, but realistically unless you want to shell out 10k for the equivalent of a Mac Studio with 512GB of RAM, you are best using other services or a small distilled model based on this one.
- jazzyjackson 1y agoYou can pay Amazon to do it for you at about a penny per 10 thousand tokens. There's a couple of guides for setting it up "manually" on ec2 instances so you're not paying the Bedrock per-token-prices, here's [1] that states four g6e.48xlarge instances (192 vCPUs, 1536GB RAM, 8x L40S Tensor Core GPUs that come with 48 GB of memory per GPU) Quick google tells me that g6e.48xlarge is something like 22k USD per month? [0] https://aws.amazon.com/bedrock/deepseek/ https://aws.amazon.com/bedrock/deepseek/ [1] https://community.aws/content/2w2T9a1HOICvNCVKVRyVXUxuKff/deploying-the-deepseek-v3-model-full-version-in-amazon-eks-using-vllm-and-lws https://community.aws/content/2w2T9a1HOICvNCVKVRyVXUxuKff/de...
- whynotmaybe 1y agoI'm using GPT4All with DeepSeek-R1-Distill-QWen-7B (which is not R1-0528) on a Ryzen 5 3600 with 32Gb ram. With an average of 3.6 tokens/sec, answers usually take 150-200 seconds.
- jacob019 1y agoNot much to go off of here. I think the latest R1 release should be exciting. 685B parameters. No model card. Release notes? Changes? Context window? The original R1 has impressive output but really burns tokens to get there. Can't wait to learn more!
- willchen 1y agoI love how Deepseek just casually drops new updates (that deliver big improvements) without fanfare.
- hd4 1y agoOn the day Nvidia report earnings too. Pretty sure it's just a coincidence, bro.
- margorczynski 1y agoYeah the timing seems strange. Considering how much money will move hands based on those results this might be some kind of play to manipulate the market at least a bit.
- consumer451 1y agoI believe that they are funded by a hedge fund. So, there are no coincidences here.
- Maxatar 1y agoHow does releasing it today affect the market compared to releasing it last week?
- doctoboggan 1y agoHard to say exactly how it will affect the market, but IIRC when deepseek was first released Nvidia stock took a big hit as people realized that you could develop high performing LLMs without access to Nvidia hardware.
- jimmyl02 1y agoI thought the reaction was more so that you can train SOTA models without an extremely large quantity of hyper-expensive GPU clusters? But I would say that the reaction was probably vastly overblown as what Deepseek really showed was there are much more efficient ways of doing things (which can also be applied with even larger clusters). If this checkpoint is trained using non-Nvidia GPUs that would definitely be a much bigger situation but it doesn't seem like there has been any associated announcements.
- _lvbh 1y agoNo information to be found about it. Hopefully we get benchmarks soon. Reminds me of the days when Mistral would just tweet a torrent magnet link
- aibrother 1y agogetting a similar vibe yeah. given how adjacent they are, wouldn't be surprised if this was an intentional nod from DeepSeek
- swyx 1y agoi think usually deepseek posts a paper after a model release about a day later. no idea why they cant just wait a bit to coordinate stuff. bit messy in the news cycle.
- Destiner 1y agohonestly a power move. it's almost as if they don't care about creating a proper buzz.
- wyre 1y agoFrom what I understand, isn’t DeepSeek just a pet project from a Chinese hedge fund? They have much less reason to create a buzz compared to openAI, Anthropic, or Google.
- TeMPOraL 1y agoNone of those players you mention actually need to create a buzz. People will do it for them for free. DeepSeek joined this group after releasing R1. Despite constant protestations of hype among the tech crowd, GenAI really is big enough of a deal that new developments don't need to be pushed onto market; people are voluntarily seeking them out.
- wongarsu 1y agoOpenAI does a lot of work hyping themselves up and creating buzz around things they do or have a vague idea that they might try to do in the future. Not to make people aware of GenAI, but to make sure OpenAI continues to be perceived as the AI company. The company that leads and revolutionizes, with everyone just copying them and trying to match them. That perception is a significant part of their value and probably their biggest moat
- heyhuy 1y ago[flagged]
- airjason 1y ago[flagged]
- htrp 1y agoYou're gonna need at least 8 h100 80s for this....
- canergly 1y agoI want to see it in groq asap !
- porphyra 1y agoGroq doesn't even have any true deepseek models --- I thought they only had `deepseek-r1-distill-llama-70b` which was distilled onto llama 70b [1]. [1] https://console.groq.com/docs/models https://console.groq.com/docs/models
- jacob019 1y agoGroq has a weak selection of models, which is frustrating because their inference speed is insane. I get it though, selection + optimization = performance.
- sergiotapia 1y agothe only reason they are fast is because the models they host are severely quantized so i've heard.
- jacob019 1y agoHuh. I heard a podcast with the founder talking about their custom hardware, but quantization would explain it.
- christianqchung 1y agoQuantization alone does not explain it. It's mostly custom hardware[0]. [0] https://groq.com/the-groq-lpu-explained/ https://groq.com/the-groq-lpu-explained/
- behnamoh 1y agothey responded to my tweet last year and said they didn't quantize the models.
- jacob019 1y agoWell that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers https://openrouter.ai/deepseek/deepseek-r1-0528/providers May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.
- jazzyjackson 1y agoNo sign of what source material it was trained on though right? So open weight rather than reproducible from source. I remember there's a project "Open R1" that last I checked was working on gathering their own list of training material, looks active but not sure how far along they've gotten: https://github.com/huggingface/open-r1 https://github.com/huggingface/open-r1
- chrsw 1y agoBased on commit history Open R1 still active and they're still making progress. Long may it continue, it's an ambitious project.
- behnamoh 1y ago> No sign of what source material it was trained on though right? out of curiosity, does anyone do anything "useful" with that knowledge? it's not like people can just randomly train models..
- ToValueFunfetti 1y agoWould be useful for answering "is this novel or was it in the training data", but that's not typically what the point of open source is
- marci 1y agoWhen you're trully open source, you can make ethings like this: Today we introduce OLMoTrace, a one-of-a-kind feature in the Ai2 Playground that lets you trace the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace is a manifestation of Ai2’s commitment to an open ecosystem – open models, open data, and beyond. https://allenai.org/blog/olmotrace https://allenai.org/blog/olmotrace
- cesarvarela 1y agoAbout half the price of o4 mini high for not that much worse performance, interesting edit: most providers are offering a quantized version...
- mjcohen 1y agoDeepseek seems to be one of the few LLMs that run on a iPod Touch because of the older version of ios.
- titaniumtown 1y ago... What?
- cropcirclbureau 1y agoHey! You! You can't just say that and not explain. Come back.
- MrPowerGamerBR 1y agoIf I had to guess, they were talking about the DeepSeek iOS app: https://apps.apple.com/br/app/deepseek-assistente-de-ia/id6737597349 https://apps.apple.com/br/app/deepseek-assistente-de-ia/id67...
- deepsquirrelnet 1y agoI think it’s cool to see this kind of international participation in fierce tech competition. It’s exciting. It’s what I think capitalism should be. This whole “building moats” and buying competitors fascination in the US has gotten boring, obvious and dull. The world benefits when companies struggle to be the best.
- karencarits 1y agoWhat use cases are people using local LLMs for? Have you created any practical tools that actually increase your efficiency? I've been experimenting a bit but find it hard to get inspiration for useful applications
- jsemrau 1y agoI have a signal tracer that evaluates unusual trading volumes. Given those signals, my local agent receives news items through API to make an assessment what happens. This helps me tremendously. If I would do this through a remote app, I'd have to spend a several dollars per day. So I have this on existing hardware.
- karencarits 1y agoThank you, this is a great example!
- dyauspitr 1y agoDo you want to share it?
- sudomarcma 1y agoAny companies with any type of sensitive data will love to have anything to do with LLM done locally.
- thenameless7741 1y agoA recent example: a law firm hired this person [0] to build a private AI system for document summarization and Q&A. [0] https://xcancel.com/glitchphoton/status/1927682018772672950 https://xcancel.com/glitchphoton/status/1927682018772672950
- lvturner 1y agoAlso worth it for the speed of AI autocomplete in coding tools, the round trip to my graphics card is much faster than going out over the network.
- Lopezkatheryn 1y ago[dead]
- AJAlabs 1y ago671B parameters! Well, it doesn't look like I'll be running that locally.
- amy_petrik 1y agothere is a small community of people that do indeed run this locally. typically on CPU/RAM (lots and lots of RAM), insofar as that's cheaper than GPU(s).
- mickey475778 1y ago[dead]
- techlatest_net 1y ago[dead]
- danielhanchen 1y agoFor those interested, I made some 1 bit dynamic quants at https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF https://huggingface.co/unsloth/DeepSeek-R1-0528-GGUF 74% smaller 713GB to 185GB. Use the magic incantation -ot ".ffn_.*_exps.=CPU" to offload MoE layers to RAM, allowing non MoEs to fit < 24GB VRAM on 16K context! The rest sits in RAM & disk.