6 ms·
Out of sheer curiosity: What’s required for the average Joe to use this, even at a glacial pace, in terms of hardware? Or is it even possible without using smar
by transcriptase 1y ago
Out of sheer curiosity: What’s required for the average Joe to use this, even at a glacial pace, in terms of hardware? Or is it even possible without using smart person magic to append enchanted numbers and make it smaller for us masses?
- jacob019 1y agoI'm sure it will be on OpenRouter within the next day or so. Not really practical to run a 685B param model at home.
- behohippy 1y agoAbout 768 gigs of ddr5 RAM in a dual socket server board with 12 channel memory and an extra 16 gig or better GPU for prompt processing. It's a few grand just to run this thing at 8-10 tokens/s
- wongarsu 1y agoAbout $8000 plus the GPU. Let's throw in a 4080 for about $1k, and you have the full setup for the price of 3 RTX5090. Or cheaper than a single A100. That's not a bad deal. For the hobby version you would presumably buy a used server and a used GPU. DDR4 ECC Ram can be had for a little over $1/GB, so you could probably build the whole thing for around $2k
- phonon 1y agoThis is the state of the art for such a setup. Really good performance! https://github.com/kvcache-ai/ktransformers https://github.com/kvcache-ai/ktransformers
- JKCalhoun 1y agoBeen putting together a "mining rig" [1] (or rather I was before the tariffs, ha ha.) Going to try to add a 2nd GPU soon. (And I should try these quantized versions.) Mobo was some kind of mining rig from AliExpress for less than $100. GPU is an inexpensive NVIDIA TESLA card that I 3D printed a shroud for (added fans). Power supply a cheap 2000 Watt Dell server PS off eBay.... [1] https://bsky.app/profile/engineersneedart.com/post/3lmg4kiz4fk2v https://bsky.app/profile/engineersneedart.com/post/3lmg4kiz4...
- terhechte 1y agoYou can run the 4bit quantized version of it on a M3 Ultra 512GB. That's quite expensive though. Another alternative is a fast CPU with 500GB of DDR5 RAM. That of course, is also not cheap and slower than the M3 Ultra. Or, you buy multiple Nvidia cards to reach ~500GB of VRam. That is probably the most expensive option but also the fastest
- lodovic 1y agoIf you use the excess memory for AI only it's cheaper to rent . A single H100 costs less than $2 per hour. (incl power)
- diggan 1y agoVast.ai has a bunch of 1x H100 SXM available, right now the cheapest at $1.554/hr. Not affiliated, just a (mostly) happy user, although don't trust the bandwidth numbers, lots of variance (not surprising though, it is a user-to-user marketplace).
- qingcharles 1y agoEvery time someone asks me what hardware to buy to run these at home I show them how many thousands of hours at vast.ai you could get for the same cost. I don't even know how these Vast servers make money because there is no way you can ever pay off your hardware from the pennies you're getting.
- omneity 1y agoWorth mentioning that a single H100 (80-96GB) is not enough to run R1. You're looking at 6-8 GPUs on the lower end, and factor in the setup and download time. An alternative is to use serverless GPU or LLM providers which abstract some of this for you, albeit at a higher cost and slow starts when you first use your model for some time.
- girvo 1y agoIt is enough to run the dynamically quantised 1.56 bit version I believe, which is fun to play around with.
- deleted 1y ago[deleted]
- hu3 1y agoIt's probably going to be free at OpenRouter. There's already a 685B parameter DeepSeek V3 for free there. https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free
- latchkey 1y agoIt is free to use, but you're feeding OR data and someone is profiting off that.
- ankit219 1y agoThats how a lot of application layer startups are going to make money. There is a bunch of high quality usage data. Either you monetize it yourself (cursor), get acquired (windsurf) or provide that data to others at a fee (lmsys, mercor). This is inevitable and a market for this is just going to increaase. If you want to prevent this as an org, there arent many ways out. Either use open source models you can deploy, or deal directly with model providers where you can sign specific contracts.
- inquirerGeneral 1y ago[dead]
- 85392_school 1y agoYou're actually sending data to random GPUs connected to one of the Bittensor subnets that run LLMs.
- latchkey 1y agoThat can, today, collect that data and sell it. There is work being done to add TEE, but it isn't live yet.
- dist-epoch 1y agoNot every prompt is privacy sensitive. For example you could use it to summarize a public article.
- hadlock 1y agoAs mentioned you can run this on a server board with 768+ gb memory in cpu mode. Average joe is going to be running quantized 30b (not 600b+) models on an $300/$400/$900 8/12/16gb GPU
- rahimnathwani 1y agoI'm not sure that's enough RAM to run it at full precision (FP8). This guy ran a 4-bit quantized version with 768GB RAM: https://news.ycombinator.com/item?id=42897205 https://news.ycombinator.com/item?id=42897205
- SkyPuncher 1y agoPractically, smaller, quantized versions of R1 can be run on a pretty typically Macbook Pro setup. Quantized versions are definitely less performant, but they will absolutely run. Truthfully, it's just not worth it. You either run these things so slowly that you're wasting your time or you have to buy 4- or 5-figures of hardware that's going to sit, mostly unused.
- danielhanchen 1y agoWe made DeepSeek R1 run on a local device via offloading and 1.58bit quantization :) https://unsloth.ai/blog/deepseekr1-dynamic https://unsloth.ai/blog/deepseekr1-dynamic I'm working on the new one!
- screaminghawk 1y agoI use this a lot! Thanks for your work and looking forward to the next one
- danielhanchen 1y agoThank you!! New versions should be much better!
- CamperBob2 1y agoYour 1.58-bit dynamic quant model is a religious experience, even at one or two tokens per second (which is what I get on my 128 MB Raptor Lake+4090). It's like owning your own genie... just ridiculously smart. Thanks for the work you've put into it!
- danielhanchen 1y agoOh thank you! :) Glad they were useful!
- nxobject 1y agoLikewise - for me, it feels how I imagined getting a microcomputer in the 70s was like. (Including the hit to the wallet… an Apple II cost the 2024 equivalent of ~$5k, too.)
- danielhanchen 1y ago:) The good ol days!
- behnamoh 1y ago> 1.58bit quantization of course we can run any model if quantize it enough. but I think the OP was talking about the unquantized version.
- mechagodzilla 1y agoI have a $2k used dual-socket xeon with 768GB of DDR4 - It runs at about 1.5 tokens/sec for the 4-bit quantized version.
- threeducks 1y ago> even at a glacial pace If speed is truly not an issue, you can run Deepseek on pretty much any PC with a large enough swap file, at a speed of about one token every 10 minutes assuming a plain old HDD. Something more reasonable would be a used server CPU with as many memory channels as possible and DDR4 ram for less than $2000. But before spending big, it might be a good idea to rent a server to get a feel for it.
- z2 1y agoHardware: any computer from the last 20 or so years. Software: client of choice to https://openrouter.ai/deepseek/deepseek-r1-0528 https://openrouter.ai/deepseek/deepseek-r1-0528 Sorry I'm being cheeky here, but realistically unless you want to shell out 10k for the equivalent of a Mac Studio with 512GB of RAM, you are best using other services or a small distilled model based on this one.
- jazzyjackson 1y agoYou can pay Amazon to do it for you at about a penny per 10 thousand tokens. There's a couple of guides for setting it up "manually" on ec2 instances so you're not paying the Bedrock per-token-prices, here's [1] that states four g6e.48xlarge instances (192 vCPUs, 1536GB RAM, 8x L40S Tensor Core GPUs that come with 48 GB of memory per GPU) Quick google tells me that g6e.48xlarge is something like 22k USD per month? [0] https://aws.amazon.com/bedrock/deepseek/ https://aws.amazon.com/bedrock/deepseek/ [1] https://community.aws/content/2w2T9a1HOICvNCVKVRyVXUxuKff/deploying-the-deepseek-v3-model-full-version-in-amazon-eks-using-vllm-and-lws https://community.aws/content/2w2T9a1HOICvNCVKVRyVXUxuKff/de...
- whynotmaybe 1y agoI'm using GPT4All with DeepSeek-R1-Distill-QWen-7B (which is not R1-0528) on a Ryzen 5 3600 with 32Gb ram. With an average of 3.6 tokens/sec, answers usually take 150-200 seconds.