9 ms·
DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
- deleted 2y ago[deleted]
- superkuh 2y agoNo... this headline is incorrect. You can't do that. I think they've confused the performance of running one of the small distills to existing smaller models. Two Arc cards cannot fit a 4 bit k-quant of a 671b model. But a portable (no install) way to run llama.cpp on intel GPUs is really cool.
- rgbrgb 2y agoyep, title is inaccurate. it's a distill into Qwen 7B DeepSeek-R1-Distill-Qwen-7B-Q4_K_M.gguf
- zamadatix 2y agoThe document contains multiple sections. The initial section does reference DeepSeek-R1-Distill-Qwen-7B-Q4_K_M.gguf as the example model but if you continue reading further you'll see a section referencing running DeepSeek-R1-Q4_K_M.gguf plus claims several other variations have been tested. It's a bit less exciting when you see they're just talking about offloading parts from the large amount of DRAM.
- genewitch 2y agoSo you thought there was some magical way to get >600B parameters in a couple of GPUs? Also, LM studio lets you run smaller models in front of larger ones, so I could see having a few GPU in front really speeding up using R1 for inference.
- hmottestad 2y agoThe MoE architecture allows you to keep the entire active model on a single GPU. If two consecutive tokens use the same export then the second token is going to be much faster.
- utopcell 2y agoWhat is the probability of that happening?
- zamadatix 2y agoDeepSeek V3/R1 uses 8 routed experts out of 256, so not all as often as one would like. That said, having even just a single GPU will greatly speed up prompt processing which is worth it even if the inference speed was the same. Ktransformers has a document about using CPU + a single 4090D to reach decent tokens/s but I'm not sure how much of the perf is due to the 4090D vs other optimizations/changes for the CPU side https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/DeepseekR1_V3_tutorial.md https://github.com/kvcache-ai/ktransformers/blob/main/doc/en... The final step of going to 6 experts instead of 8 feels like cheating (not a lossless optimization).
- genewitch 2y agowhere does 256 come from? it's repeated in here and elsewhere that a single expert is 37B sized, so you'd have to have way more than "several hundred billion parameters", to hold 256 of those? Maybe i don't understand the architecture, but if that's the case, then everyone repeating 37B doesn't, either.
- zamadatix 2y agoI think this diagram from the DeepSeekMoE paper explains it the clearest: https://i.imgur.com/CRKttob.png https://i.imgur.com/CRKttob.png The one on the right is how the feed forward layers of DeepSeek V3/R1 work, blue and green are experts, and everything in that right section is what counts as "active parameters". K (K=8 for these models, but you can customize that if you want) experts of 256 per layer are activated at a time. The 256 comes from the model file, it's just how many they chose to build it with. In these models there is also 1 shared expert which is always active in the layer. The router picks which k routed experts to use each forward pass and then a gating mechanism combines the outputs. If you sum the 1 shared expert + K routed experts + router + output networks you end up with 37 B parameters active for each feed forward layer pass. The individual experts are therefore much smaller than the total (probably something like 4 B parameters each? I've never really checked that directly). Or, for the short answer: "37 B is the active parameters of 9 experts + 'overhead', not the parameters of a single expert".
- colorant 2y agoSee this section https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quickstart/llamacpp_portable_zip_gpu_quickstart.md#flashmoe-for-deepseek-v3r1 https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic...
- Cheer2171 2y agoYou don't have to go that far down the page to see it is paging to system RAM: Requirements: 380GB CPU Memory 1-8 ARC A770 500GB Disk
- superkuh 2y agoYep. That's why the headline is incorrect. 380GB of the model on CPU system RAM and 32GB on some ARC GPUs. The ratio, 380/32, is obvious. Most of the processing is being done on the CPU. The GPU are little bit icing in this context. Fast, sure, but having to wait for the CPU layers (that's how layer splits work with llama.cpp). I think changing the end of headline to "Xeon w/380GB RAM" would stop it from being incorrect and misleading.
- Cheer2171 2y ago"with" does not mean "entirely on" Edit: but what you added in your edit is right, it would be more accurate to append the system ram requirement
- ryao 2y agoWhat if it does not need to read from system RAM for every token by reusing experts whenever they just happen to be in VRAM from being used for the previous token? If the selected experts do not change often, this is doable on paper.
- hmottestad 2y agoThat’s probably the main performance benefit of using the GPU. If you’re changing the active expert for every single token then it wouldn’t be any faster than just running it on the CPU. Once you can reuse the active expert for two tokens you’re already going to be a lot faster than just the CPU. More GPUs let you keep more experts active at a time.
- hexaga 2y agoExpert distribution should be approximately random token-by-token, so not likely.
- deleted 2y ago[deleted]
- ryao 2y agoIt is theoretically possible. Each token only needs 37B parameters and if the same experts are chosen often, it would behave closer to a 37B model than a 671B model, since reusing experts can skip loads from system RAM. You might still be right since I have not confirmed that the selected experts change infrequently doing prompt processing / token generation, and someone could have botched the headline. However, treating Deepseek like llama 3 when reasoning about VRAM requirements is not necessarily correct.
- superkuh 2y agoMoE is pretty enabling after you've spent all the extra $$$$ to stuff your server CPU memory channels with ram so it's possible to run at all. But it's still spending a lot of money which makes this a lot less novel or interesting than "just on 1~2 Arc A770" implies. Especially for the marginal performance that even 8-12 channels of CPU memory bandwidth gets you.
- utopcell 2y agoIs this amount of RAM really that expensive? 6x 64GiB DDR4 DIMMs are < $1,000.
- utopcell 2y agoActually, 384GiB is already <$400 [1]. [1] https://www.amazon.com/NEMIX-RAM-DDR4-2666MHz-PC4-21300-Reduced/dp/B0BJFN1JYS https://www.amazon.com/NEMIX-RAM-DDR4-2666MHz-PC4-21300-Redu...
- superkuh 2y agoA slow-end DDR4 speed older generation Xeon system is unlikely to be used by Intel for this benchmark. It's far more likely they used an expensive DDR5 modern Xeon with as many memory channels as they could. Single user LLM inference is memory bandwidth bottlenecked. I just can't see Intel using old/deprecated hardware. And if someone not Intel were to build a Xeon DDR4 system it wouldn't reach the DDR5 tokens/s speeds reported here. The reason they used a Xeon is memory channels. Non-server CPUs only have 2 but modern Xeons have 8 to 12 depending on generation/type. And the Xeons with the most are the most $$$$ and it ends up cheaper to just get a GPU or dedicated accelerator.
- deleted 2y ago[deleted]
- ryao 2y agoWhere is the benchmark data?
- zamadatix 2y agoSince the Xeon alone could run the model in this set up it'd be more interesting if they compared the performance uplift with using 0/1/2..8 Arc A770 GPUs. Also, it's probably better to link straight to the relevant section https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quickstart/llamacpp_portable_zip_gpu_quickstart.md#flashmoe-for-deepseek-v3r1 https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic...
- colorant 2y agoYes, you are right. Unfortunately HN somehow truncated my original URL link.
- zamadatix 2y agoSounds like submission "helper" tools are working about as well as normal :). Did you have the chance to try this out yourself or did you just run across it recently?
- hmottestad 2y agoIf you’re running just one GPU your context is limited to 1024 tokens, as far as I could tell. I couldn’t see what the context size is for more cards though.
- colorant 2y agohttps://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quickstart/llamacpp_portable_zip_gpu_quickstart.md#flashmoe-for-deepseek-v3r1 https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic... Requirements (>8 token/s): 380GB CPU Memory 1-8 ARC A770 500GB Disk
- colorant 2y agoAlso see the demo from Jason Dai's post: https://www.linkedin.com/posts/jasondai_with-the-latest-ipex-llm-llamacpp-portable-activity-7303194182729244673-FcxL https://www.linkedin.com/posts/jasondai_with-the-latest-ipex...
- aurareturn 2y agoCPU inference is both bandwidth and compute constrained. If your prompt has 10 tokens, it’ll do ok, like in the LinkedIn demo. If you need to increase the context, compute bottleneck will kick in quickly.
- colorant 2y agoPrompt length mainly impacts prefill latency (FTFF), not the decoding speed (TPOT)
- moffkalast 2y agoDecoding speed won't matter one bit if you have to sit there for 5 minutes waiting for the model to ingest a prompt that's two sentences long.
- colorant 2y agoWith ~1000 input, the TTFT is ~10 seconds
- faizshah 2y agoAnyone got a rough estimate of the cost of this setup? I’m guessing it’s under 10k. I also didn’t see tokens per second numbers.
- jamesy0ung 2y agoWhat exactly does the Xeon do in this situation, is there a reason you couldn't use any other x86 processor?
- VladVladikoff 2y agoI think it’s that most non Xeon motherboards don’t have the memory channels to have this much memory with any sort of commercially viable dimms.
- genewitch 2y agoPcie lanes
- hedora 2y agoI was about to correct you because this doesn't use PCIe for anything, and then I realized Arc was a GPU (and they support up to 8 per machine). Any idea how many Arc's it takes to match an H100?
- npodbielski 2y agoI am reading from time to time about multi GPU solution and last time I found some real life information about this (it was two 7900 xtx) the result was that performance is the same at best often it is slower. So even if you manage to slap like 8 cheap cards onto motherboard, even if you would somehow make it work (people have problems with such setups), even if this would work continuously without much problems (crashes, power consumption) performance would be just OK. I am not sure if spending 10k on such setup would be better than buying 10k card with 40gbs of RAM.
- pshirshov 2y agoOllama works fine with multi-gpu setups. Since rocm 6.3 everything is stable and you can mix different GPU generations. The performance is good enough for the models to be useful. The only thing which doesn't work well is running on iGPUs. It might work but it's very unstable.
- deleted 2y ago[deleted]
- 7speter 2y agoI’ve been following the progress Intel Arc support in Pytorch is making, at least in Linux, and it seems like if things stay on track, we may see the first version of pytorch with full Xe/Arc support by around June. I think I’m just going to wait until then instead of dealing with anything ipex or openvino.
- colorant 2y agoThis is based on llama.cpp
- deleted 2y ago[deleted]
- CamperBob2 2y agoArticle could stand to include a bit more information. Why are all the TPS figures x'ed out? What kind of performance can be expected from this setup (and how does it compare to the dual Epyc workstation recipe that was popularized recently?)
- codetrotter 2y ago> the dual Epyc workstation recipe that was popularized recently Anyone have a link to this one?
- hedora 2y agohttps://news.ycombinator.com/item?id=42897205 https://news.ycombinator.com/item?id=42897205
- colorant 2y ago>8TPS at this moment on a 2-socket 5th Xeon (EMR)
- anacrolix 2y agoNow we just need a model that can actually code
- ohgr 2y agoI'll settle with a much lower bar: an engineer that can tell the code the model generates is shit.
- brokegrammer 2y agoMost engineers can do that because it's way easier to find flaws in code you didn't write vs in ones that you write. My code is always perfect in my own eyes until someone else sees it.
- ohgr 2y agoFrom experience, most engineers can do neither.
- yongjik 2y agoDid DeepSeek learn how to name their models from OpenAI.
- vlovich123 2y agoThe convention is weird but it's pretty standard in the industry across all models, particularly GGUF. 671B parameters, quantized to 4 bits. The K_M terminology I believe is more specific to GGUF and describes the specific quantization strategy.
- chriscappuccio 2y agoBetter to run the Q8 model on an epyc pair with 768GB, you'll get the same performance
- ltbarcly3 2y agoThe Q8 model is totally different?
- manmal 2y agoMy experience with quantizations is that anything below 6 is noticeably worse. Coherence suffers. I’ve rarely gotten anything really useful out of a Q4 model, code wise. For transformations they are great though, eg convert JSON to Markdown and vice versa.
- hnfong 2y agoAs other commenters have mentioned, the performance of this set up is probably not really great since there's not enough VRAM and lots of bits have to be moved between CPU and GPU RAM. That said, there are sub-256GB quants of DeepSeek-R1 out there (not the distilled versions). See https://unsloth.ai/blog/deepseekr1-dynamic https://unsloth.ai/blog/deepseekr1-dynamic I can't quantify the difference between these and the full FP8 versions of DSR1, but I've been playing with these ~Q2 quants and they're surprisingly capable in their own right. Another model that deserves mention is DeepSeek v2.5 (which has "fewer" params than V3/R1) - but still needs aggressive quantization before it can run on "consumer" devices (with less than ~100GB VRAM), and this is recently done by a kind soul: https://www.reddit.com/r/LocalLLaMA/comments/1irwx6q/deepseekv25_dynamic_quants_anyone/ https://www.reddit.com/r/LocalLLaMA/comments/1irwx6q/deepsee... DeepSeek v2.5 is arguably better than Llama 3 70B, so it should be of interest to anyone looking to run local inference. I really think more people should know about this.
- SlavikCA 2y agoI tried that Unsloth R1 quantization on my dual Xeon Gold 5218 with 384 GB DDR4-2666 (about half of memory channels used, so not most optimal). Type IQ2_XXS / 183GB, 16k context: CPU only: 3 t/s (tokens per second) for PP (prompt processing) and 1.44 t/s for response. CPU + NVIDIA RTX 70GB VRAM: 4.74 t/s for PP and 1.87 t/s for response. I wish Unsloth produce similar quantization for DeepSeek V3, - it will be more useful, as it doesn't need reasoning tokens, so even with same t/s it will faster overall.
- idonotknowwhy 2y agoThanks a lot for the v2.5! I'll give that a whirl. Hopefully it's as coherent as v3.5 when quantized so small. > I can't quantify the difference between these and the full FP8 versions of DSR1, but I've been playing with these ~Q2 quants and they're surprisingly capable in their own right. I run the Q2_K_XL and it's perfectly good for me. Where it lacks vs FP8 is in creative writing. If you prompt it with for a story a few times, then compare with FP8, you'll see what I mean. For coding, the 1.58bit clearly makes more errors than the Q2XXS and Q2_K_XL
- pinoy420 2y ago
- mrbonner 2y agoI see there are a few options to run inference for LLM and Stable Diffusion outside Nvidia. There is Intel Arc, Apple Ms and now AMD Ryzen AI Max. It is obvious that running in Nvidia would be the most optimal way. But given the availability of high VRAM Nvidia cards at reasonable price, I can't stop thinking about getting one that is not Nvidia. So, if I'm not interested in training or fine tuning, would any of those solutions actually works? On a Linux machine?
- 999900000999 2y agoIf you actually want to seriously do this, go with Nvidia. This article is basically Intel saying remember us, we made a GPU! And they make great budget cards, but the ecosystem is just so far behind. Honestly this is not something you can really do on a budget.
- Gravityloss 2y agoI'm sure this question has been asked before, but why not launch a GPU with more but slower ram? That would fit bigger models while still affordable...
- varelse 2y ago[dead]
- ChocolateGod 2y agoBecause then you would have less motivation to buy the more expensive GPUs.
- fleischhauf 2y agothey absolutely can build gpus with larger vram, they just don't have the competition to have to do so. it's much more profitable this way.
- TeMPOraL 2y agoWhy would you need it for? Not gaming for sure. AI, you say? Then fork up the cash. That's Nvidia's current MO. There's more demand for GPUs for AI than there are GPUs available, and most of that demand still has stupid amounts of money behind it (being able to get grants, loans or investment based on potential/hype) - money that can be captured by GPU vendors. Unfortunately, VRAM is the perfect discriminator between "casual" and "monied" use. (This is not unlike the "SSO tax" - single sign-on is pretty much the perfect discriminator between "enterprise use" and "not enterprise use".)
- notum 2y agoCensoring of token/s values in the sample output surely means this runs great!
- DeathArrow 2y agoAny chance of using a couple of 3090 with Deepseek and fit the whole thing in the video RAM? I'm thinking to something like a software or "fake" NVLink.
- andrewstuart 2y agoWith the arrival of APUs for AI everyone is going to lose interest in GPUs real fast. Why buy an overpriced Nvidia 4090 when you can get an AMD Halo Strix or Apple M3 Studio APU with 512GB or 128GB of Ram? Nvidia has kept prices high and performance low for as long as it can and finally competition is here. Even Intel can make APUs with tons of RAM. Nvidia hopefully is squirming.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]