4 ms·
It better be! AMD @ $2k: https://digitalspaceport.com/how-to-run-deepseek-r1-671b-fully-locally-on-2000-epyc-rig/ https://digitalspaceport.com/how-to-run-deepse
by ynniv 2y ago
It better be! AMD @ $2k: https://digitalspaceport.com/how-to-run-deepseek-r1-671b-fully-locally-on-2000-epyc-rig/ https://digitalspaceport.com/how-to-run-deepseek-r1-671b-ful...
- utopcell 2y agoWhat a teaser article! All this info for setting up the system, but no performance numbers.
- yvdriess 2y agoThat's because the OP is linking to the quickstart guide. There are benchmark numbers on the github's root page, but it does not appear to include the new deepseek yet: https://github.com/intel/ipex-llm/tree/main?tab=readme-ov-file#ipex-llm-performance https://github.com/intel/ipex-llm/tree/main?tab=readme-ov-fi...
- utopcell 2y agoAm I missing something ? I see a lot of the small-scale models results but no results for DeepSeek-R1-671B-Q4_K_M on their github repos.
- aurareturn 2y agoThis article keeps getting posted but it runs a thinking model at 3-4 tokens/s. You might as well take a vacation if you ask it a question. It’s a gimmick and not a real solution.
- hnuser123456 2y agoIf you value local compute and don't need massive speed, that's still twice as fast as most people can type.
- aurareturn 2y agoHuman typing speed is magnitudes slower than our eyes scanning for the correct answer. ChatGPT o3 mini high thinks at about 140 tokens/s by my estimation and I sometimes wish it can return answers quicker. Getting a simple prompt answer would take 2-3 minutes using the AMD system and forget about longer context.
- evilduck 2y agoReasoning models spend a whole bunch of time reasoning before returning an answer. I was toying with QWQ 32B last night and ran into one question I gave it where it spent 18 minutes at 13tok/s in the <think> phase before returning a final answer. I value local compute but reasoning models aren’t terribly feasible at this speed since you don’t really need to see the first 90% of their thinking output.
- walrus01 2y agoIt's meant to be a test/development setup for people to prepare the software environment and tooling for running the same on more expensive hardware. Not to be fast.
- aurareturn 2y agoI remember people trying to run the game Crysis using CPU rendering. They got it to run and move around. People did it for fun and the "cool" factor. But no one actually played the game that way. It's the same thing here. CPUs can run it but only as a gimmick.
- refulgentis 2y ago> It's the same thing here. CPUs can run it but only as a gimmick. No, that's not true. I work on local inference code via llama.cpp, on both GPU and CPU on every platform, and the bottleneck is much more ram / bandwidth than compute. Crappy Pixel Fold 2022 mid-range Android CPU gets you roughly same speed as 2024 Apple iPhone GPU, with Metal acceleration that dozens of very smart people hack on. Additionally, and perhaps more importantly, Arc is a GPU, not a CPU. The headline of the thing you're commenting on, the very first thing you see when you open it, is "Run llama.cpp Portable Zip on Intel GPU" Additionally, the HN headline includes "1 or 2 Arc 7700"
- aurareturn 2y agoIt's both compute and bandwidth constrained - just like trying to run Crysis on CPU rendering. A770 has 16GB of RAM. You're shuffling data to the GPU at a rate of 64GB/s, which is magnitudes slower than the internal VRAM of the GPU. Hence, this setup is memory bandwidth constrained. However, once you want to use it to do anything useful like a longer context size, the CPU compute will be a huge bottleneck for time-to-first-token as well as tokens/s. Trying to run a model this large, and a thinking one at that, on CPU RAM is a gimmick.
- refulgentis 2y ago
- miklosz 2y agoExactly! I run it on my old T7910 Dell workstation (2x 2697A V4, 640GB RAM) that I build for way less than a $1k. But so what, it's about ~2 tokens / s. Just like you said, it's cool that it's run at all, but that's it.