4 ms·
https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quickstart/llamacpp_portable_zip_gpu_quickstart.md#flashmoe-for-deepseek-v3r1 https://github.com/intel/i
by colorant 2y ago
https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quickstart/llamacpp_portable_zip_gpu_quickstart.md#flashmoe-for-deepseek-v3r1 https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic...
Requirements (>8 token/s):
380GB CPU Memory
1-8 ARC A770
500GB Disk
- colorant 2y agoAlso see the demo from Jason Dai's post: https://www.linkedin.com/posts/jasondai_with-the-latest-ipex-llm-llamacpp-portable-activity-7303194182729244673-FcxL https://www.linkedin.com/posts/jasondai_with-the-latest-ipex...
- aurareturn 2y agoCPU inference is both bandwidth and compute constrained. If your prompt has 10 tokens, it’ll do ok, like in the LinkedIn demo. If you need to increase the context, compute bottleneck will kick in quickly.
- colorant 2y agoPrompt length mainly impacts prefill latency (FTFF), not the decoding speed (TPOT)
- moffkalast 2y agoDecoding speed won't matter one bit if you have to sit there for 5 minutes waiting for the model to ingest a prompt that's two sentences long.
- colorant 2y agoWith ~1000 input, the TTFT is ~10 seconds
- faizshah 2y agoAnyone got a rough estimate of the cost of this setup? I’m guessing it’s under 10k. I also didn’t see tokens per second numbers.
- ynniv 2y agoIt better be! AMD @ $2k: https://digitalspaceport.com/how-to-run-deepseek-r1-671b-fully-locally-on-2000-epyc-rig/ https://digitalspaceport.com/how-to-run-deepseek-r1-671b-ful...
- utopcell 2y agoWhat a teaser article! All this info for setting up the system, but no performance numbers.
- yvdriess 2y agoThat's because the OP is linking to the quickstart guide. There are benchmark numbers on the github's root page, but it does not appear to include the new deepseek yet: https://github.com/intel/ipex-llm/tree/main?tab=readme-ov-file#ipex-llm-performance https://github.com/intel/ipex-llm/tree/main?tab=readme-ov-fi...
- utopcell 2y agoAm I missing something ? I see a lot of the small-scale models results but no results for DeepSeek-R1-671B-Q4_K_M on their github repos.
- aurareturn 2y agoThis article keeps getting posted but it runs a thinking model at 3-4 tokens/s. You might as well take a vacation if you ask it a question. It’s a gimmick and not a real solution.
- hnuser123456 2y ago
- deleted 2y ago[deleted]
- GTP 2y ago> 1-8 ARC A770 To get more than 8 t/s, is one Intel Arc A770 enough?
- colorant 2y agoYes, but the context length will be limited due to VRAM constraint