Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
colorant
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
FlashMoE: DeepSeek-R1 671B and Qwen3MoE 235B with 1~2 Intel B580 GPU in IPEX-LLM
(github.com)
2 points
by
colorant
1y ago
|
1 comments
2.
▲
by
colorant
1y ago
Hardware Requirements: - 380GB CPU memory for DeepSeek V3/R1 671B INT4 model - 128GB CPU memory for Qwen3MoE 235B INT4 model - 1-2 ARC A770 or B580 - 500GB Disk space
3.
▲
by
colorant
2y ago
With ~1000 input, the TTFT is ~10 seconds
4.
▲
by
colorant
2y ago
The ipex-llm implementation extends llama.cpp and includes additonal CPU-GPU hybrid optimizations for sparse MoE
5.
▲
by
colorant
2y ago
Currently >8 token/s; there is a demo in this post: https://www.linkedin.com/posts/jasondai_run-671b-deepseek-r1...
6.
▲
by
colorant
2y ago
Yes, but the context length will be limited due to VRAM constraint
7.
▲
by
colorant
2y ago
Prompt length mainly impacts prefill latency (FTFF), not the decoding speed (TPOT)
8.
▲
by
colorant
2y ago
This is based on llama.cpp
9.
▲
by
colorant
2y ago
>8TPS at this moment on a 2-socket 5th Xeon (EMR)
10.
▲
by
colorant
2y ago
See this section https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic...
11.
▲
by
colorant
2y ago
Also see the demo from Jason Dai's post: https://www.linkedin.com/posts/jasondai_with-the-latest-ipex...
12.
▲
by
colorant
2y ago
Yes, you are right. Unfortunately HN somehow truncated my original URL link.
13.
▲
by
colorant
2y ago
https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic... Requirements (>8 token/s): 380GB CPU Memory 1-8 ARC A770 500GB Disk
14.
▲
DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon
(github.com)
297 points
by
colorant
2y ago
|
112 comments
15.
▲
by
colorant
2y ago
llamafile cannot use Intel GPU (including integrated GPU on your PC)
16.
▲
IPEX-LLM Portable Zip for Ollama on Intel GPU
(github.com)
2 points
by
colorant
2y ago
|
2 comments
17.
▲
by
colorant
2y ago
There is ipex-llm support for Ollama on Intel GPU ( https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quic... )
18.
▲
BigDL-LLM: running LLM on your laptop using INT4
(github.com)
1 points
by
colorant
3y ago
|
1 comments
19.
▲
by
colorant
3y ago
BigDL-LLM is a library for running LLM (language language model) on your local laptop using INT4 with very low latency on CPU. (It is built on top of the excellent work of llama.cpp, gptq, bitsandbytes, etc., and supports any Hugging Face T
20.
▲
Seamlessly Scaling AI for Distributed Big Data
(medium.com)
1 points
by
colorant
6y ago
|
0 comments
21.
▲
Distributed TF/PyTorch/OpenVINO Inference with a Simple Pub/Sub API
(analytics-zoo.github.io)
3 points
by
colorant
7y ago
|
0 comments
22.
▲
by
colorant
8y ago
- Data wrangling and analysis using PySpark - Deep learning model development using TensorFlow or Keras - Distributed training/inference on Spark and BigDL - All within a single unified pipeline and in a user-transparent fashion!
23.
▲
Show HN: Analytics Zoo – Distributed TensorFlow and Keras on Apache Spark
(github.com)
2 points
by
colorant
8y ago
|
1 comments