4 ms·
Truly mind-blowing to see Mixtral 47B running on a smartphone.
by helloericsf 2y ago
Truly mind-blowing to see Mixtral 47B running on a smartphone.
- westurner 2y ago> TL;DR PowerInfer is a CPU/GPU LLM inference engine leveraging activation locality for your device.
- wmf 2y agoSo is it running on the phone or not?
- westurner 2y agoFrom the abstract https://arxiv.org/abs/2406.06282 https://arxiv.org/abs/2406.06282 : > Notably, PowerInfer-2 is the first system to serve the TurboSparse-Mixtral-47B model with a generation rate of 11.68 tokens per second on a smartphone. For models that fit entirely within the memory, PowerInfer-2 can achieve approximately a 40% reduction in memory usage while maintaining inference speeds comparable to llama.cpp and MLC-LLM What does this do with a 40 TOPS+ NPU/TPU?