3 ms·
Project Github link:https://github.com/SJTU-IPADS/PowerInfer https://github.com/SJTU-IPADS/PowerInfer
by helloericsf 2y ago
Project Github link:https://github.com/SJTU-IPADS/PowerInfer https://github.com/SJTU-IPADS/PowerInfer
- talldayo 2y agoThe video/litepaper they linked is also a great read: https://powerinfer.ai/v2/ https://powerinfer.ai/v2/
- helloericsf 2y agoTruly mind-blowing to see Mixtral 47B running on a smartphone.
- westurner 2y ago> TL;DR PowerInfer is a CPU/GPU LLM inference engine leveraging activation locality for your device.
- wmf 2y agoSo is it running on the phone or not?
- westurner 2y agoFrom the abstract https://arxiv.org/abs/2406.06282 https://arxiv.org/abs/2406.06282 : > Notably, PowerInfer-2 is the first system to serve the TurboSparse-Mixtral-47B model with a generation rate of 11.68 tokens per second on a smartphone. For models that fit entirely within the memory, PowerInfer-2 can achieve approximately a 40% reduction in memory usage while maintaining inference speeds comparable to llama.cpp and MLC-LLM What does this do with a 40 TOPS+ NPU/TPU?