3 ms·
Depends on your definition of edge or smart device. My phone can run a 7B parameter model at 12 tokens per second, which is probably faster than most humans ar
by coder543 3y ago
Depends on your definition of edge or smart device.
My phone can run a 7B parameter model at 12 tokens per second, which is probably faster than most humans are comfortable reading, and definitely faster than a virtual assistant would speak.
Out of curiosity, I tested a 3B parameter model, and it runs at about 21 tokens per second on my phone.
- omneity 3y agoGenerating text fast enough for a human to read it imo is only the bare minimum. New classes of use cases would be possible if you could generate 100s of tokens. For example, on-device classification of emails/texts, LLM-powered recommendation system based on your local data (to go beyond simply parsing dates in text for example), context-aware text or email auto-responder (I'm sorry, can't reply as I'm driving/in a meeting, I'm not home next week, can you deliver to this address instead etc.) ... Many of these use cases are possible today with either specialized models, or are old school and rule-based. Being able to have an LLM apply soft judgment on a device that generates so much contextual information, and completely locally/privately, is bound to make smartphones an entirely new kind of device.
- coder543 3y agoMany of the use cases you’re describing can be done offline, such as when the phone is charging overnight, although not all of them. An email autoresponder could still work in real time at these token rates, and it would still be faster than most humans at responding to an email. 7 hours * 3600sec/hr * 21token/sec = 530,000 tokens per night on this hardware, assuming no thermal throttling. (I don’t have data to say what the sustained rate would be, throttling could happen.)
- omneity 3y agoAgreed on overnight batch processing. Although my vision for it is to have a sort of local service that can provide "intelligence" on demand for other apps, which might request it concurrently, at which point double digit throughput might become limiting. There are other reasons to want a higher throughput. To perform retrieval or for a chain-of-thought approach, you typically need to run several prompts per user prompt, effectively impacting user-perceived performance of your LLM based solution.
- quaintdev 3y agoHow to run it on phone??
- coder543 3y agoI use MLC Chat to run Llama 7b. Not incredibly useful, but it is fun to experiment with. https://apps.apple.com/us/app/mlc-chat/id6448482937 https://apps.apple.com/us/app/mlc-chat/id6448482937