4 ms·
Anyone packaged one of these in an iPhone App? I am sure it is doable, but I am curious what tokens/sec is possible these days. I would love to ship "private" A
by justanotheratom 1y ago
Anyone packaged one of these in an iPhone App? I am sure it is doable, but I am curious what tokens/sec is possible these days. I would love to ship "private" AI Apps if we can get reasonable tokens/sec.
- Alifatisk 1y agoIf you ever ship a private AI app, don't forget to implement the export functionality, please!
- deleted 1y ago[deleted]
- idonotknowwhy 1y agoYou mean conversations? Just the jsonl of the standard hf dataset format to import into other systems?
- Alifatisk 1y agoYeah I mean conversations.
- nico 1y agoWhat kind of functionality do you need from the model? For basic conversation and RAG, you can use tinyllama or qwen-2.5-0.5b, both of which run on a raspberry pi at around 5-20 tokens per second
- justanotheratom 1y agoI am looking for structured output at about 100-200 tokens/second on iPhone 14+. Any pointers?
- nico 1y agoThe qwq-2.5-0.5b is the tiniest useful model I've used, and pretty easy to fine-tune locally on a Mac. Haven't tried it on an iPhone, but given it runs at about 150-200 tokens/second on a Mac, I'm kinda doubtful it could do the same on an iPhone. But I guess you'd just have to try
- zamadatix 1y agoThere are many such apps, e.g. Mollama, Enclave AI or PrivateLLM or dozens of others, but you could tell me it runs at 1,000,000 tokens/second on an iPhone and I wouldn't care because the largest model version you're going to be able to load is Gemma 3 4B q4 (12 B won't fit in 8 GB with the OS + you still need context) and it's just not worth the time to use. That said, if you really care, it generates faster than reading speed (on an A18 based model at least).
- woodson 1y agoSome of these small models still have their uses, e.g. for summarization. Don’t expect them to fully replace ChatGPT.
- zamadatix 1y agoThe use case is more "I'm willing to have really bad answers that have extremely high rates of making things up" than based on the application. The same goes for summarization, it's not like it does it well like a large model would.
- nolist_policy 1y agoFWIW, I can run Gemma-3-12b-it-qat on my Galaxy Fold 4 with 12Gb ram at around 1.5 tokens / s. I use plain llama.cpp with Termux.
- Casteil 1y agoDoes this turn your phone into a personal space heater too?