11 ms·
Show HN: OfflineLLM – a Vision Pro app running TinyLlama on device
Hey, I built this in a day while at Founders Inc Apple Vision Pro residency. Try it out, let me know what you think.
- reactordev 3y agoI think you should reevaluate the price target…
- arjvik 3y agoto be higher or lower?
- dghlsakjg 3y agoWhen it comes to the $3k novelty face computer you can likely get away with higher, at least until someone else does it too
- coder543 3y agoNo… not for an app focused on TinyLlama, which I haven’t been able to find a single use case for that an end user would care about. It’s essentially a toy, or optimistically a useful tool for LLM research at very small sizes. Someone is developing an app call cnvrs, which I’ve been using through TestFlight, and it supports TinyLlama and many other models, currently for free. MLCChat is another free app that focuses on Mistral-7B, and that one is in the App Store for sure. Neither is Vision Pro specific… but as someone who actually owns a Vision Pro, I’d rather have an iPad app with useful models than pay for a Vision Pro app with TinyLlama. And I also say this as someone who tried multiple checkpoints of TinyLlama as it was developed, and followed it closely. It was an awesome research project!
- wahnfrieden 3y agoI'm also working on this, but with OpenAI BYOK in addition to local LLM via Llama.cpp: https://ChatOnMac.com https://ChatOnMac.com for iOS/macOS and hopefully visionOS soon.
- whywhywhywhy 3y agoThe entitlement of iOS, iPad and seemingly Vision Pro ecosystems is bizarre. Must be something about how the systems are designed that has devalued applications in the eyes of the users. Mac and Windows ecosystems you'd have no issue even charging up to $7 a month for a AI frontend.
- coder543 3y agoNow it's “entitlement” to say that I don’t want to pay for access to a language model that is completely useless in a chat format like this (a model that I have plenty of experience with), when I already have access to more useful models through other apps? Wow. Your comment sets the bar for entitlement really low. So, surely you spend all of your money buying things that you know are useless out of some obligation to not seem entitled? You can see how ridiculous that sounds, so the most charitable interpretation of your comment is that you didn’t actually read my comment before responding. The feedback I provided was a lot more useful than trying to guilt trip someone into spending money on something they know they won’t get any value out of. If the author switches to a better language model, it will make their app far more attractive to potential buyers, and they can do this, as shown by the existence of other apps that already have. We are fortunate that TinyLlama is not the best model available. > Mac and Windows ecosystems you'd have no issue even charging up to $7 a month for a AI frontend. I absolutely would have an issue with paying $7/mo for a TinyLlama-only frontend, no matter the platform. Maybe you’ve never actually used TinyLlama? What have you found it useful for in a general chat environment? How did the accuracy compare to models like Mistral-7B that run just fine on Vision Pro?
- deleted 3y ago[deleted]
- xp84 3y agoI don’t think it’s that… cool to try to pressure people to give their work away. Nobody is forcing anyone to buy. Pressuring for free or nearly-free gets you garbage like what we have today for mobile games, search, social media, etc.
- dannyw 3y agoGP doesn’t necessarily mean lower, I could see $15 being a fair price for this.
- reactordev 3y agoeveryone assuming I mean cheaper...
- xp84 3y agoClarity is always welcome and appreciated in such comments. Others on the thread did specifically mean cheaper, though.
- codepixel 3y agohow much would you pay?
- fosterfriends 3y agoPrice is fine, I wouldn't stress it at this stage.
- bastawhiz 3y agoWhen I saw your comment I expected it cost $100... Bro it's $7. I pay more each month for email. Good software costs money and $7 is a pittance. What the heck were you expecting, a dollar??
- deleted 3y ago[deleted]
- reactordev 3y ago$10-$15 actually.
- MarioMan 3y agoIt doesn’t look Apple Vision optimized, but this barebones app is free while also running local LLMs on VisionOS, iOS, and iPadOS: https://apps.apple.com/us/app/mlc-chat/id6448482937 https://apps.apple.com/us/app/mlc-chat/id6448482937 I’ve messed with it mostly out of curiosity to see how fast Apple silicon can run inference on mobile.
- btbuildem 3y ago[flagged]
- LorenDB 3y agoI bet most people who paid $3500 for a face computer are happy to pay $6.99 for an AI app. Myself, I'll just keep running Ollama on my Linux laptop and let the Apple fans spend their money.
- deleted 3y ago[deleted]
- brrrrrm 3y ago[flagged]
- jdamon96 3y agonice, congrats on shipping
- codepixel 3y agothanks! move fast and ship is my mantra
- fosterfriends 3y agoLove the spirit! App looks very cool, keep up the great work :)
- LorenDB 3y agoYou might also evaluate Gemma 2B and Qwen 0.5B as alternative tiny models, FWIW.
- behnamoh 3y agoBig downvote to Gemma for reasons that are already discussed.
- deleted 3y ago[deleted]
- supermatt 3y agodiscussed where?
- codepixel 3y agooooh thanks for the tip, will try this
- refulgentis 3y agoStable LM 3B Zephyr, it's the only model below 7B that can handle RAG: i.e. understand "hey those are documents, use them to answer these questions" It'll work too, it was quite delightful to open Test Flight, install my Flutter app not designed for Vision Pro at all, and everything "just worked".
- BrutalCoding 3y agoIs this Flutter app something you created? If so, is it open source? I’m in that same space and I generally just like to learn from other people’s work. If not, all good. I don’t have a Vision Pro myself but I got a similar app which runs on all platforms including iPadOS, thus I guess my app should work on that too. Thanks for the reminder!
- refulgentis 3y ago
- jsheard 3y agoOut of curiosity, how much memory can a single app actually use on the Vision Pro? I know it physically has 16GB of RAM but mobile OSes usually don't let an app use anything close to the entire memory, and that arbitrary limit will dictate how big of a model you can load.
- wahnfrieden 3y agohttps://developer.apple.com/documentation/bundleresources/entitlements/com_apple_developer_kernel_increased-memory-limit https://developer.apple.com/documentation/bundleresources/en... is how you get lots of RAM on iOS/iPadOS but it's not marked for visionOS so I can't tell
- jsheard 3y agoApparently 16GB iPads set the line at 5GB per app by default, and while that flag lets you request more they don't make any guarantees of how much extra quota you will get, so I'm not sure how useful that is for loading a big model if it might randomly fail. I suppose it's probably safe to assume the Vision Pro is similar. https://9to5mac.com/2021/06/25/apps-can-request-access-to-more-ram-with-ios-15-entitlement-exceeding-normal-system-memory-limits https://9to5mac.com/2021/06/25/apps-can-request-access-to-mo...
- samstave 3y ago>request more they don't make any guarantees of how much extra quota you will get How is the logic made for what requests get what? Maybe a new feature in future iOs would be a user setting for specific apps to slider how much the user wants to allot - and a toggle to "shed other apps as needed" and "shed other apps when temp reaches X"
- wahnfrieden 3y agoApple will never do that. And the algo is undocumented.
- 3y ago
- aussieguy1234 3y agoI can't see much more benefit to this compared to SillyTavern on my desktop. What i'd like to see in this space is an actual 3D avatar assistant that you can talk to using your voice as if they were another person.
- codepixel 3y agoyes i'm working on this 3D avatar idea as well. it's actually really mind blowing in my opinion, just need to bring your own imagination. this is just the start, i will add memory, RAG, voice interface, and other features to this.
- samstave 3y agoJust make sure you watch "HER" and "Ex Machina" and that other new one about Huma-Driod relations, for inspiration and caution...
- aussieguy1234 3y agoI've watched both movies. Her was an audio only chatbot and this is already doable. SillyTavern + OpenAI Whisper + Silero TTS and you've basically got Her. I've already done it and it works quite well, Whisper is much much better than the speech recognition Google offers even when running locally on a CPU. Ex Machina was an actual physical robot. Not possible yet, but since GPTs became smart, huge investments are being made in robotics, the most recent annoucement today: https://futurism.com/the-byte/humanoid-robot-maker-deal-openai https://futurism.com/the-byte/humanoid-robot-maker-deal-open... Once this happens a robot will be basically able to do any job a human can do.
- samstave 3y agoCool... My actual point was, Huma-Droid Relations: Her: Human foolishly falls in love with an AI bot (already happened in real life) E.M: AI bot gets a body, lies her way out of prison and releases her self on society (GPt already lied its way through Mechanical Turk Captchas.) Point being, that the 3D avatar will be like the all the AI warning we have of Holographic Personal AI Assistants... and some people will fall in love with them... and some of the assistance will either be/be used for Evil... :-) I didnt doubt you had seen them, though.
- d--b 3y agoIMO, small LLMs are not good enough yet. I understand that people prefer running stuff on-device for privacy and cost reason, but a model that makes mistakes all the time is not worth the tradeoff.
- mmcwilliams 3y agoThis is anecdata but "good enough" is relative. I've finetuned TinyLlama with the same dataset and technique as Llama2 7B for on-device purposes (not for cost or privacy but for physical hardware that have to run offline and with low power consumption) and it produces higher task alignment in 1/4 the inference time. As a general purpose model it isn't great but small models have their place in the ecosystem.
- eurekin 3y agoCare to elaborate on the finetune? It's surprisingly very hard to come across a useful finetuning examples.
- mmcwilliams 3y agoSure, very generally we're doing PEFT starting with insights from examples very much like this one [0] and have gradually built our own tooling and customized the approach a lot as the underlying Huggingface libraries have progressed even in the last 6 months. I will say that one of the most important parts of the process that I've found is in the prompt structuring, the use of special tokens based on how the base models were trained and customizing the tokenizer where necessary. That work in particular is not covered adequately by the examples I was able to find when I started, in my opinion. [0] https://medium.com/@kshitiz.sahay26/fine-tuning-llama-2-for-news-category-prediction-a-step-by-step-comprehensive-guide-to-fine-tuning-48c06dee28a9 https://medium.com/@kshitiz.sahay26/fine-tuning-llama-2-for-...
- codepixel 3y agoyes, i can see that, they are fun to play with though, many of the responses are interesting, and yes they will get more powerful fast, so swapping for another model will be possible and soon i will support this
- whywhywhywhy 3y agoYour screenshots need a bit more thought into them you're essentially just showing some empty rectangles with no context as to what you're selling. Also I'd be looking towards voice as the input/output for llm on Vision Pro.
- codepixel 3y agookay fair, the actual ui is improved then on the screenshots, so i will be updating them asap
- emadm 3y agoShould look at MLX optimisations too, Stable LM 1.6b which is about the same size and quantises to 4bit really well runs 100 tok/s on a M2 Mac mini. https://x.com/awnihannun/status/1750986911827832992?s=20 https://x.com/awnihannun/status/1750986911827832992?s=20
- woadwarrior01 3y agoI'm just about to ship an update to the iOS version my offline LLM app which will replace its current 3B default model (RedPajama Chat) with Stable LM 1.6B. Works extremely well even when quantized. I initially wanted to ship it with TinyLlama Chat, but TinyLlama and its fine tunes are quite subpar and many of my beta testers complained that it's much worse than even the old 3B model and then I found StableLM 2 Zephyr 1.6B. :) https://imgur.com/a/Imd2l9o https://imgur.com/a/Imd2l9o