28 ms·
I made an app that runs Mistral 7B 0.2 LLM locally on iPhone Pros
- novolunt 3y ago[dead]
- sevagh 3y agoAre you allowed to monetize Mistral's weights?
- hananova 3y agoWhy do none of these apps allow you to set the system prompt? I find these LLM apps kind of useless without being able to refine the way in which the model will respond to later questions.
- wahnfrieden 3y agoI made a free / mostly open source one for iOS that lets you edit the system prompt https://chatonmac.com https://chatonmac.com
- ionwake 3y agoAmazing! Does it submit any data online ?
- wahnfrieden 3y agoNo. I definitely do not want any liability of user-generated content or PII or similar. I have no analytics, besides the standard Apple opt-in crash/reporting (not using any 3rd-party service and not sending anything to my own servers). It downloads configuration from GitHub and HuggingFace directly. It also has OpenAI integration, directly to their servers via BYOK.
- ricktdotorg 3y agotrying this out! BTW and FYI i need to reduce the font size on my iOS device to be smaller than i like in order to use your add/replace API key key pages. if the font is "larger than normal" i can't see/focus on the box to enter or paste in the API key. just increase your iOS system font size to trigger this. thanks in advance for fixing, will try out the app!
- wahnfrieden 3y agoThanks for the detailed report - will fix asap, along with releasing the macOS v1.0. I've just soft launched this so far but have more to come so please let me know anything else.
- YetAnotherNick 3y agoMistral instruct doesn't have system prompt AFAIK. Also llama chat system prompt is very useless in my testing.
- rrr_oh_man 3y agoIt seems to work well on GPT4All (macOS) with system prompts? Can you link to any doc why it shouldn't work?
- refulgentis 3y agoIt does; and if it's LLaMa 2 7B Chat stock from Facebook, that was a little rushed imho, doesn't seem as baked in. (GPTs matters but it's _very_ bizarre who it thinks it's coming from)
- nl 3y agoMistral Instruct does use a system prompt. You can see the raw format here: https://www.promptingguide.ai/models/mistral-7b#chat-template-for-mistral-7b-instruct https://www.promptingguide.ai/models/mistral-7b#chat-templat... and you can see how LllamaIndex uses it here (as an example): https://github.com/run-llama/llama_index/blob/1d861a9440cdc9e1b2335320a7a2bf26667b90b2/llama_index/llms/llama_utils.py#L27 https://github.com/run-llama/llama_index/blob/1d861a9440cdc9...
- sp332 3y agoSo the system prompt is just part of the first prompt in a conversation? How is that different from not having a system prompt?
- brittlewis12 3y agowould love for you to give cnvrs a shot! - save characters (system prompt + temperature, and a name & cosmetic color) - download & experiment with models from 1b, 3b, & 7b, and quant options q2k, q4km, q6k - save, search, continue, & export past chats along with smaller touches: - custom theme colors - haptics and more coming soon! https://testflight.apple.com/join/ERFxInZg https://testflight.apple.com/join/ERFxInZg
- sockaddr 3y agoDo not download this. I downloaded this on my 14 Pro and it completely locked up the system to the point where even the power button wouldn’t work. I couldn’t use my phone for about 10 minutes.
- scottbartell 3y agoI've used it for a couple weeks on my 15 Pro and I haven't experienced anything like that. (IMO it's well worth the download) The developer is also pretty responsive and actively looking for feedback (which is why it's currently free on TestFlight)
- brittlewis12 3y agoI’m very sorry about your experience. That’s definitely not what I was aiming for, and I can imagine that was a nasty surprise. Any hang like that is unacceptable, full stop. My understanding is Metal is currently causing hangs on devices when there is barely enough RAM to fit the model and prompt, but not quite enough to run. Will work on falling back to CPU to avoid this kind of experience much more aggressively than today. Thank you for taking the time to both try it out and to share your experience; I will use it to ensure it’s better in the future.
- sockaddr 3y agoThanks for the response. Unfortunately on my device the behavior makes it impossible to report a bug using a screenshot as requested in the app. I can give you more device info if you want to narrow down the cause.
- deleted 3y ago[deleted]
- quickthrower2 3y ago[flagged]
- chupapimunyenyo 3y agoNot everyone chases money
- fragmede 3y agoEveryone has expenses that must get paid in order to live in this world.
- deleted 3y ago[deleted]
- hakdbha 3y ago[dead]
- deleted 3y ago[deleted]
- quickthrower2 3y agoThe app already has a price, so I assume this is for a reason.
- golergka 3y agoMoney is a good tool to chase any other thing.
- hakdbha 3y ago[dead]
- deleted 3y ago[deleted]
- simonw 3y agoDoes it save all conversations and let me revisit them later? I use MLC Chat to run Mistral 7B on my iPhone at the moment, but the lack of conversation history is a real nuisance: https://apps.apple.com/us/app/mlc-chat/id6448482937 https://apps.apple.com/us/app/mlc-chat/id6448482937
- brittlewis12 3y agoyou can absolutely access and continue all your past chats in cnvrs! would love to hear what you think: https://testflight.apple.com/join/ERFxInZg https://testflight.apple.com/join/ERFxInZg
- coder543 3y agoEDIT: Attempting to converse with any Q4_K_M 7B parameter model on a 15 Pro Max... the phone just melts down. It feels like it is producing about one token per minute. MLC-Chat can handle 7B parameter models just fine even on a 14 Pro Max, which has less RAM, so I think there is an issue here. EDIT 2: Even using StableLM, I am experiencing a total crash of the app fairly consistently if I chat in one conversation, then start a new conversation and try to chat in that. On a related note, since chat history is saved... I don't think it's necessary to have a confirmation prompt if the user clicks the "new chat" shortcut in the top right of a chat. ----- That does seem much nicer than MLC Chat. I really like the selection of models and saving of conversations. It looks like you’re still using the old version of TinyLlama. The 1.0 release is out now: https://huggingface.co/TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF https://huggingface.co/TheBloke/TinyLlama-1.1B-Chat-v1.0-GGU... Microsoft recently re-licensed Phi-2 to be MIT instead of non-commercial, so I would love to see that in the list of models. Similarly, there is a Dolphin-Phi fine tune. The topic of discussion here is Mistral-7B v0.2, which is also missing from the model list, unfortunately. There are a few Mistral fine tunes in the list, but obviously not the same thing. I also wish I could enable performance metrics to see how many tokens/sec the model was running at after each message, and to see how much RAM is being used. On the whole, this app seems really nice!
- brittlewis12 3y ago
- xvector 3y agoSome other local LLM iOS apps: - MLC Chat: https://llm.mlc.ai https://llm.mlc.ai - LLM Farm: https://llmfarm.site https://llmfarm.site - Enchanted (not local, just a frontend): https://github.com/AugustDev/enchanted https://github.com/AugustDev/enchanted But I don't think any of these support Mistral 0.2 which is a pretty big deal.
- Firmwarrior 3y agoIs it weird if I carry a phone with this and a solar charger around at all times, in case I suddenly get hurled back in time?
- LewisVerstappen 3y agoI think it would be far better to just store a bunch of epubs on your phone in case you get hurled back. Textbooks on physics, chem, etc.
- genman 3y agoI think a great caution should be used with modern physics and chemistry - it may be a way to get yourself killed for sorcery. But if you want to say alive then I'll recommend including few books about creating modern medicine from scratch - like creating aspirin from willow bark and penicillin from molded bread.
- Firmwarrior 3y agoI don't think it's remotely feasible, but imagine how rich you could get by synthesizing viagra
- jrflowers 3y agoA machine that tells you that the Golden State Warriors won the 2012 Stanley Cup by bowling a perfect 300 would be invaluable in 1602
- elzbardico 3y agoIn that case no need for a LLM, just a wikipedia dump with a full-text index is enough.
- BHSPitMonkey 3y agoI don't think Wikipedia contains facts quite like the one in GP's example...
- TacticalCoder 3y agoAre these LLMs you can run locally giving answers deterministically just as with, say, StableDiffusion? In StableDiffusion if you reuse the exact same version of SD / model and same query and seed, you always get the same result (at least I think so).
- tionis 3y agoYes, you can set the temperature to 0, then they should be deterministic.
- dilawar 3y agoSomeone mentions temperature in the context of algorithms, can't stop thinking, cool, simulated annealing. Haven't seen temperature used in any other family of algo before this.
- potatoman22 3y agoI'm interested, how does LLM temperature relate to simulated annealing?
- amluto 3y agoIf you squint, it’s the same thing. Simulated annealing generally attempts to sample from the Boltzmann distribution. (Presumably because actual annealing is a thermodynamic thing, and you can often think of annealing in a way that the system is a sample from the Boltzmann distribution.) And softmax is exactly the function that maps energies into the corresponding normalized probabilities under the Boltzmann distribution. And transformers are generally treated as modeling the probabilities of strings, and those probabilities are expressed as energies under the Boltzmann distribution (i.e., logits are on a log scale), and asking your favorite model a question works by sampling from the Boltzmann distribution based on the energies (log probabilities) the model predicts, and you can sample that distribution at any temperature you like.
- JimDabell 3y agoEven with Stable Diffusion, determinism is “best effort”- there are flags you can set in Torch to make it more deterministic at a performance cost, but it’s explicitly disclaimed: https://pytorch.org/docs/stable/notes/randomness.html https://pytorch.org/docs/stable/notes/randomness.html
- lulznews 3y agoHow much disk space does it need?
- UnlockedSecrets 3y agoSize 3.3 GB source: The page that is linked
- geuis 3y agoI have a 2020 16in MacBook Pro. I think it's the last generation of Intel chips. I've been struggling to get some of the LLM models like Mixtral to run on it. I hate the idea of needing to buy another $3k laptop less than 4 years after spending that much on my current machine. But if I want to get serious about developing non-chatgpt services, do I need a new M2 or M3 chip to get this stuff running locally?
- muricula 3y agoDoes your mac support an external GPU? A mid to high end nvidia card may or may not outperform the M3 GPU at a lower or similar price. You can also stick it in a PC or resell it separately.
- xfitm3 3y agoDo you recommend any specific external GPU? I had one from Black Magic, it was not that great performance wise.
- kiratp 3y agoNo Nvidia drivers for MacOS.
- lights0123 3y agoCould dual boot Windows or Linux
- fnordpiglet 3y agoeGPU isn’t supported on Apple silicon
- sp332 3y agoAs GP said, the early 2020 MBP had an Intel CPU.
- K0balt 3y agoMy 64gb m2 mbp is faster running inference than my dual 3090 desktop rig, and at 64g of unified memory it can hold slightly bigger models than the 48gb of vram of the desktop. The performance of the m2/m3 with a big unified memory is very impressive. Not much difference between m2/m3 though, if all other things are the same.
- scosman 3y agoHow did you make it? llama.cpp?
- castles 3y agoalmost certainly - edit: https://github.com/ggerganov/llama.cpp/tree/master/examples/llama.swiftui https://github.com/ggerganov/llama.cpp/tree/master/examples/...
- alekseiprokopev 3y agoHere is how to do that on Android: https://github.com/ggerganov/llama.cpp/#android https://github.com/ggerganov/llama.cpp/#android
- jameshart 3y agoI don't think running raw llama.cpp under termux in a shell on your phone, after downloading and compiling it from scratch,, is really comparable to 'I made an app'.
- deleted 3y ago[deleted]
- halJordan 3y ago[flagged]
- enticeing 3y agoI don't read GP comment as saying "Android is always better", unless there's something I missed? More options is a good thing, nobody said a manual process was better than an app.
- _andrei_ 3y agoWhat we're seeing here might be classic case of the iOS Freedom Choking Syndrome: when a device's lack of freedom spreads to its owner and chokes cerebral circulation.
- throwaway743 3y agoPlease keep the snarky cliche responses on other forums where they belong
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- perryizgr8 3y agoAre these apps using the neural compute parts of Apple's chips? Or ar they just using the regular CPU/GPU cores?
- brittlewis12 3y agoTL;DR: No, nearly all these apps will use GPU (via Metal), or CPU, not Neural Engine (ANE). Why? I suggest a few main reasons: 1) No Neural Engine API 2) CoreML has challenges modeling LLMs efficiently right now. 3) Not Enough Benefit (For the Cost... Yet!) This is my best understanding based on my own work and research for a local LLM iOS app. Read on for more in-depth justifications of each point! --- 1) No Neural Engine API - There is no developer API to use the Neural Engine programmatically, so CoreML is the only way to be able to use it. 2) CoreML has challenges modeling LLMs efficiently right now. - Its most-optimized use cases seem tailored for image models, as it works best with fixed input lengths[1][2], which are fairly limiting for general language modeling (are all prompts, sentences and paragraphs, the same number of tokens? do you want to pad all your inputs?). - CoreML features limited support for the leading approaches for compressing LLMs (quantization, whether weights-only or activation-aware). Falcon-7b-instruct (fp32) in CoreML is 27.7GB [3], Llama-2-chat (fp16) is 13.5GB [4] — neither will fit in memory on any currently shipping iPhone. They'd only barely fit on the newest, highest-end iPad Pros. - HuggingFace‘s swift-transformers[5] is a CoreML-focused library under active development to eventually help developers with many of these problems, in addition to an `exporters` cli tool[6] that wraps Apple's `coremltools` for converting PyTorch or other models to CoreML. 3) Not Enough Benefit (For the Cost... Yet!) - ANE & GPU (Metal) have access to the same unified memory. They are both subject to the same restrictions on background execution (you simply can't use them in the background, or your app is killed[7]). - So the main benefit from unlocking the ANE would be multitasking: running an ML task in parallel with non-ML tasks that might also require the GPU: e.g. SwiftUI Metal Shaders, background audio processing (shoutout Overcast!), screen recording/sharing, etc. Absolutely worthwhile to achieve, but for the significant work required and the lack of ecosystem currently around CoreML for LLMs specifically, the benefits become less clear. - Apple's hot new ML library, MLX, only uses Metal for GPU[8], just like Llama.cpp. More nuanced differences arise on closer inspection related to MLX's focus on unified memory optimizations. So perhaps we can squeeze out some performance from unified memory in Llama.cpp, but CoreML will be the only way to unlock ANE, which is lower priority according to lead maintainer Georgi Gerganov as of late this past summer[9], likely for many of the reasons enumerated above. I've learned most of this while working on my own private LLM inference app, cnvrs[10] — would love to hear your feedback or thoughts! Britt --- [1] https://github.com/huggingface/exporters/pull/37 https://github.com/huggingface/exporters/pull/37 [2] https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html https://apple.github.io/coremltools/docs-guides/source/flexi... [3] https://huggingface.co/tiiuae/falcon-7b-instruct/tree/main/coreml/text-generation/falcon-7b-64-float32.mlpackage/Data/com.apple.CoreML/weights https://huggingface.co/tiiuae/falcon-7b-instruct/tree/main/c... [4] https://huggingface.co/coreml-projects/Llama-2-7b-chat-coreml/tree/main/llama-2-7b-chat.mlpackage/Data/com.apple.CoreML/weights https://huggingface.co/coreml-projects/Llama-2-7b-chat-corem... [5] https://github.com/huggingface/swift-transformers https://github.com/huggingface/swift-transformers [6] https://github.com/huggingface/exporters https://github.com/huggingface/exporters [7] https://developer.apple.com/documentation/metal/gpu_devices_and_work_submission/preparing_your_metal_app_to_run_in_the_background https://developer.apple.com/documentation/metal/gpu_devices_... [8] https://github.com/ml-explore/mlx/issues/18 https://github.com/ml-explore/mlx/issues/18 [9] https://github.com/ggerganov/llama.cpp/issues/1714#issuecomment-1693191133 https://github.com/ggerganov/llama.cpp/issues/1714#issuecomm... [10] https://testflight.apple.com/join/ERFxInZg https://testflight.apple.com/join/ERFxInZg
- etaioinshrdlu 3y agoDoes Apple enforce strict safety and content rules on these types of apps?
- Dig1t 3y agoWhat does that mean exactly? Like your phone won’t print text that says something offensive?
- alchemist1e9 3y ago[flagged]
- valianteffort 3y agoWhat the hell are you on about? Apple has rules and guidelines on (user) generated content, GP was asking whether it applied here.
- alchemist1e9 3y agoUnderstood but at some point it becomes the responsibility of the user of the hammer if they use it in an attack or hurt someone else or themselves with it. LLMs are LLMs anyone who is using it who doesn’t understand it is language model and is a machine and how at the high level it works, probably shouldn’t use it, and it shouldn’t be Apple’s responsibility to keep hammers out of the hands of everyone due to the few who can’t hit a nail and stub their thumb.
- deleted 3y ago[deleted]
- freedomben 3y agoApple forcefully takes on that responsibility, and their customers love it (as clearly evidenced by their domination of the market). If you don't want a hammer that gets reviewed and screened by Apple before you can use it, then you're on the wrong platform.
- johngalt2600 3y agoWhere to leave feedback? I am trying the Mistral dolphin model but getting GGML ASSERT errors referencing Users/tito lol (not me). Using iPhone 14 Pro Max.
- winstonschen 3y agoWhich app are you trying? The app posted here is Offline Chat, which doesn’t have a choice of model and works fine on iPhone 14 Pro.
- brittlewis12 3y agoSorry for the confusing experience, and thank you for sharing this! I’ve just submitted a new update for review with a few small but hopefully noticeable changes, thanks in no small part to your feedback: 1. StableLM Zephyr 3b Q4_K_M is now the built-in model, replacing the Q6_K variant. 2. More aggressive RAM headroom calculation, with forced fallback to CPU rather than failing to load as you observed, or crashing outright in some nasty edge cases. 3. New status indicator for Metal when model is loaded (filled bolt for enabled, vs slashed bolt for disabled.) 4. Metal will now also be enabled for devices with 4GB RAM or less, but only when the selected model can comfortably fit in RAM. Previously, only devices with at least 6GB had Metal enabled. Thank you so much for taking the time to test and share your experience! Feel free to reach out anytime at britt [at] bl3 [dot] dev. Britt
- furyofantares 3y agoedit: my bad, I misread the price and it's really hard to see the price after you bought it to double check. $10 for something that (I think) doesn't work on most phones but isn't gated to ones it works on feels hostile. Probably there's no way to gate, in that case I'd suggest not charging for it. Or I guess adding a daily usage limit that's lifted with an IAP. I'll admit I was off-put by the price to begin with, which probably amplifies what a slap in the face it feels like to pay and get something that doesn't work at all.
- sp332 3y agoIt's $1.99 and the description says: The app requires a Pro iPhone with a minimum of 6GB of RAM. Only the following devices meet the requirement: - iPhone 15 Pro, iPhone 14 Pro, iPhone 13 Pro, iPhone 12 Pro. - iPads: Please check. RAM varies based on model and year.
- halJordan 3y agoSo if they know it wont work, and do not put that info into the store's compatibility matrix then it's still a bait/switch to me. Compare to the Resident Evil page which does set the store limits on what devices can dl it.
- deleted 3y ago[deleted]
- winstonschen 3y agoThere is no way to specify iPhone models or memory capacity when submitting an app to App Store. Believe me - I spent several days trying.
- dasickis 3y agoYou can set the minimum deployment to iOS 17 & then if someone has iPhone X*, 11 or SE then you can alert them to get a refund when they open the app either with a device check or total memory check. That'll set it so you remove most of the issues of older devices. Source: https://support.apple.com/guide/iphone/models-compatible-with-ios-17-iphe3fa5df43/ios https://support.apple.com/guide/iphone/models-compatible-wit...
- dazzaji 3y agoI’m intrigued and currently downloading this app. Love the idea of having offline direct access to this model. One small-ish thing though: Looks like the URL for the privacy policy (http://opusnoma.com/privacy http://opusnoma.com/privacy) linked from the App Store page goes nowhere. Actually, opusnoma.com is likewise offline.
- tanepiper 3y agoNow stick a "Don't Panic" sticker it your phone...
- Alifatisk 3y agoTo think that we went from ClosedAi (OpenAi)s chatGPT to now being able to do this on out phones offline is incredible.
- andersa 3y agoWhile it's a great achievement for sure, quantized Mistral 7b is not even remotely comparable to ChatGPT.
- deleted 3y ago[deleted]
- Const-me 3y agoI made a free and open source Windows equivalent: https://github.com/Const-me/Cgml/releases/tag/1.1 https://github.com/Const-me/Cgml/releases/tag/1.1
- Kelteseth 3y agoCan you link a ready to use example model I can just download and toy around with?
- Const-me 3y agoThe model is on BitTorrent, see readme for the frontend app: https://github.com/Const-me/Cgml/tree/master/Mistral/MistralChat https://github.com/Const-me/Cgml/tree/master/Mistral/Mistral... The torrent file is also inside the MistralChat.zip archive.
- Const-me 3y agoMinor update https://github.com/Const-me/Cgml/releases/tag/1.1a https://github.com/Const-me/Cgml/releases/tag/1.1a Can’t edit that comment anymore, too late.
- antirez 3y agoI see the efforts required to create the little app, but inference via llama.cpp or core Ml is trivial and the models are open weights, so it makes more sense to have a free app for this: most of the value is in the LLM which is free.
- scanny 3y agoI think there is some cost associated with iPhone app development ($100-$300 plus submission costs), as opposed to android, when it comes to publishing, it seems fair enough for an individual to charge a dollar or two to recoup that.
- kccqzy 3y agoI'd argue in this space besides the model weights, a lot of the value comes from a nice, not-too-fancy but nevertheless intuitive and delightful UI. I mean I've used the free MLC Chat app which runs Mistral 7B fine, and because it's free, I have very low expectations of its UI design. If someone is making a new app with a nicer UI, I really don't mind paying a buck or two.
- antirez 3y agoThat makes sense, but if that's the case I would like to see something a bit more vague polish. For instance: 1. Ability to forward messages to app. 2. Ability to run the same questions to the different LLMs installed. 3. Some updated list of GGUF files I can download with a description of the model highlight. 4. Advanced things like check token preplexity to identify parts of the chat the LLM is most unsure and highlight them? I could continue because it's full of obvious things like that. Do this effort, and I'll pay you 10$ for the app, not 2$. But if it's just not brutal but very low effort, it seems that it's not going anyway since the free apps with exact capabilities will emerge and will not be so different.
- madlag 3y agoI love the idea, that's the future. However you should be aware that the explanation of second law of thermodynamics generated by the LLM you used in your app store screenshot is wrong: the LLM has it backwards. Energy transfers to less stable states from more stable states, and not the reverse. (I use LLMs for science education apps like https://apps.apple.com/fr/app/explayn-learn-chemistry/id6448284993 https://apps.apple.com/fr/app/explayn-learn-chemistry/id6448..., so I am quite used to spot that kind of errors in LLM outputs...)
- Horffupolde 3y agoHow do you define stability in that context?
- madlag 3y agoStability is actually defined by having a lower energy level. That explains why energy can only flow from a less stable system to a more stable system : the more stable system does not have available energy to give.
- wannabag 3y agoOh, that's an interesting app and in French too... is that something you plan to have on Android as well?
- madlag 3y agoYes, it's Unity based, so quite easy. There is another version on Quest too, so running on Android : https://www.meta.com/fr-fr/experiences/6113695908674751/ https://www.meta.com/fr-fr/experiences/6113695908674751/ .
- Const-me 3y agoIs that explanation better? https://github.com/Const-me/Cgml/blob/master/Mistral/MistralChat/screenshot.png https://github.com/Const-me/Cgml/blob/master/Mistral/Mistral... Same Mistral Instruct 0.2 model, different implementation.
- 3y ago
- foxhop 3y agoare you using a quantized version of the model and if so which one?
- hospitalJail 3y agoHow long does it take to answer "2+2=" I have a 3060 on a laptop and it is faster than gpt4. I use it all the time. How can you possibly use this? I tried doing CPU based on a bleeding edge laptop and couldn't use it.
- nojvek 3y agoThis is the kind of thing I’d expect Mistral to ship. They shouldn’t be just chasing the api revenue stream like OpenAI
- woadwarrior01 3y agoI've had a successful offline LLM app[1] on the App Store since June, last year. Works on all iPhones since iPhone 11 and ships with a 3B RedPajama Chat model and has an optional download for 7B Llama 2 based model on newer iPhones and Apple Silicon iPads. I'm currently working on an update to bring more 3B and 7B models to the iOS app. [1]: https://apps.apple.com/us/app/private-llm/id6448106860 https://apps.apple.com/us/app/private-llm/id6448106860
- trao 3y agoI just bough the app and learnt that this app is just a reskin of the MLC-LLM iOS app. Save yourself the $1.99 and get that app for free instead. https://apps.apple.com/us/app/mlc-chat/id6448482937 https://apps.apple.com/us/app/mlc-chat/id6448482937
- bloody-crow 3y agoThis might be the best reason to consider a Pro model next time I'm upgrading my iPhone.
- woadwarrior01 3y agoiPhone 15 and iPhone 14 Pro, 14 Pro Max have exactly the same CPU and amount of RAM (Apple A16 Bionic and 6GB). This is also true for iPhone 14 and iPhone 13 Pro, Pro Max (Apple A15 Bionic and also 6GB).
- bloody-crow 3y agoI don't play games or do anything too resource-demanding on my phone normally. Pro models typically have more memory than non-pro models and running LLMs on device might be the only scenario where it can realistically make a difference for me.
- woadwarrior01 3y agoSmaller 3B LLMs (like phi-2) work fine on newer non pro models, at full context lengths. Running 7B models on even 8GB iPhone 15 Pro and Pro Max phones involves reducing the context lengths to 1k or fewer tokens, because the full context length KV cache won't fit on these devices.
- ggrelet 3y agoWhat happens if it's launched on a non-Pro earlier iPhone?
- urbandw311er 3y agoAre there any models out there that don’t come trained or tweaked or system prompted into somebody else’s idea of ethical or professional conduct? I tested out a bunch of these apps and asked them to write an explicit story to see if they would, and despite this being entirely legal, none would do so. Are we entering some new Orwellian era?