5 ms·
The bottleneck with compute at the edge is (and will be) model size (both app download time and storage space on device). Stable Diffusion sits at about 2GB fo
by ronyfadel 4y ago
The bottleneck with compute at the edge is (and will be) model size (both app download time and storage space on device).
Stable Diffusion sits at about 2GB for fp16, Whisper Medium at 1.53GB, LLAMA is 120GB.
Sure, Apple can ship an optimized model (<2-4GB) as part of the OS, but what if a capable app maker wants to ship one? Users will not be happy with an app sized at >1GB.
- bloudermilk 4y agoI would gladly install a 1GB+ app on my 512GB iPhone that never gets above 50GB because of iCloud’s remote storage optimizations.
- ronyfadel 4y agoThat is not even taking into consideration that shipping your weights to the client is akin to giving your product away.
- aix1 4y agoWait, how is that fundamentally different from shipping binary code? Isn't that also akin to "giving your product away"?
- ronyfadel 4y agoYou can easily pirate an ML model, host it and provide a backend service that uses it, and no one would suspect anything. I don't think you can pirate e.g. the Facebook app binary and make it your own.
- KuzMenachem 4y agoYou can put some obstacles (https://developer.apple.com/documentation/coreml/generating_a_model_encryption_key https://developer.apple.com/documentation/coreml/generating_...) in the way though.
- CGamesPlay 4y agoResearch developments are already showing that our models are woefully inefficient in their current state (compare the performance of GPT-3 140B against Alpaca 30B). Not only will hardware get better, the minimum model sizes for good inference will become smaller in the future.
- upbeat_general 4y agoTons of popular iOS games are >1GB.
- mrtksn 4y agoHyper casual games are around 300Mb these days, proper AAA games are multiple GB. People still download those, as you can tell by the billions of dollars they make. The problem with OpenAI's business model is that it's actually quite expensive for them to maintain centralised processing. With Apple, there are billions of very powerful computers deployed to users and these computers mostly stay idle apart from occasionally running some bloated JS to show a button ar something. If Apple manages to run a good enough model on device with acceptable performance and energy impact, then suddenly OpenAI and Microsoft will be just burning away money with no expectation of recouping if they provide the service for free, if they make it paid they will be making money in a niche.
- jonplackett 4y agoIt might actually be _good_ for apple to have a genuine reason to get a new phone. There really isn’t that much difference between iPhone 12 and 14 If a nee one comes out with LLM Siri + hardware that makes it possible that would be a massive upgrade cycle.
- singularity2001 4y agoYes, by the time iphones can smoothly run llama 30G, the state of the art gpt-x will probably be a terabyte. Skynet will forever live in the cloud with just assistant agents living on the devices.
- dragonwriter 4y ago> Yes, by the time iphones can smoothly run llama 30G, the state of the art gpt-x will probably be a terabyte. Yeah, but if every individual is running decent chat-capable LLM, and businesses are running their own on their own devices, and those can communicated with each other, who needs to rely on Skynet?
- KennyBlanken 4y agoA number of games are in the 60-150GB range...
- mikkelam 4y agoApple is artifically shipping phones with low storage sizes to upsell phones with more space. If there was a big economical advantage for them to have larger models on the phone I expect it would be easy to solve.
- ayewo 4y agoApple typically solves this with device segmentation. They can announce an iPhone Pro Ultra model that comes with higher RAM and storage capacity along with a souped up Neural Engine, similar to how they do today with how there are differences between the iPhone and the iPhone Pro screen and camera. Even better, they could bundle the base model for LLAMA with iOS and ship incremental model updates to those iPhone Ultra users (possibly on a monthly subscription).
- jonplackett 4y agoI wonder if they’d create a simple version that lives on device that can call on a ‘cleverer’ version for more extensive tasks - like how ChatGPT is using plugins. Most interactions probably don’t require full power
- LeanderK 4y agoyeah, some sort of caching. Small models for most tasks with a way bigger models just one get-request away. It's a smart idea that hasn't been tried yet. But I bet that you could shrink the models significantly. It doesn't really have to know so much if it can google, but I we don't really know how to train such models (we just want the "common sense" without so much knowledge).