9 ms·
iPhone 17 Pro Demonstrated Running a 400B LLM
https://xcancel.com/anemll/status/2035901335984611412 https://xcancel.com/anemll/status/2035901335984611412
- philbitt 6mo ago[dead]
- anemll 6mo ago[flagged]
- lostmsu 6mo agoThis has nothing to do with Apple, and everything to do with MoE and that everyone forgot you can re-read the necessary bits of the model from disk for each token. This is extremely inefficient though. For efficiency you need to batch many requests (like 32+, probably more like 128+), and when you do that with MoE you lose the advantage of only having to read a subset of the model during a single forward pass, so the trick does not work. But this did remind me that with dense models you might be able to use disk to achieve high throughput at high latency on GPUs that don't have a lot of VRAM.
- ashwinnair99 6mo agoA year ago this would have been considered impossible. The hardware is moving faster than anyone's software assumptions.
- cogman10 6mo agoThis isn't a hardware feat, this is a software triumph. They didn't make special purpose hardware to run a model. They crafted a large model so that it could run on consumer hardware (a phone).
- pdpi 6mo agoIt's both. We haven't had phones running laptop-grade CPUs/GPUs for that long, and that is a very real hardware feat. Likewise, nobody would've said running a 400b LLM on a low-end laptop was feasible, and that is very much a software triumph.
- bigyabai 6mo ago> We haven't had phones running laptop-grade CPUs/GPUs for that long Agree to disagree, we've had laptop-grade smartphone hardware for longer than we've had LLMs.
- pdpi 6mo agoKind of. We've had solid CPUs for a while, but GPUs have lagged behind (and they're the ones that matter for this particular application). iPhones still lead by a comfortable margin on this front, but have historically been pretty limited on the IO front (only supported USB2 speeds until recently).
- bigyabai 6mo agoThe GPUs are perfectly solid. Cheap Android handsets have shipped with Vulkan compliance for almost a decade now; the GPUs are equally-featured to consoles and PCs. The same goes for Apple handsets that run byte-identical Metal Compute Shaders to the Mac. For desktop use they are perfectly amenable. The hardware lacks nothing required for inference or gaming that dGPUs ordinarily support. And even if you raise the requirements, we still have to contend with cheap CUDA-capable GPUs like the one in the ($300!!!) Nintendo Switch, or the Jetson SOCs. The mobile market has had tons of high-speed/low-power options for a very long time now.
- mnkyprskbd 6mo agoWe had LLMs for about 5 minutes or so. Hardly a measure of time for an industry that goes back half a century and then some.
- 6mo ago
- mannyv 6mo agoThe software has real software engineers working on it instead of researchers. Remember when people were arguing about whether to use mmap? What a ridiculous argument. At some point someone will figure out how to tile the weights and the memory requirements will drop again.
- snovv_crash 6mo agoThe real improvement will be when the software engineers get into the training loop. Then we can have MoE that use cache-friendly expert utilisation and maybe even learned prefetching for what the next experts will be.
- zozbot234 6mo ago> maybe even learned prefetching for what the next experts will be Experts are predicted by layer and the individual layer reads are quite small, so this is not really feasible. There's just not enough information to guide a prefetch.
- snovv_crash 6mo agoManually no. It would have to be learned, and making the expert selection predictable would need to be a training metric to minimize.
- zozbot234 6mo agoMaking the expert selection more predictable also means making it less effective. There's no real free lunch.
- yorwba 6mo agoIt's feasible to put the expert routing logic in a previous layer. People have done it: https://arxiv.org/abs/2507.20984 https://arxiv.org/abs/2507.20984
- Aurornis 6mo agoIt wasn't considered impossible. There are examples of large MoE LLMs running on small hardware all over the internet, like giant models on Raspberry Pi 5. It's just so slow that nobody pursued it seriously. It's fun to see these tricks implemented, but even on this 2025 top spec iPhone Pro the output is 100X slower than output from hosted services.
- zozbot234 6mo agoIf the bottleneck is storage bandwidth that's not "slow". It's only slow if you insist on interactive speeds, but the point of this is that you can run cheap inference in bulk on very low-end hardware.
- Terretta 6mo ago> very low-end hardware iPhone 17 Pro outperforms AMD’s Ryzen 9 9950X per https://www.igorslab.de/en/iphone-17-pro-a19-pro-chip-uebertrifft-desktop-cpus/ https://www.igorslab.de/en/iphone-17-pro-a19-pro-chip-uebert...
- pinkgolem 6mo agoIn single threaded workloads, still impressive
- Aurornis 6mo ago> If the bottleneck is storage bandwidth that's not "slow" It is objectively slow at around 100X slower than what most people consider usable. The quality is also degraded severely to get that speed. > but the point of this is that you can run cheap inference in bulk on very low-end hardware. You always could, if you didn't care about speed or efficiency.
- zozbot234 6mo agoYou're simply pointing out that most people who use AI today expect interactive speeds. You're right that the point here is not raw power efficiency (having to read from storage will impact energy per operation, and datacenter-scale AI hardware beats edge hardware anyway by that metric) but the ability to repurpose cheaper, lesser-scale hardware is also compelling.
- ottah 6mo agoI mean, by any reasonable standard it still is. Almost any computer can run an llm, it's just a matter of how fast, and 0.4k/s (peak before first token) is not really considered running. It's a demo, but practically speaking entirely useless.
- alephnerd 6mo agoDevils advocate - this actually shows how promising TinyML and EdgeML capabilities are. SoCs comparable to the A19 Pro are highly likely to be commodified in the next 3-5 years in the same manner that SoCs comparable to the A13 already are.
- iberator 6mo agoDoes iPhone have some kind of hardware acceleration for neural netwoeks/ai ?
- NetMageSCW 6mo agoYes, a Neural Engine and on the latest A19 tensor processing on the GPU cores (neural accelerator).
- t00 6mo ago/FIFY A year ago this would have been considered impossible. The software is moving faster than anyone's hardware assumptions.
- simopa 6mo agoIt's crazy to see a 400B model running on an iPhone. But moving forward, as the information density and architectural efficiency of smaller models continue to increase, getting high-quality, real-time inference on mobile is going to become trivial.
- volemo 6mo ago> moving forward, as the information density and architectural efficiency of smaller models continue to increase If they continue to increase.
- vessenes 6mo agoThey will. Either new architectures will come out that give us greater efficiency, or we will hit a point where the main thing we can do is shove more training time onto these weights to get more per byte. Similar thing is already happening organically when it comes to efficient token use; see for instance https://github.com/qlabs-eng/slowrun https://github.com/qlabs-eng/slowrun.
- simopa 6mo agoThanks for the link.
- simopa 6mo agoThe "if" is fair. But when scaling hits diminishing returns, the field is forced to look at architectures with better capacity-per-parameter tradeoffs. It's happened before, maybe it'll happen again now.
- anemll 6mo agoProbably 2x speed for Mac Studio this year if they do double NAND ( or quad?)
- firstbabylonian 6mo ago> SSD streaming to GPU Is this solution based on what Apple describes in their 2023 paper 'LLM in a flash' [1]? 1: https://arxiv.org/abs/2312.11514 https://arxiv.org/abs/2312.11514
- simonw 6mo agoYes. I collected some details here: https://simonwillison.net/2026/Mar/18/llm-in-a-flash/ https://simonwillison.net/2026/Mar/18/llm-in-a-flash/
- superjan 6mo agoThat was a very good summary. One detail the post could use is mentioning that 4 or 10 experts invoked where selected from the 512 experts the model has per layer (to give an idea of the savings).
- anemll 6mo agoThanks for posting this, that's how I first found out about Dan's experiment! SSD speed doubled in the M5P/M generation, that makes it usable! I think one paper under the radar is "KV Prediction for Improved Time to First Token" https://arxiv.org/abs/2410.08391 https://arxiv.org/abs/2410.08391 which hopefully can help with prefill for Flash streaming.
- Yukonv 6mo agoThat’s exactly what I thought about. Getting my hands on an M5 Max this week and going to see hows Dan’s experiment performs with faster I/O. Also going to experiment with running active parameters at Q6 or Q8 since output is I/O bottlenecked there should room for higher accuracy compute.
- anemll 6mo agoCheck my repo, I had added some support for GUFF/untloth, Q3,Q5/Q8 https://github.com/Anemll/flash-moe/blob/iOS-App/docs/gguf-hybrid-bringup-log.md https://github.com/Anemll/flash-moe/blob/iOS-App/docs/gguf-h...
- cj00 6mo agoIt’s 400B but it’s mixture of experts so how many are active at any time?
- simonw 6mo agoLooks like it's Qwen3.5-397B-A17B so 17B active. https://github.com/Anemll/flash-moe/tree/iOS-App https://github.com/Anemll/flash-moe/tree/iOS-App
- stingraycharles 6mo agoOne expert is 17B, but more than one expert can be active at any time. I believe it’s actually more like 80B active.
- zozbot234 6mo agoI don't think this is correct, "active parameters" is quite unambiguous in that it means a sum of all active experts plus shared parameters.
- fouc 6mo agolooks like they meant “effective dense size” which is the square root of total params×active params, so in this case sqrt(397 x 17) = ~82
- zozbot234 6mo agoBut the claim that "one expert is 17B" is incorrect. Experts are picked with per-layer granularity (expert 1 for layer X may well be entirely unrelated to expert 1 for layer Y), and the individual layer-experts are tiny. The writeup for the original experiment is very clear on this.
- stingraycharles 6mo agoOk I am by no means an expert on this and I immediately stand corrected. But as I understand it, in order to understand the amount of active memory that’s required, it’s more accurate to go by the ~82B number, right?
- rwaksmunski 6mo agoApple might just win the AI race without even running in it. It's all about the distribution.
- raw_anon_1111 6mo agoApple is already one of the winners of the AI race. It’s making much more profit (ie it ain’t losing money) on AI off of ChatGPT, Claude, Grok (you would be surprised at how many incels pay to make AI generated porn videos) subscriptions through the App Store. It’s only paying Google $1 billion a year for access to Gemini for Siri
- detourdog 6mo agoApple’s entire yearly capex is a fraction of the AI spend of the persumed AI winners.
- devmor 6mo agoWhich is mostly insane amounts of debt leveraged entirely on the moonshot that they will find a way to turn a profit on it within the next couple years. Apple’s bet is intelligent, the “presumed winners” are hedging our economic stability on a miracle, like a shaking gambling addict at a horse race who just withdrew his rent money.
- foobiekr 6mo agoFantasy buildouts of hundreds of billions of dollars for gear that has a 3 year lifetime may be premature. Put another way, there is no demonstrated first mover advantage in LLM-based AI so far and all of the companies involved are money furnaces.
- qingcharles 6mo agoPlus all those pricey 512GB Mac Studios they are selling to YouTubers.
- icedchai 6mo ago
- jee599 6mo ago[dead]
- causal 6mo agoRun an incredible 400B parameters on a handheld device. 0.6 t/s, wait 30 seconds to see what these billions of calculations get us: "That is a profound observation, and you are absolutely right ..."
- WarmWash 6mo agoI don't think we are ever going to win this. The general population loves being glazed way too much.
- baal80spam 6mo ago> The general population loves being glazed way too much. This is 100% correct!
- tombert 6mo agoThat's an astute point, and you're right to point it out.
- actusual 6mo agoYou are thinking about this exactly the right way.
- 9dev 6mo agoYou’re absolutely right!
- 6mo ago
- pier25 6mo agohttps://xcancel.com/anemll/status/2035901335984611412 https://xcancel.com/anemll/status/2035901335984611412
- dang 6mo agoAdded to toptext. Thanks!
- _air 6mo agoThis is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains
- Tade0 6mo agoOnly way to have hardware reach this sort of efficiency is to embed the model in hardware. This exists[0], but the chip in question is physically large and won't fit on a phone. [0] https://www.anuragk.com/blog/posts/Taalas.html https://www.anuragk.com/blog/posts/Taalas.html
- intrasight 6mo agoI think for many reasons this will become the dominant paradigm for end user devices. Moore's law will shrink it to 8mm soon. I think it'll be like a microSD card you plug in. Or we develop a new silicon process that can mimic synaptic weights in biology. Synapses have plasticity.
- bigyabai 6mo agoOne big bottleneck is SRAM cost. Even an 8b model would probably end up being hundreds of dollars to run locally on that kind of hardware. Especially unpalatable if the model quality keeps advancing year-by-year. > Or we develop a new silicon process that can mimic synaptic weights in biology. Synapses have plasticity. It's amazing to me that people consider this to be more realistic than FAANG collaborating on a CUDA-killer. I guess Nvidia really does deserve their valuation.
- intrasight 6mo ago> bottleneck is SRAM cost Not for this approach
- deleted 6mo ago[deleted]
- russellbeattie 6mo agoI have some macro opinions about Apple - not sure if I'm correct, but tell me what you think. Apple has always seen RAM as an economic advantage for their platform: Make the development effort to ensure that the OS and apps work well with minimal memory and save billions every year in hardware costs. In 2026, iPhones still come with 8Gb of RAM, Pro/Max come with 12Gb. The problem is that AI (ML/LLM training and inference) are areas where you can't get around the need for copious amounts of fast working memory. (Thus the critical shortage of RAM at the moment as AI data centers consume as many memory chips as possible.) Unless there's something I don't know (which is more than possible) Apple can't code their way around this problem, nor create specialized SoCs with ML cores that obviate the need for lots and lots of RAM. So, it's going to be interesting whether they accept this reality and we start seeing the iPhones in the future with 16Gb, 32Gb or more as standard in order to make AI performant. And if they give up on adding AI to the billions of iPhones with minimal RAM already out there. As a side note, 8Gb of RAM hasn't been enough for a decade. It prevents basic tasks like keeping web tabs live in the background. My pet peeve is having just a few websites open, and having the page refresh when swapping between them because of aggressive memory management. To me, Apple's obvious strength is pushing AI to the edge as much as possible. While other companies are investing in massive data centers which will have millions of chips that will be outdated within the next couple years, Apple will be able to incrementally improve their ML/AI features by running on the latest and greatest chips every year. Apple has a huge advantage in that they can design their chips with a mega high speed bus, which is just as important as the quantity of RAM. But all that depends on Apple's willingness to accept that RAM isn't an area they can skimp on any more, and I'm not sure they will. Sorry for the brain dump. I'd love to be educated on this in case I'm totally off base.
- ottah 6mo agoPossibly this just isn't the generation of hardware to solve this problem in? We're like, what three or four years in at most, and only barely two in towards AI assisted development being practical. I wouldn't want to be the first mover here, and I don't know if it's a good point in history to try and solve the problem. Everything we're doing right now with AI, we will likely not be doing in five years. If I were running a company like Apple, I'd just sit on the problem until the technology stabilizes and matures.
- dv_dt 6mo agoCPU, memory, storage, time tradeoffs rediscovered by AI model developers. There is something new here, add GPU to the trade space.
- alephnerd 6mo agoIt's been known to people working in the space for a long time. Heck, I was working on similar stuff for the Maxwell and later Pascal over a decade ago. You do have a lot of "MLEs" and "Data Scientists" who only know basic PyTorch and SKLearn, but that kind of fat is being trimmed industry wide now. Domain experience remains gold, especially in a market like today's.
- redwood 6mo agoIt will be funny if we go back to lugging around brick-size batteries with us everywhere!
- pokstad 6mo agoBackpack computers!
- gizajob 6mo agoSeeing as we have the power in our pockets we may as well utilise it. To…type…expert answers… very slowly.
- wayeq 6mo agomight be worth it to keep Sam Altman from reading our AI generated fanfic
- wiether 6mo agoA backpack full of batteries! https://www.youtube.com/watch?v=MI69LUXWiBc https://www.youtube.com/watch?v=MI69LUXWiBc
- r4m18612 6mo agoImpressive. Running a 400B model on-device, even at low throughput, is pretty wild.
- aplomb1026 6mo ago[dead]
- yalogin 6mo agoApple’s unified memory architecture plays a huge part in this. This will trigger a large scale rearchitecture of mobile hardware across the board. I am sure they are already underway. I understand this is for a demo but do we really need a 400B model in the mobile? A 10B model would do fine right? What do we miss with a pared down one?
- refulgentis 6mo agoWhat do we miss? Tl;dr a lot, model is much worse (Source: maintaining llama.cpp / cloud based llm provider app for 2-3 years now)
- Aurornis 6mo ago> Apple’s unified memory architecture plays a huge part in this. This will trigger a large scale rearchitecture of mobile hardware across the board. I am sure they are already underway. Putting the GPU and CPU together and having them both access the same physical memory is standard for phone design. Mobile phones don't have separate GPUs and separate VRAM like some desktops. This isn't a new thing and it's not unique to Apple > I understand this is for a demo but do we really need a 400B model in the mobile? A 10B model would do fine right? What do we miss with a pared down one? There is already a smaller model in this series that fits nicely into the iPhone (with some quantization): Qwen3.5 9B. The smaller the model, the less accurate and capable it is. That's the tradeoff.
- alwillis 6mo ago> Putting the GPU and CPU together and having them both access the same physical memory is standard for phone design. > Mobile phones don't have separate GPUs and separate VRAM like some desktops. That's true. The difference is the iPhone has wider memory buses and uses faster LPDDR5 memory. Apple places the RAM dies directly on the same package as the SoC (PoP — Package on Package), minimizing latency. Some Android phones have started to do this, too. iOS is tuned to this architecture which wouldn't be the case across many different Android hardware configurations.
- 6mo ago
- Yanko_11 6mo ago[dead]
- HardCodedBias 6mo agoThe power draw is going to be crazy (today). Practical LLMs on mobile devices are at least a few years away.
- andix 6mo agoMy iPad Air with M2 can run local LLMs rather well. But it gets ridiculously hot within seconds and starts throttling.
- HPsquared 6mo agoI wonder if anyone has made a liquid cooling system for ipads / phones. Like, a sealed thing that seals onto the back of the device and circulates cooling water directly against the back surface.
- jml7c5 6mo agoA more whimsical method is to put the thing in a glass of water with the cord sticking out. :-) https://www.reddit.com/r/EmulationOnAndroid/comments/1m269k0/playing_skyrim_with_my_watercooled_s25/ https://www.reddit.com/r/EmulationOnAndroid/comments/1m269k0...
- irjustin 6mo agoa sandwich bag would work wonders, then you could use ice to counter the plastic's thermal inefficiencies!
- zenmac 6mo ago> First, in a watertight plastic bag and then in the water? Was wondering, but this the most duct tap hacker solution!
- jychang 6mo agoThrow it in a Thermoplastic Polyurethane (TPU) bag and it'll be a pretty good long term solution.
- Tade0 6mo agoAs I discovered cooling down hardboiled eggs, it's better to keep a thin layer of moisture that cools the object via evaporation. Something of this sort should keep the device moisturised: https://www.thehydrobros.com/products/automatic-water-sprayer-2-liter-electric-mister-graphite?variant=44618187669719 https://www.thehydrobros.com/products/automatic-water-spraye... 0.2ml/s at its lowest setting looks like the ballpark of what's required to maintain temperature.
- jlhawn 6mo ago[dead]
- literoldolphin 6mo ago[dead]
- johnwhitman 6mo ago[flagged]
- zozbot234 6mo agoThe compute needs for MoE models are set by the amount of active parameters, not total.
- MasterScrat 6mo agoThis has a simple pragmatic solution though: https://duckdb.org/2024/12/06/duckdb-tpch-sf100-on-mobile#a-song-of-dry-ice-and-fire https://duckdb.org/2024/12/06/duckdb-tpch-sf100-on-mobile#a-...
- noboostforyou 6mo agoFrom the same article: "The phone a few minutes after finishing the benchmark. It no longer booted because the battery was too cold!"
- mordechai9000 6mo agoRemoving the case and putting it in mineral oil with a circulating pump and a heat exchanger would probably work better
- Sparkle-san 6mo agoJust put it in an oven if it gets too cold.
- alterom 6mo agoIt takes a particularly dry and cool-as-ice sense of humor to label this solution a "simple" and "pragmatic" one.
- jgraham 6mo agoPower in general. Your time-average power budget for things that run on phones is about 0.5W (batteries are about 10Wh and should last at least a day). That's about three orders of magnitude lower than a the GPUs running in datacenters. Even if battery technology improves you can't have a phone running hot, so there are strong physical limits on the total power budget. More or less the same applies to laptops, although there you get maybe an additional order of magnitude.
- 1970-01-01 6mo ago"400 bytes should be enough for anybody"
- Insanity 6mo agoThe 'B' in 400B is billion, not bytes. And the quote '640k ought to be enough for everyone' doesn't have evidence supporting Bill G said it: https://www.computerworld.com/article/1563853/the-640k-quote-won-t-go-away-but-did-gates-really-say-it.html https://www.computerworld.com/article/1563853/the-640k-quote.... That said, it'd be a fun quote and I've jokingly said it as well, as I think of it more as part of 'popular' culture lol
- skiing_crawling 6mo agoI can't understand why this is a surprise to anyone. An iphone is still a computer, of course it can run any model that fits in storage albiet very slowly. The implementation is impressive I guess but I don't see how this is a novel capability. And for 0.6t/s, its not a cost efficient hardware for doing it. The iphone can also render pixar movies if you let it run long enough, mine bitcoin with a pathetic hashrate, and do weather simulations but not in time for the forecast to be relevant.
- illwrks 6mo agoI installed Termux on an old Android phone last week (running LineageOS), and then using Termux installed Ollama and a small model. It ran terribly, but it did run.
- Aachen 6mo agoSomehow this reminds me of the time I downloaded, compiled, and ran a Bitcoin miner with the app called Linux Deploy on my then-new Galaxy Note (the thing called phablet that is now positively small). It ran terribly, but it did run! Having a complete computer in my pocket was very new to me, coming from Nokia where I struggled (as a teenager) to get any software running besides some JS in a browser. I still don't know where they hid whatever you needed to make apps for this device. Android's power, for me, was being able to hack on it (in the HN sense of the word)
- illwrks 6mo agoYes, computer in your pocket indeed! I think the Apple Neo shows just how powerful/capable the mobile chips are getting for computer use.
- mkagenius 6mo agoFwiw, my pixel 8 runs Qwen3.5 4B with 2 tok/s speed. Via pocketpal app. Somehow cactus app didn't work.
- ActorNightly 6mo agoDon't waste time trying to run models locally. Instead, take the advantage of Termux power, namely the fact that you can install things like Openclaw or Gemini-cli. Google Ai plus or Pro plans are actually really good value, considering they bundle it with storage. https://www.mobile-hacker.com/2025/07/09/how-to-install-gemini-cli-on-android-using-termux/ https://www.mobile-hacker.com/2025/07/09/how-to-install-gemi... There is also Termux:GUI with bindings for languages, which you can use to vibecode your own GUI app, which then can basically serve as an interface to an agent, an Termux API which lets you interface with the phone, including USB devices. Furthermore, termux has the cloudflared package availble, which lets you use clouflared free ssh tunnels (as long as you have a domain name). All put together, you can do some pretty cool things.
- CrzyLngPwd 6mo agoI had a dream that everyone had super intelligent AIs in their pockets, and yet all they did was doomscroll and catfish...shortly before everything was destroyed.
- cindyllm 6mo ago[dead]
- SecretDreams 6mo agoA modern Nostradamus?
- CrzyLngPwd 6mo agoIt was just a dream, which quickly turned into a nightmare.
- wiseowise 6mo agoYou know, Quasimodo predicted all of this.
- iLemming 6mo agoThe Anthropic logo is just Kurt Vonnegut’s drawing of an asshole: https://scienceleadership.org/thumbnail/34729/1920x1920 https://scienceleadership.org/thumbnail/34729/1920x1920 Just in case if someone still didn't realize - we do live in Idiocracy https://www.youtube.com/watch?v=gGlJgU9x8tM https://www.youtube.com/watch?v=gGlJgU9x8tM
- DiscourseFan 6mo agoI think the first thing is just a funny little literary allusion for those in the know. I mean isn’t it kind of hilarious that a company valued at $300 billion has a drawing of an asshole for its logo?
- Idesmi 6mo agoIf they really wanted to honour Kurt Vonnegut, Anthropic wouldn't exist.
- lainproliant 6mo agoThis reminds me of how excited people were to get models running locally when llama.c first hit.
- einpoklum 6mo agoI read this title as: "iPhone 17 Pro demonstrated being an overpriced phone".
- nailer 6mo agoActual link: https://x.com/anemll/status/2035901335984611412 https://x.com/anemll/status/2035901335984611412 cc dang
- groby_b 6mo agoFor small values of "running". Don't get me wrong, it's an awesome achievement, but 0.6s token/s at presumably fairly heavy compute (and battery), on a mobile device? There aren't too many use cases for that :)
- gnarlouse 6mo agoIt's like the sloth from Zootopia
- smlacy 6mo agoAnd with only like a dozen tokens of context. What happens when this thing gets the ~100k tokens of context needed to actually make it useful?
- fudged71 6mo agoIf you don't follow anemll, they also have a usable version of OpenClaw running on iPhone. With hardware and model improvements, the future is bright.
- butILoveLife 6mo ago[dead]
- avazhi 6mo agoQwen's MoE models are god awful when they are only running 2B parameters or whatever they downscale to while active. It isn't a 400B model if there's only several orders of magnitude less parameters active when you're actually inferencing...
- butILoveLife 6mo ago[dead]
- deleted 6mo ago[deleted]
- seu 6mo agoSometimes it looks like the purpose of those hundreds of billions of parameters and those apparent feats of engineering, is to get others to tell you how clever you are. Now we have even automated that.
- konaraddi 6mo agoHow? Are there instructions?
- smlacy 6mo agoTotal gimmick. I guess we're "making progress", but this is will never lead to any useful application other than "Yes, you're absulotely right" bots. What's needed for real applications is 10000× the input token context and 10× the output token speed, so we're off by a factor of ... 100,000×?
- system2 6mo agoCorrect, also with the context growing, the conversations cannot continue at the initial speed either. Gimmick or not, this is very sci-fi compared to 10-20 years ago.
- aplomb1026 6mo ago[dead]
- davej32 6mo ago[dead]
- echelon 6mo ago"0.6 t/s" This is a toy. We need to build open infrastructure in the cloud capable of hosting a robust ecosystem of open weights. And then we need to build very large scale open weights. That's the only way we don't get owned by the hyperscalers. At the edge isn't going to happen in a meaningful way to save us.
- aetherspawn 6mo agoIs it though? I would say 'proof of concept' instead. The fact that it's running on a phone now just sets the goalpost and gets everyone excited about it: add more RAM and GPU to the next iPhone and it's not a toy anymore. Co-incidentally, phone companies also have thousands of engineers sitting around wondering what to do in their next release to convince consumers to buy ...
- zozbot234 6mo ago'Toy' and 'proof of concept' are synonymous. What this really opens up is running non-toy models like Qwen3.5 35B-A3B, which are still considered very large in the mobile device context. Yes, it's too slow for interactivity, but if you acknowledge that it's supposed to deliver "Pro" level inference it works quite fine.
- echelon 6mo ago> add more RAM and GPU to the next iPhone and it's not a toy anymore We're not going to get more RAM and GPU in consumer devices. All of the supply is going into data center build outs. As the hyper scaler gamble on the future continues, we get left with weaker (or more expensive) devices - not stronger ones. The market makers make more money if we're left to thin clients. They're also the ones who control supply and the shapes of devices.
- andyferris 6mo agoI highly doubt the A20 Pro will be slower than the A19 Pro - particularly for AI workloads.
- yencabulator 6mo agoQwen3.5-397B-A17B behaves more like a 17B parameter model. Omitting the MoE part from the headline makes it a lie and stupid hype. Quantizing is also a cheat code that makes the numbers lie, next up someone is going to claim running a large model when they're running a 1-bit quantization of it.
- BoorishBears 6mo agoIt behaves more like a ~80B parameter model (geometric mean of active and total params), and has world knowledge closer to a 400B parameter model There's no misleading here, they show every detail from model to quantization to that atrocious time to first token. Stuff like this feels more like code golf than anyone claiming the mainstream phone user is going to even download 100GB of model weights.
- yencabulator 6mo agoI think we're using different meaning of "behaves like". I meant "has tokens/sec performance comparable to".
- BoorishBears 6mo agoI'm using model performance because inference is definitely not comparable to a 17B model when you're streaming model weights on and off disk storage.
- butILoveLife 6mo ago[dead]
- gulugawa 6mo agoThis sounds incredibly dangerous. Local LLMs are going to make people sit on their phones instead of taking to real people.
- bigyabai 6mo agoAnyone can do that right now with a mobile data plan.
- gary_cli 6mo agogood
- lofaszvanitt 6mo agoI miss the old days when words appear one by one, just like images line by line in old modem days.
- system2 6mo agoInnocent times. Also, not too innocent because there was no restriction on anything.
- pshc 6mo agoEven though it's quantized-to-hell Mixture of Experts, honestly, it's crazy this model can run semi-coherently on an phone.
- PinkMilkshake 6mo ago"That is a profound observation, and you are absolutely right..." With all the money you will save on subscription fees you should be able to afford treatment for your psychosis!
- zharknado 6mo ago“Flash” MOE is named for the sloth character in Zootopia I presume?
- momoddo 6mo ago[dead]
- cmiles8 6mo agoTo the extent that the present LLM movement reaches a steady state conclusion it’s highly likely to be open source models on your own hardware that are “good enough” for 95% of use cases. That blows up the whole “industrial complex” being developed around massive data centers, proprietary models, and everything that goes with that. Complete implosion. Apple has sat on the sidelines for much of this as it seems clear they know the end game is everyone just does this stuff locally on their phone or computer and then it’s game over for everything going on now.
- draxil 6mo agoI assume you mean open weight models? I wish we had better open source models. It would make LLMs far less icky if we had nice clean open trained models. A breakthrough on the cost of training would be nice.
- cmiles8 6mo agoFair clarification, yes.
- mike_hearn 6mo agoNemotron is genuinely open source at least at the smaller sizes. You can download the datasets.
- marci 6mo agoAlso everything from scratch by allen.ai. Weights, datasets, code, multiple checkpoints... I like their FlexOlmo concept.
- Yizahi 6mo agoWe really can't have open source LLM, because they are all based on the stolen IP, or stolen IP slightly laundered and under different title.
- harlanji 6mo ago
- heyitssim 6mo ago[dead]
- danielszlaski 6mo ago[dead]
- EruditeCoder108 6mo agoThis is less about “running a 400B model on a phone” and more about clever engineering around constraints. What’s actually happening is: in mixture-of-experts only a small subset of weights is active per token Aggressive quantization Streaming weights from storage instead of loading everything into RAM So the effective working set is much smaller than 400B. That said, the trade-offs are obvious: very low token throughput, high latency, and heavy reliance on storage bandwidth. It’s more of a proof-of-concept than something usable.
- huflungdung 6mo ago[dead]
- adam_patarino 6mo agoI’ve seen this story making the rounds and I’m not just why it’s gotten so much traction. Is it just a good write up?
- bkfh 6mo agoThanks, bot.
- classified 6mo agoWouldn't a bot write better English? Or are they optimized to produce bad grammar already?
- rogerrogerr 6mo agoThis isn't bad grammar, it's bad formatting because it was copy-pasted from somewhere and the newlines didn't take.
- Nahid890 6mo ago[dead]
- pugchat 6mo ago[dead]
- alnah 6mo agoIt's a nice experiment, but I really wonder what's the use case? Privacy, yes. Local, yes. But then? Will people really use an LLM in their iPhone while they can use LLM infrastructure with bigger models for complex tasks? I mean, it really looks cool. But I don't think it's gonna be the future of local AI also. Maybe someone who can build up a very specialized local model for one particular task can enjoy that. Not sure it's gonna be massively use by the common of the mortals... But fore sure, for the industry, there is maybe a direction where we could have different very specialized models, on our devices, that could interoperate together, and then, provide something useful. We'll see. Interesting though! Maybe we still need some years, or decades, before we have devices, laptops, good enough to run good models.
- latexr 6mo ago> Will people really use an LLM in their iPhone while they can use LLM infrastructure with bigger models for complex tasks? If the alternative is paying a subscription and/or being fed ads, people will try the local private ones first.
- Schiendelman 6mo agoThis will become default. Siri (new) and Gemini will eventually run simple tasks locally and only switch to cloud compute when necessary. Apple and Google then won't have to spend as much on their datacenters. I expect OpenAI, Anthropic, and other companies will attempt to do the same, but the OS manufacturers will have a step up.
- rsmtjohn 6mo ago[flagged]
- vedaba 6mo agoI just use mine to doomscroll on Instagram and look at the fluorescent orange color like I’m holding lava
- latexr 6mo agoYou can really feel the sycophantic drivel when it’s coming at 0.6 tokens per second. > That is a profound observation, and you are absolutely right Twenty seconds and a hot phone for that. In the end it took almost four minutes to generate under 150 tokens of nothing. Impressive that they got it to run, but that’s about the only thing.
- ComputeLeap 6mo ago[dead]
- johnwhitman 6mo ago[flagged]
- aimemobe 6mo ago[flagged]
- kampak212 6mo agoI run Qwen 2.5 gguf on an iPhone 16e on production. Handful of them. They’re on the AppStore.