9 ms·
"Scaling up performance from M5 and offering the same breakthrough GPU architecture with a Neural Accelerator in each core, M5 Pro and M5 Max deliver up to 4x f
by Tangokat 7mo ago
"Scaling up performance from M5 and offering the same breakthrough GPU architecture with a Neural Accelerator in each core, M5 Pro and M5 Max deliver up to 4x faster LLM prompt processing than M4 Pro and M4 Max, and up to 8x AI image generation than M1 Pro and M1 Max."
Are they doubling down on local LLMs then?
I still think Apple has a huge opportunity in privacy first LLMs but so far I'm not seeing much execution. Wondering if that will change with the overhaul of Siri this spring.
- jahller 7mo agolooks like this will be their angle for the whole agentic AI topic
- Sharlin 7mo ago"Apple Intelligence is even more capable while protecting users’ privacy at every step." Remains to be seen how capable it actually is. But they're certainly trying to sell the privacy aspect.
- re-thc 7mo ago> Remains to be seen how capable it actually is. It's the best. We all turned it off. 100% privacy.
- whizzter 7mo agoWe had a workshop 6 months ago and while I've always been sceptical of OpenAI,etc's silly AGI/ASI claims, the investments have shown the way to a lot of new technology and has opened up a genie that won't be put back into the bottle. Now extrapolating in line with how Sun servers around year 2000 cost a fortune and can be emulated by a 5$ VPS today, Apple is seeing that they can maybe grab the local LLM workloads if they act now with their integrated chip development. But to grab that, they need developers to rely less on CUDA via Python or have other proper hardware support for those environments, and that won't happen without the hardware being there first and the machines being able to be built with enough memory (refreshing to see Apple support 128gb even if it'll probably bleed you dry).
- fny 7mo agoI feel like the push by devs towards Metal compatibility has been 10x than AMD. I assume that's because the majority of us run MacBooks.
- davidmurdoch 7mo agoWho is "us" in this case? Majority of devs that took the stack overflow survey use Windows: https://survey.stackoverflow.co/2025/technology/#1-computer-operating-systems https://survey.stackoverflow.co/2025/technology/#1-computer-...
- deleted 7mo ago[deleted]
- pdpi 7mo agoI think it's reasonable to say that the people responding to surveys on Stack Overflow aren't the same people who work on pushing the state of the art in local LLM deployment. (which doesn't prove that that crowd is Apple-centric, of course)
- davidmurdoch 7mo agoPerhaps. Though Windows has been the majority share even when stack overflow was at it's peak, and before.
- petercooper 7mo agoIt's not the whole answer, but SO came from the .NET world and focused on it first so it had a disproportionately MS heavy audience for some time. GitHub had the same issue the other way around. Ruby was one of GitHub's top five languages for its first decade for similar reasons.
- AdamN 7mo agoThat's the broad developer community. 90%+ of the engineers at Big Tech and the technorati startups are on MacOS with 5% on Linux and the other 5% on Windows.
- aurareturn 7mo agoAre they doubling down on local LLMs then? Neural Accelerator was present in iPhone 17 and M5 chip already. This is not new for M5 Pro/Max. Apple's stated AI strategy is local where it can and cloud where it needs. So "doubling down"? Probably not. But it fits in their strategy.
- butILoveLife 7mo agoI think its just marketing, and the marketing is working. Look how many people bought Minis and ended up just paying for API calls anyway. (Saw it IRL 2x, see it on reddit openclaw daily) I don't mind it, I open Apple stock. But I'm def not buying into their rebranding of integrated GPU under the guise of Unified Memory.
- Hamuko 7mo agoI've tried to use a local LLM on an M4 Pro machine and it's quite painful. Not surprised that people into LLMs would pay for tokens instead of trying to force their poor MacBooks to do it.
- giancarlostoro 7mo agoWhat are the other specs and how's your setup look? You need a minimum of 24GB of RAM for it to run 16GB or less models.
- SV_BubbleTime 7mo agoThis is typically true. And while it is stupid slow, you can run models of hard drive or swap space. You wouldn’t do it normally, but it can be done to check an answer in one model versus another.
- Hamuko 7mo ago48 GB MacBook Pro. All of the models I've tried have been slow and also offered terrible results.
- giancarlostoro 7mo agoTry a software called TG Pro lets you override fan settings, Apple likes to let your Mac burn in an inferno before the fans kick in. It gives me more consistent throughput. I have less RAM than you and I can run some smaller models just fine, with reasonable performance. GPT20b was one.
- 7mo ago
- lynx97 7mo agoThe topic is MacBook, so my criticism is a little off. However, I really dont believe in this "local LLM" promise from Apple. My phone already gets noticeably warm if I answer 5 WhatsApp messages. And looses 5% of battery during the process. I highly doubt Apple will have a useable local LLM that doesn't drain my battery in minutes, before 2030.
- cosmic_cheese 7mo agoSomething is not right if WhatsApp is seriously draining your phone like that. Admittedly I’m not a big WhatsApp user my iPhone hasn’t had any trouble like that with it.
- jakeydus 7mo agoYeah is OP using an iPhone X?
- blackqueeriroh 7mo agoBro that’s WhatsApp. Meta is known for their dirty mobile code
- Aurornis 7mo agoThe hardware capabilities that make local LLMs fast are useful for a lot of different AI workloads. Local LLMs are a hot topic right now so that’s what the marketing team is using as an example to make it relatable.
- kilroy123 7mo agoI've been so disappointed in Apple's lack of execution on this. There is so much potential for fantastic local models to run and intelligently connect to cloud models. I just don't get why they're dropping the ball so much on this.
- NetMageSCW 7mo agoBecause it won’t sell enough hardware to matter to them. They aren’t dropping the ball, they are being smart and prudent.
- kilroy123 7mo agoDownvote all you want. Point blank, they are dropping the ball.
- game_the0ry 7mo ago> Are they doubling down on local LLMs then? Honestly, I think that's the move for apple. They do not seem to have any interest in creating a frontier lab/model -- why would they give the capex and how far behind they are. But open source models (Kimi, Deepseek, Qwen) are getting better and better, and apple makes excellent hardware for local LLMs. How appealing would it be to have your own LLM that knows all your secrets and doesnt serve you ads/slop, versus OpenAI and SCam Altman having all your secrets? I would seriously consider it even if the performance was not quite there. And no need for subscription + cli tool. I think apple is in the best position to have native AI, versus the competition which end up being edge nodes for the big 4 frontier labs.
- iAMkenough 7mo agoRE Frontier models/hardware: I'm interested to see what happens with their "private cloud compute" marketing concept now that they're moving from running Siri AI experiences on Apple servers to Google servers instead.
- Spooky23 7mo agoYou can deliver confidential compute on GCP.
- andy_ppp 7mo agoIt is simply marketing nonsense - what they really mean (I think) is they support matrix multiplication (matmul) at the hardware level which given AI is mostly matrix multiplications you'll get much faster inference (and some increase in training too) on this new hardware. I'm looking forward to seeing how fast a local 96gb+ LLM is on the M5 Max with 128gb of RAM.
- manmal 7mo agoWe've already established in this thread that memory bandwidth isn't that much greater than M4 Max - 12%? However, I wonder if batched inference will benefit greatly from the vastly improved compute. My guess is that parallel usage of the same model will be a couple times faster. So, single "threaded" use not that much better, but say you want to run a lot of batch jobs, it'd be way faster?
- andy_ppp 7mo agoIs this a reply to a different comment?
- ivankra 7mo agoBut memory bandwidth (bottleneck for LLM inference) is only marginally improved, 614 GB/s vs 546 GB/s for M4/M5 Max - where is this 4x improvement coming from? I think I'll pass on upgrading.
- general_reveal 7mo agoIt’s not necessarily doubling down on local. The reality is your LLM should be inferencing every tick … the same way your brain thinks every. Fucking. Nano. Second. So yes, the LLM should be inferencing on your prompt, but it should also be inferencing on 25,000 other things … in parallel. Those are the compute needs. We just need compute everywhere as fast as possible.
- Lalabadie 7mo agoThere already are a bunch of task-specific models running on their devices, it makes sense to maintain and build capacity in that area. I assume they have a moderate bet on on-device SLMs in addition to other ML models, but not much planned for LLMs, which at that scale, might be good as generalists but very poor at guaranteeing success for each specific minute tasks you want done. In short: 8gb to store tens of very small and fast purpose-specific models is much better than a single 8gb LLM trying to do everything.
- Munachi1869 7mo agoProbably possible for pure coding models. I see on-device models becoming viable and usable in like 2-3 years on device
- deleted 7mo ago[deleted]
- Someone1234 7mo agoApple's AI strategy really kind of threads the needle cleverly. "AI" (LLMs) may or may not have a bubble-pop moment, but until it does Apple get to ride it on these press releases and claims. But if the big-pop occurs, then Apple winds up with really fantastic hardware that just happens to be good at AI workloads (as well as general computing). For example, image classification (e.g. face recognition/photo tagging), ASR+vocoders, image enhancement, OCR, et al, were popular before the current boom, and will likely remain popular after. Even if LLM usage dries up/falls out of vogue, this hardware still offers a significant user benefit.
- ChrisGreenHeur 7mo agothose things could likely just run fine on the gpu though
- Someone1234 7mo agoThey could run fine on the CPU too. But these are mobile devices, therefore battery usage is another significant metric. Dedicated hardware is more energy efficient than general hardware, and GPU in particular is a power-hog.
- vel0city 7mo agoExactly. It's the same thing as video or audio encoding and decoding. Sure the CPU could do it, potentially use the GPU, but having actual hardware encoders and decoders for the most common codecs will save a lot of energy.
- Nevermark 7mo agoNot if GPU RAM is a limiter. Which it is for most models. Unified memory is a serious architectural improvement. How many GPUs does it take to match the RAM, and make up for the additional communication overhead, of a RAM-maxed Mac? Whatever the answer, it won’t fit in a MacBook Pro’s physical and energy envelopes. Or that of an all-in-one like the Studio.
- lamontcg 7mo ago
- jmyeet 7mo agoApple absolutely has a massive opportunity here because they used a shared memory architecture. So as most people in or adjacent to the AI space know, NVidia gatekeeps their best GPUs with the most memory by making them eye-wateringly expensive. It's a form of market segmentation. So consumer GPUs top out at 16GB (5090 currently) while the best AI GPUs (H200?) is 141GB (I just had to search)? I think the previou sgen was 80GB. But these GPUs are north of $30k. Now the Mac Studio tops out currently at 512GB os SHARED memory. That means you can potentially run a much larger model locally without distributing it across machines. Currently that retails at $9500 but that's relatively cheap, in comparison. But, as it stands now, the best Apple chips have significantly lower memory bandwidth than NVidia GPUs and that really impacts tokens/second. So I've been waiting to see if Apple will realize this and address it in the next generation of Mac Studios (and, to a lesser extend, Macbook Pros). The H200 seems to be 4.8TB/s. IIRC the 5090 is ~1.8TB/s. The best Apple is (IIRC) 819GB/s on the M3 Ultra. Apple could really make a dent in NVidia's monopoly here if they address some of these technical limitations. So I just checked the memory bandwidth of these new chips and it seems like the M5 is 153GB/s, M5 Pro is ~300 and M5 Max is ~600. I was hoping for higher. This isn't a big jump from the M4 generation. I suspect the new Studios will probably barely break 1TB/s. I had been hoping for higher.
- SirMaster 7mo ago>So consumer GPUs top out at 16GB (5090 currently) 5090 has 32GB, and the 4090 and 3090 both have 24GB.
- deleted 7mo ago[deleted]
- ericd 7mo agoHard to get 6000+ bit memory bus HBM bandwidth out of a 512 or 1024 bit memory bus tied to DDR... I think it's also just tough to physically tie in 512 gigs close enough to the GPU to run at those speeds. But yeah, I wish there was a very competitive local option, too, short of spending $50k+.
- fridder 7mo ago
- meisel 7mo agoWhat % of users actually care that much about local LLMs? It appears to still be an inferior (though maybe decent) service compared to ChatGPT etc., and requires very top-end hardware. Is privacy _that_ important to people when their Google search history has been a gateway to the soul for years? I wonder if these machines would cost significantly less (or put the cost to other things, e.g. more CPU cores) without this emphasis on LLMs.
- barrell 7mo agoPrivacy is definitely not a cern for the layman, but it is for lots of people, especially pro users. I also haven’t made a google search in years. I also haven’t seen any improvements in the frontier models in years, and I’m anxiously awaiting local models to catch up.
- NetMageSCW 7mo ago> I also haven’t made a google search in years. That’s makes you so far out at the end of the curve even professionals can’t see you.
- m3kw9 7mo agoA useful llm that needs 64gb of ram and mid double digit cores is not useful for 99% of their customers. The LLMs they have on iphone 17's certainly cannot do anything useful other than summerization and stuff. It's a hardware constraint that they have.
- tiffanyh 7mo ago> Are they doubling down on local LLMs then? Apple is in the hardware business. They want you to buy their hardware. People using Cloud for compute is essentially competitive to their core business.
- causal 7mo ago"Doubling down on already being the best hardware for local inference"
- lakrici88284 7mo ago[dead]
- neya 7mo ago> I still think Apple has a huge opportunity in privacy first LLMs This correlation of Apple and privacy needs to rest. They have consistently proven to be otherwise - despite heavily marketing themselves as "privacy-first" https://www.theguardian.com/technology/2019/jul/26/apple-contractors-regularly-hear-confidential-details-on-siri-recordings https://www.theguardian.com/technology/2019/jul/26/apple-con...
- chaostheory 7mo agoNot for everything. Apple has initially focused on edge AI that runs locally per device. It didn’t work out well the first try, but I would still bet on them trying again once compute catches up. Besides, they still have a better track record than the other tech giants.
- 4fterd4rk 7mo agoI think it's a little telling that the best you can do is a seven year old article.
- lern_too_spel 7mo agoNo other company makes you tell them every application you install on your device. No other company makes you tell them every location you read from your GPS sensor.
- blackqueeriroh 7mo agoPlease, source this ridiculous claim
- neya 7mo agoFirst page, first result on Google: https://andreafortuna.org/2025/11/30/hidden-metadata-reveals-what-your-iphone-silently-records-about-you https://andreafortuna.org/2025/11/30/hidden-metadata-reveals...
- icar 7mo agoDidn't they announce a partnership with Google Gemini?
- ignoramous 7mo ago> doubling down on local LLMs Do think it'll be common to see pros purchasing expensive PCs approaching £25k or more if they could run SoTA multi-modal LLMs faster & locally.
- woadwarrior01 7mo ago> Are they doubling down on local LLMs then? Neural Accelerators (aka NAX) accelerates matmults with tile sizes >= 32. From a very high level perspective, LLM inference has two phases: (chunked) prefill and decode. The former is matmults (GEMM) and the latter is matrix vector mults (GEMV). Neural Accelerators make the former (prefill) faster and have no impact on the latter.
- blueTiger33 7mo agohave you seen that github repo where they unlock the true power of NE?
- recov 7mo agoHave a link?
- caycep 7mo agoGiven all the supply issues w/ Nvidia, I think Apple's AI strategy should be - local AI everything (not just LLMs), but also make Metal competitive w/ CUDA. Their ace in the hole is the unified memory model.
- maherbeg 7mo agoHonestly, they can keep waiting for another year or two for on-device models at the size they're looking for to be powerful enough.
- rafark 7mo ago> Are they doubling down on local LLMs then? I love the push to local llms. But it’s hilarious how apple a few years ago was so reluctant to even mention “AI” in its keynotes and fast forward a couple years they’ve fully embraced it. I mean I like that they embraced it rather than be “different” (stubborn) and stay behind the tech industry. It’s the smart choice. I just think it’s funny.
- Nevermark 7mo ago• Having NPU cores since the M1, would seem to verify that running models has been a game plan for a while. LLMs coming along can only have increased that focus. • Studios with Ultra Mx, now 4-way RDMA over Thunderbolt 5, and enormous RAM and SSD options, suggest a strong focus. I don't know what else that RAM would be intended for. Four Studio Ultras (total of 360 GPU cores with M5 Ultras?) with 2TB of unified RAM is a local model beast. • They refashioned their GPU cores to better support both graphic and neural processing, despite already having focused NPU cores. I would say they have been leaning into local models for several years. I expect we will see more models being optimized for smaller sizes, as demand for them increases. With hardware performance and neural focus trending up, and model requirements/quality trending down, the next few years will be interesting times. What would make me happy: Ultra x 2 (i.e. 2xUltra, 4xMax, 8xPro, 16xM5) packaging in the Studio. With 8-way RDMA. Mac Kong. Perhaps Apple will start making server cards again.
- mvkel 7mo agoIt's more that they can't think of anything else that could possibly need that much compute.