4 ms·
Apple has totally failed to deliver interesting AI experiences so far ... and I still think they're going to be the dominant provider of AI in 5 years. We're ju
by habosa 3mo ago
Apple has totally failed to deliver interesting AI experiences so far ... and I still think they're going to be the dominant provider of AI in 5 years. We're just one or two advances in chips / models / both away from being able to run very good local models for free on mid-tier Apple devices. The privacy, cost, and latency story there will be too much for OpenAI/Anthropic/Google to beat.
Just writing this down so I can be praised/mocked in 5 years.
- giancarlostoro 3mo agoIts a shame RAM prices are getting in the way of Apple. The 512 GB RAM Mac Studios would have been worthwhile.
- QuercusMax 3mo agoI'm also of this opinion, but also that it doesn't have to be Apple (but they are well positioned). What I've seen with running local models on my 48GB M4 MBP is really impressive - it's not the same level as hosted stuff, but it's better than what I was using a year or two ago.
- dabbz 3mo agoI suspect we'll see a hybrid before an all or nothing. Local models for computer control or delegating, online models for things that need strong reasoning, planning, and knowledge access. Again, I'd be more than happy to be wrong. I just see models growing faster than the hardware can.
- xyst 3mo agoApple has failed to live up to the Steve Jobs era and initial iPhone hype. It's not just "AI experiences", it's computing in general. Maybe the consumer sector is just dead/dying. Maybe consumers are just running out of cash because filthy VCs are destroying communities and forcing the 99% into poverty. The one thing that is marginally exciting: the Apple SoC or M series chips. It's unfortunate they are locked behind crappy macOS and other proprietary apple crap.
- overfeed 3mo ago> it's computing in general. Unsurprising. Apple seriously thought the iPad would replace computers and usher in a "post-PC" word during their "What is a computer?" ad campaign era. Now they are sticking phone chips in laptop chassis.
- mlsu 3mo agoFor the general consumer, they were basically right though. Most people don't use laptops except for work. The primary computing device is the phone, and phones have basically become become mini-ipads in form factor since that ad aired.
- dwaite 3mo ago> Apple seriously thought the iPad would replace computers ...for some users. See their "Mac is a truck" analogy. And it has. My parents haven't owned a Windows or Mac machine in six years, since they got rid of the one I gifted to them a decade ago. Its all iPad and iPhone.
- ingenieroariel 3mo agoI wrote: "we should all be buying a fully loaded Mac Studio (128GB of ram, 20 CPU cores, a lot of GPU and Neural cores.)" April, 2023 We are both late and early. https://news.ycombinator.com/item?id=35527692 https://news.ycombinator.com/item?id=35527692
- bigyabai 3mo agoYou should not buy a fully loaded Mac Studio for AI unless you absolutely NEED macOS. You will be wasting so much electricity idling on prefill while your GPU pulls 150-250w from the wall. Buy an Nvidia Spark, then whatever cheap Mac you want to use as a thin client. There's no reason to force Apple Silicon's round peg into a square hole like AI inference.
- deleted 3mo ago[deleted]
- timy2shoes 3mo agoBenchmarking that I've seen shows that the M5 Max outperforms the DGX Spark, e.g. https://www.reddit.com/r/LocalLLaMA/comments/1tfzsd6/m5_vs_dgx_spark_vs_strix_halo_vs_rtx_6000/ https://www.reddit.com/r/LocalLLaMA/comments/1tfzsd6/m5_vs_d... or https://www.reddit.com/r/LocalLLaMA/comments/1tr7hzw/psa/ https://www.reddit.com/r/LocalLLaMA/comments/1tr7hzw/psa/. Seems to me that Apple is doing pretty well with local AI inference.
- bigyabai 3mo agoOutperforms doing what? Inference is not a homogeneous workload, memory bandwidth correlates to decode speed and layer swapping but not necessarily inference speed overall. The other half of that equation is latency, predicated on prefill performance which needs a powerful GPU and ideally ALU-level optimization to build larger KV caches quickly. Even the M5 gets smoked in this department, the M5 Max has a 50% longer TTFT on Qwen's 27b dense model at only 16k of context, which is a pretty typical starting context to use for agentic editing in normal apps like OpenCode/Claude Code: https://raw.githubusercontent.com/Osmantic/MMBT-Messy-Model-Bench-Tests/0adec7c5e9a91ef19b0f1385e9d9d2b589fee45c/hardware-tests/qwen3.6-q8-fleet-2026-05-17/aggregate/canonical-headline.json https://raw.githubusercontent.com/Osmantic/MMBT-Messy-Model-... For agentic, 50-256k token on-device coding sessions, the Spark will be faster and consume less power running larger models. Without an external GPU (which Apple doesn't support), Apple Silicon will always be bottlenecked during prefill. Apple's failure to address this with their GPU architecture is a big reason why Apple Silicon viewed as a waste of time and money for professional datacenter deployment.
- ransom1538 3mo ago"Apple has totally failed to deliver interesting AI experiences so far ..." You think this is a mistake...
- nvme0n1p1 3mo agohttps://www.crikey.com.au/2025/01/08/apple-new-artificial-intelligence-rewords-scam-messages-look-legitimate/ https://www.crikey.com.au/2025/01/08/apple-new-artificial-in... Of course. Do you think this was on purpose? All part of Apple's brilliant master plan?
- _verandaguy 3mo agoI'm perfectly happy with Apple not becoming an "everything we do is AI-centric" business. I'm fatigued by it all at this point. It's streamlining the interesting and fun parts out of my job (by practical necessity of use there), and if I used it half as much outside of work I'm sure it'd do the same there too.
- operatingthetan 3mo ago>I'm fatigued by it all at this point. This is the prevailing opinion of people even outside of tech.
- _verandaguy 3mo agoI know public opinion polling supports that, but the parts of my social circle which are outside of tech seem to be, at worst, apathetic (and at best enthusiastic, though that's not a big fraction). That said, I think it's a good thing that this sentiment is coming to the forefront.
- nsagent 3mo agoI was actually surprised to hear my brother-in-law deride LLMs as being useless for areas he has expertise in when I visited for the 4th of July. He was complaining that he would ask how to perform a certain repair on a car, and the LLMs he tried (ChatGPT & Grok) would give him a long involved process and he'd ask why not do it this simpler way and it would say, oh you're right! He just found it gave bad advice and realized (rightly) that in areas he has less expertise in he has no way to judge how good the outputs are. This is from a guy who loves tech, historically worshipped Elon, loves his Tesla, and (rightfully again) didn't buy into SpaceX because he thought it was overvalued. In the past when I visited for holidays he was liable to have a positive outlook on LLMs and their utility. Seems telling that he's starting to see the cracks.
- operatingthetan 3mo agoYep. In my areas of expertise I can easily catch it coming up with wrong information, bad calculations etc. So laypeople are probably being led astray quite often.
- thesurlydev 3mo agoDid you mean 5 months? :)
- ceroxylon 3mo agoI don't mind the subtle ML integrations that they have put in the photos app: plant ID, recognizing faces, removing background, OCR text search (even for handwriting!), etc.
- rdn 3mo agoI don't like when I'm trying to copy a photo online and it copies some piece of text in the photo instead
- eastbound 3mo agoYou’re listing every feature I’d like to remove; Meanwhile Android users have great holiday pictures at the bottom of the Pyramids because they can remove the people from their pictures.
- j45 3mo agoI'm not sure that's what this article is about. Apple is doing something very different. Their AI experience for end users definitely has been a little behind. Apple Silicon, however, has been quite unique for the last 4-6 years and it's increasing overlap with LLMS. The model/chip optimizations are definitely improvements, the thing that is really standing out the past 2 years is how much the open source model community has been making possible, especially when you know a group of use cases.
- ttul 3mo agoHere’s the two main reasons why local inference won’t compete any time soon with the cloud: 1. Most useful LLM work is done in parallel. A Mac Mini can run one LLM inference thread at a time. The cloud can spool up dozens and spread that inference across efficiently batched operations over a fleet of hardware. 2. Faster inference hardware such as the chips from Cerebras and Groq cannot be run locally. But the advantages of running >5x the token throughput per thread can’t be overstated. Add in the multi-threading advantage and it’s a knock-out punch for local LLMs. Local inference has a role: if you’re working with extremely private matters or you want an uncapped model that will talk dirty or generate NSFW photos, local is the only option. I think Apple and others will continue to also run a lot of useful workloads locally such as text editing suggestions, speech to text, text to speech, and image manipulation. As local hardware improves, these capabilities will get better too. But, for most LLM work, the cloud will continue to dominate for a long time to come, if not forever.
- varispeed 3mo agoYou can always buy multiple Macs. I think Apple's great differentiator could be making frontier class models local. I you want more threads, buy more Macs. I don't want to run any workflows on someone else's computers.
- 7speter 3mo agoAs far as I know, there isn’t an interface like nvlink that allows these macs to work in tandem; they would just send their data over ethernet, maybe thunderbolt/usb-c?
- fragmede 3mo agoWhy yes you can! https://www.jeffgeerling.com/blog/2025/15-tb-vram-on-mac-studio-rdma-over-thunderbolt-5/ https://www.jeffgeerling.com/blog/2025/15-tb-vram-on-mac-stu...
- anon373839 3mo ago> A Mac Mini can run one LLM inference thread at a time. That’s not accurate. With MLX, at least, parallel inference is both possible and useful. Model serving tools like LM Studio and oMLX support parallel generation with continuous batching, and the total throughput increases with it.
- rk06 3mo agoThere is common wisdom: during gold rush, sell shovels. Apple's shovel (ahem, Mac mini) is the highest quality.with Companies burning money left, right and center, Apple can dispense with advertising altogether