3 ms·
Apple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will
by vmware508 2mo ago
Apple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions.
Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.
- Gecko4072 2mo agoThey will cost an insane amount as well. Maybe less than subscriptions or tokens. But running massive models on laptops with batteries and poor cooling doesn’t make much sense.
- LeBit 2mo agoUntil hiding PII from the cloud LLM is a resolved issue, running local LLMs will remain a necessity. There are workplaces that refuse to use LLMs because they fear the devs will expose sensitive data without care.
- ignoramous 2mo ago> will cost an insane amount as well We will get to a point where prosumer laptops that etch SoTA LLMs in removable silicon will be as expensive as cars.
- dannyw 2mo agoIf you're a company with a considerable API bill, buying a few Blackwells and a rack could make a lot of economic sense, while still giving your security team full control of the infrastructure
- schleck8 2mo agoYou'd need the 256 gb memory model which will be expensive because apple has trouble getting capacity (got turned down by cxmt). And even then you can only run a 2 bit quant which is noticeably worse than 8 bit
- Flavius 2mo ago> run free LLMs locally at native speed This reads like a hallucination. What does native speed even mean?
- kyxsc 2mo agofor example, models running at like 100-150 tokens/second (or faster!) vs 15 t/s (fable/sol are ~60 t/s, and OpenAI just announced their Cerebras partnership(?) for "ultrafast" mode of 750 t/s) models aren't able to run that fast right now on our consumer/prosumer hardware. M5 Max for example has a memory bandwidth of 600 GB/s. a 5090 has 3x that, so running the same model on a 5090 is that much faster (provided the model is within 30GB). running a bigger model on an M5 Ultra is still much slower than running it on a Blackwell chip with sufficient vram, CUDA being a major difference. if apple can bridge this gap, interesting things will happen... and just imagine if M7 Ultra has comparable speeds to Blackwell (or even Rubin)!
- toasty228 2mo agoMeanwhile the GB300 used by hosted llms: GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional https://pi3g.com/nvidia-gb300-specifications-including-memory-bandwidth-and-llm-benchmarks-based-on-2026-systems/ https://pi3g.com/nvidia-gb300-specifications-including-memor... If you think M7 will hit even 15% of these speeds you're very optimistic.
- webbrain 2mo ago
- re-thc 2mo ago> Apple will release M7 MacBook Pros / Mac Minis next year The latest on Apple is TSMC is stuck on the next iPhone due to lack of RAM. Good luck getting any Macs. Memory shortage is getting worse.
- toasty228 2mo agoSure buddy, all you'll end up with is a $10k machine that run gimped models at like 30tok/s for about 5m before the fan kicks in and it starts to sound like a turboprop, while offering maybe 30% of the context size of hosted models.
- scotty79 2mo agoI don't know why you'd want to burden your laptop with a large model. But I can totally see a new "developer workstation" product that's just a semi-large box that's optimized for running frontier open weights models for one to few users.
- gehsty 2mo agoLocal vs remote compute is a constant thread in tech history - mainframes and desktops then local and cloud compute (think Google Photos bs Apple photos - one indexes on device the other indexes in cloud). Now we have the next chapter local vs cloud LLM models. There will always be a market for frontier labs in the cloud based models - these models will always be able to be bigger, and that will likely translate to doing things local models can’t. Logically also we’ll likely get to a point where RAM drops in price as production ramps up, and local LLM is both capable and cost effective. This feels like it is coming for Siri / Gemini / Alexa personal assistant type use cases. So I think the local LLM will become a thing in laptops and phones in a year or two, offering PA type use cases. Professional LLM services will likely remain at the frontier (and in the cloud) for the foreseeable.
- layer8 2mo agoThe RAM shortage situation won’t be sorted out within the next year.
- rammler 2mo agoBold to believe it will be sorted at all
- fearmerchant 2mo agoIf the margins are there it will get sorted.
- Havoc 2mo agoThat’s not how that works. The hosted models don’t stay still in size and capability while Apple advances. Both will advance their frontier and there will still be a gap and developers will still prefer the stronger option.
- WarmWash 2mo agoI'm still waiting for Linux to topple Windows
- bigbadfeline 2mo agoThe processors running the Linux kernel outnumber those running Windows 4:1.
- nater5000 2mo agoI'll give you credit for at least offering a specific, somewhat unique take. But this is a pretty dumb take lol
- kube-system 2mo agoEvery single MacBook built in the past half-decade already has an LLM built into the latest version of their OS. But there's a significant difference in hardware required between running a 3B parameter model and a 700B-1T+ parameter model.