5 ms·
Last Chinese new year we would not have predicted a Sonnet 4.5 level model that runs local and fast on a 2026 M5 Max MacBook Pro, but it's now a real possibilit
by bertili 8mo ago
Last Chinese new year we would not have predicted a Sonnet 4.5 level model that runs local and fast on a 2026 M5 Max MacBook Pro, but it's now a real possibility.
- lostmsu 8mo agoWill 2026 M5 MacBook come with 390+GB of RAM?
- bertili 8mo agoMost certainly not, but the Unsloth MLX fits 256GB.
- embedding-shape 8mo agoCurious what the prefilled and token generation speed is. Apple hardware already seem embarrassingly slow for the prefill step, and OK with the token generation, but that's with way smaller models (1/4 size), so at this size? Might fit, but guessing it might be all but usable sadly.
- regularfry 8mo agoThey're claiming 20+tps inference on a macbook with the unsloth quant.
- embedding-shape 8mo agoYeah, I'm guessing the Mac users still aren't very fond of sharing the time the prefill takes, still. They usually only share the tok/s output, never the input.
- alex43578 8mo agoQuants will push it below 256GB without completely lobotomizing it.
- lostmsu 8mo ago> without completely lobotomizing it The question in case of quants is: will they lobotomize it beyond the point where it would be better to switch to a smaller model like GPT-OSS 120B that comes prequantized to ~60GB.
- lambda 8mo agoIn general, quantizing down to 6 bits gives no measurable loss in performance. Down to 4 bits gives small measurable loss in performance. It starts dropping faster at 3 bits, and at 1 bit it can fall below the performance of the next smaller model in the family (where families tend to have model sizes at factors of 4 in number of parameters) So in the same family, you can generally quantize all the way down to 2 bits before you want to drop down to the next smaller model size. Between families, there will obviously be more variation. You really need to have evals specific to your use case if you want to compare them, as there can be quite different performance on different types of problems between model families, and because of optimizing for benchmakrs it's really helpful to have your own to really test it out.
- Wowfunhappy 8mo ago> In general, quantizing down to 6 bits gives no measurable loss in performance. ...this can't be literally true or no one (including e.g. OpenAI) would use > 6 bits, right?
- alex43578 8mo agoNVIDIA is showing training at 4 bits (NVPF4), and 4 bit quants have been standard for running LLMs at home for quite a while because performance was good enough.
- lambda 8mo agoI mean, GPT-OSS is delivered as a 4 bit model; and apparently they even trained it at 4 bits. Many train at 16 bits because it provides improved stability for gradient descent, but there are methods that allow even training at smaller quantizations efficiently. There was a paper that I had been looking at, that I can't find right now, that demonstrated what I mentioned, it showed only imperceptible changes down to 6 bit quants, then performance decreasing more and more rapidly until it crossed over the next smaller model at 1 bit. But unfortunately, I can't seem to find it again. There's this article from Unsloth, where they show MMLU scores for quantized Llama 4 models. They are of an 8 bit base model, so not quite the same as comparing to 16 bit models, but you see no reduction in score at 6 bits, while it starts falling after that. https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs/unsloth-dynamic-ggufs-on-aider-polyglot https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs/uns... Anyhow, like anything in machine learning, if you want to be certain, you probably need to run your own evals. But when researching, I found enough evidence that down to 6 bit quants you really lose very little performance, and even at much smaller quants the number of parameters tends to be more important than the quantization, all the way down to 2 bits, that it acts as a good rule of thumb, and I'll generally grab a 6 to 8 bit quant to save on RAM without really thinking about, and I try out models down to 2 bits if I need to in order to fit them into my system.
- margorczynski 8mo agoMy hope is the Chinese will also soon release their own GPU for a reasonable price.
- echelon 8mo agoI hope China keeps making big open weights models. I'm not excited about local models. I want to run hosted open weights models on server GPUs. People can always distill them.
- halJordan 8mo agoTheyll keep releasing them until they overtake the market or the govt loses interest. Alibaba probably has staying power but not companies like deepseek's owner
- hmmmmmmmmmmmmmm 8mo agoYeah I wouldn't get too excited. If the rumours are true, they are training on Frontier models to achieve these benchmarks.
- YetAnotherNick 8mo agoWhy does it matter if it can maintain parity with just 6 months old frontier models?
- hmmmmmmmmmmmmmm 8mo agoBut it doesn't except on certain benchmarks that likely involves overfitting. Open source models are nowhere to be seen on ARC-AGI. Nothing above 11% on ARC-AGI 1. https://x.com/GregKamradt/status/1948454001886003328 https://x.com/GregKamradt/status/1948454001886003328
- meffmadd 8mo agoHave you ever used an open model for a bit? I am not saying they are not benchmaxxing but they really do work well and are only getting better.
- Aurornis 8mo agoI have used a lot of them. They’re impressive for open weights, but the benchmaxxing becomes obvious. They don’t compare to the frontier models (yet) even when the benchmarks show them coming close.
- Zababa 8mo agoHas the difference between performance in "regular benchmarks" and ARC-AGI been a good predictor of how good models "really are"? Like if a model is great in regular benchmarks and terrible in ARC-AGI, does that tell us anything about the model other than "it's maybe benchmaxxed" or "it's not ARC-AGI benchmaxxed"?
- doodlesdev 8mo ago
- Aurornis 8mo agoI’m still waiting for real world results that match Sonnet 4.5. Some of the open models have matched or exceeded Sonnet 4.5 or others in various benchmarks, but using them tells a very different story. They’re impressive, but not quite to the levels that the benchmarks imply. Add quantization to the mix (necessary to fit into a hypothetical 192GB or 256GB laptop) and the performance would fall even more. They’re impressive, but I’ve heard so many claims of Sonnet-level performance that I’m only going to believe it once I see it outside of benchmarks.
- PlatoIsADisease 8mo ago'fast' I'm sure it can do 2+2= fast After that? No way. There is a reason NVIDIA is #1 and my fortune 20 company did not buy a macbook for our local AI. What inspires people to post this? Astroturfing? Fanboyism? Post Purchase remorse?
- speedgoose 8mo agoI have a Mac Studio m3 ultra on my desk, and a user account on a HPC full of NVIDIA GH200. I use both and the Mac has its purpose. It can notably run some of the best open weight models with little power and without triggering its fan.
- PlatoIsADisease 8mo ago>with little power and without triggering its fan. This is how I know something is fishy. No one cares about this. This became a new benchmark when Apple couldn't compete anywhere else. I understand if you already made the mistake of buying something that doesn't perform as well as you were expecting, you are going to look for ways to justify the purchase. "It runs with little power" is on 0 people's christmas list.
- speedgoose 8mo agoIt was for my team. Running useful LLMs on battery power is neat for example. Some simply care a bit about sustainability. It’s also good value if you want a lot of memory. What would you advice for people with a similar budget? It’s a real question.
- PlatoIsADisease 8mo agoBut you arent really running LLMs. You just say you are. There is novelty, but not practical use case. My $700, 2023, 3060 laptop runs 8B models. At the enterprise level we got 2, A6000s. Both are useful and were used for economic gain. I don't think you have gotten any gain.
- throwjjj 8mo ago[dead]