4 ms·
14-billion parameter model with 4-bit quantization seems rather small
by lastdong 7mo ago
14-billion parameter model with 4-bit quantization seems rather small
- simlevesque 7mo agoIt's not much for a frontier AI but it can be a very useful specialized LLM.
- giancarlostoro 7mo agoOn my 24GB RAM M4 Pro MBP some models run very quickly through LM Studio to Zed, I was able to ask it to write some code. Course my fan starts spinning off like the worlds ending, but its still impressive what I can do 100% locally. I can't imagine on a more serious setup like the Mac Studio.
- efxhoy 7mo agoHow is the output quality of the smaller models?
- elsombrero 7mo agonot good enough for coding anything more than simple scripts. generally, the less parameters, the less knowledge they have.
- kraig911 7mo agowhat model were you using?
- giancarlostoro 7mo agoWrote about it here: https://news.ycombinator.com/item?id=47191915 https://news.ycombinator.com/item?id=47191915
- jbellis 7mo agoYour limitation after prefill is memory bandwidth. A maxed out Studio has less than a single 3090 (really).
- sroussey 7mo agoYeah, the 3090 has faster memory, but not by a lot. The 5090 is at 1,792GB/sec and potential M5 Ultra would be 1,230GB/sec and 512GB RAM. Maybe 1TB. Not 32.
- thejazzman 7mo agoYou’re suggesting that a difference of the entirety of the M5 Max’s bandwidth is an insignificant gap!
- veidr 7mo agoNo, that difference is the 5090, not the 3090.
- butILoveLife 7mo agoFor anyone who has been watching Apple since the iPod commercials, Apple really really has grey area in the honesty of their marketing. And not even diehard Apple fanboys deny this. I genuinely feel bad for people who fall for their marketing thinking they will run LLMs. Oh well, I got scammed on runescape as a child when someone said they could trim my armor... Everyone needs to learn.
- zitterbewegung 7mo agoYesterday I ran qwen3.5:27b with an M1 Max and 64 GB of ram. I have even run Llama 70B when llama.cpp came out. These run sufficiently well but somewhat slow but compared to what the improvements with the M5 Max it will make it a much faster experience.
- giwook 7mo agoI don't know that there would be a huge overlap between the people who would fall for this type of marketing and the people who want to run LLMs locally. There definitely are some who fit into this category, but if they're buying the latest and greatest on a whim then they've likely got money to burn and you probably don't need to feel bad for them. Reminds me of the saying: "A fool and his money are soon parted".
- nine_k 7mo agoThere used to be a polite way to call this out, the "Steve Jobs's reality distortion field".
- hamdingers 7mo agoNow that every CEO has their own reality distortion field I wonder if it's even worth calling out any more.
- nine_k 7mo agoMost are not nearly as smooth and successful at the distorting.
- 7mo ago
- bilbo0s 7mo agoIt is. That's how they make loot on their 128GB MacBook Pros. By kneecapping the cheap stuff. Don't think for a second that the specs weren't chosen so that professional developers would have to shell out the 8 grand for the legit machine. They're only gonna let us do the bare minimum on a MacBook Air.
- derefr 7mo agoI think these aren't meant to be representative of arbitrary userland-workload LLM inferences, but rather the kinds of tasks macOS might spin up a background LLM inference for. Like the Apple Intelligence stuff, or Photos auto-tagging, etc. You wouldn't want the OS to ever be spinning up a model that uses 98% of RAM, so Apple probably considers themselves to have at most 50% of RAM as working headroom for any such workloads.
- duskwuff 7mo agoAlso: they're advertising the degree of improvement ("4x faster"), not an absolute level of performance.