10 ms·
While everyone has focused on Apple's power-efficiency on the M series chips, one thing that has been very interesting is how powerful the unified memory model
by InTheArena 2y ago
While everyone has focused on Apple's power-efficiency on the M series chips, one thing that has been very interesting is how powerful the unified memory model (by having the memory on-package with CPU) with large bandwidth to the memory actually is. Hence a lot of people in the local LLMA community are really going after high-memory Macs.
It's great to see NPUs here with the new Ryzen cores - but I wonder how effective they will be with off-die memory versus the Apple approach.
That said, it's nothing but great to see these capabilities in something other then a expensive NVIDIA card. Local NPUs may really help with edge deploying more conferencing capabilities.
Edited - sorry, ,meant on-package.
- vbezhenar 2y agoApple does not make on-die RAM.
- chaostheory 2y agoWhat Apple has is theoretically great on paper, but it fails to live up to expectations. Whats the point of having the RAM for running an LLM locally when the performance is abysmal compared to running it on even a consumer Nvidia GPU. It’s a missed opportunity that I hope either the M4 or M5 addresses
- bearjaws 2y agoIt's a 25w processor. How will it ever live up to a 400w GPU? Also you can't even run large models on a single 4090, but you can on M series laptops with enough RAM. The fact a laptop can run 70B+ parameter models is a miracle, it's not what the chip was built to do at all.
- chaostheory 2y agoThe problem is that it extends to both Mac Studio and Mac Pro.
- wongarsu 2y agoIt's a valid comparison in the very limited sense of "I have $2000 to spend on a way to run LLMs, should I get an RTX4090 for the computer I have or should I get a 24GB MacBook", or "I have $5000, should I get an RTX A6000 48GB or a 96GB MacBook". Those comparisons are unreasonable in a sense, but they are implied by statements like GPs "Hence a lot of people in the local LLMA community are really going after high-memory Macs".
- fckgw 2y agoNo, it is not a valid comparison to make between an entire laptop and a single PC part. "The computer I have" is doing a ton of heavy lifting here.
- wongarsu 2y ago"Should I upgrade what I have or buy something new" is a completely normal everyday decision. Of course it doesn't apply to everyone since it presumes you have something compatible to upgrade, but it is a real decision lots of people are making
- fckgw 2y agoBut it's also assuming everyone has a desktop PC capable of this stuff that can be upgraded.
- michaelt 2y agoSorta yes, sorta no. You're certainly right that with a macbook you get a whole computer, so you're getting more for your money. And it's a luxury high-end computer too! But personally, I've never seen anyone step directly from not-even-having-a-PC to buying a 4090 for $1800. Folks that aren't technically inclined by and large stick with hosted models like ChatGPT. More common in my experience is for technical folks with, say, an 8GB GPU to experiment with local ML, decide they're interested in it, then step up to a 4090 or something.
- oceanplexian 2y agoThe answer depends on what you plan to do with it. Do you need to do fine tuning on a smaller model and need the highest inference performance with smaller models? Are you planning to use it as a lab to learn how to work with tools that are used in Big Tech (i.e. CUDA)? Or do you just want to do slow inference on super huge models (e.g Grok)? Personally, I chose the Nvidia route because as a backend engineer, Macs aren’t seriously used in datacenters. The same frameworks I use to develop on a 3090 are transferable to massive, infiniband-connected clusters with TB of VRAM.
- dragonwriter 2y ago> The fact a laptop can run 70B+ parameter models is a miracle Most laptops with 64+GB of RAM can run a 70B model at 4-bit quantization. It’s not a miracle, it’s just math. M2 can do it faster than systems with slower memory bandwidth.
- deleted 2y ago[deleted]
- InTheArena 2y agoThe performance of oolama on my M1 MAX is pretty solid - and does things that my 2070 GPU can't do because of memory.
- dangus 2y agoNot that I don’t believe you but the 2070 is two generations and 5 years old. Maybe a comparison to a 4000 series would be more appropriate?
- Kirby64 2y agoThe M1 Max is also 2 generations old, and ~3 years old at this point. Seems like a fair comparison to me.
- dangus 2y agoThe 4000 series still has a bigger gap in how much of a generational leap that product was. The M3 Max has something like 33% faster overall graphics performance than the M1 Max (average benchmark) while the 4090 is something like 138% faster than the 2080Ti. Depending on which 2070 and 4070 models you compare the difference is similar, close to or exceeding 100% uplift.
- whizzter 2y agoGoogling power draw the 4090 goes up to 450w whilst the 2080ti was at 250w, adjusting for power consumption the increase is somewhere around 32%. Some architectural gains and probably optimizations in chipset workings but we're not seeing as many amazing generational leaps anymore regardless of manufacturer/designer.
- dangus 2y agoI’m still seeing over a 100% uplift comparing mobile to mobile on Nvidia products: https://gpu.userbenchmark.com/Compare/Nvidia-RTX-4090-Laptop-vs-Nvidia-RTX-2080-Mobile/m2036852vsm700275 https://gpu.userbenchmark.com/Compare/Nvidia-RTX-4090-Laptop... As far as desktop products, power consumption is irrelevant.
- zitterbewegung 2y agoBuying a m3 max with 128gb of RAM while will underperform any consumer NVIDIA GPU it will be able to load larger models in practice but slowly. I think a way for the m series chips to aggressively target GPU inference or training would need a strategy that increases the speed of the RAM to start to match GDDR6 or HBM3 or use it directly.
- deleted 2y ago[deleted]
- chaostheory 2y agoYou summed up my point better than I did
- evilduck 2y agoThat completely depends on your expectations and uses. I have a gaming rig with a 4080 with 16GB of RAM and it can't even run Mixtral (kind of the minimum bar of a useful generic LLM in my opinion) without being heavily quantized. Yeah it's fast when something fits on it, but I don't see much point in very fast generation of bad output. A refurbished M1 Max with 32GB of RAM will enable you to generate better quality LLM output than even a 4090 with 24GB of VRAM and for ~$300 less, and it's a whole computer instead of a single part that still needs a computer around it. Compared to my 4080, that GPU and the surrounding computer get you half the VRAM capacity for greater cost than the Mac. If you're building a rig with multiple GPUs to support many users or for internal private services and are willing to drop more than $3k then I think the equation swings back in favor of Nvidia, but not until then.
- soupbowl 2y agoJust buy another 16gb of ram for 80$....
- elzbardico 2y agoGPU ram?
- evilduck 2y agoRunning your larger-than-your-GPU-VRAM LLM model on regular DDR ram will completely slaughter your token/s speed to the point that the Mac comes out ahead again.
- programd 2y agoDepends on what you're doing. Just chatting with the AI? I'm getting about 7 tokens per sec for Mistral with the Q6_K on a bog standard Intel i5-11400 desktop with 32G of memory and no discrete GPU (the CPU has Intel UHD Graphics 730 built in). 2 year old low end CPU that goes for, what $150? these days. As far as I'm concerned that's conversational speed. Pop in some 8 core modern CPU and I'm betting you can double that, without even involving any GPU. People way overestimate what they need in order to play around with models these days. Use llama.cpp and buy that extra $80 worth of RAM and pay about half the price of a comparable Mac all in. Bigger models? Buy more RAM, which is very cheap these days. There's a $487 special on Newegg today with an i7-12700KF, motherboard and 32G of ram. Add another $300 worth of case, power supply, SSD and more RAM and you're under the price of a Macbook Air. There's your LLM inference machine (not for training obviously) which can run even the 70B models at home at acceptable conversational speed.
- deleted 2y ago[deleted]
- instagib 2y agoOne thing I would consider is usage throttling on a MacBook Pro. Would repeated LLM usage run into throttling? No idea what specifically everyone is pulling their performance data from or what task(s). Here is a video to help visualize the differences with a maxed out m3 max vs 16gbm1 pro vs 4090 on llm 7B/13b/70b llama 2. https://youtu.be/jaM02mb6JFM https://youtu.be/jaM02mb6JFM Here’s a Reddit comparison of 4090 vs M2 Ultra 96gb with tokens/s https://old.reddit.com/r/LocalLLaMA/comments/14319ra/rtx_40906000_vs_m2_max_with_96gb_unified_memory/ https://old.reddit.com/r/LocalLLaMA/comments/14319ra/rtx_409... M3 pro memory BW 150 gb/s M3 max 10/30 300 gb/s M3 max 12/40 400 gb/s “Llama models are mostly limited by memory bandwidth. rtx 3090 has 935.8 gb/s rtx 4090 has 1008 gb/s m2 ultra has 800 gb/s m2 max has 400 gb/s so 4090 is 10% faster for llama inference than 3090 and more than 2x faster than apple m2 max https://github.com/turboderp/exllama https://github.com/turboderp/exllama using exllama you can get 160 tokens/s in 7b model and 97 tokens/s in 13b model while m2 max has only 40 tokens/s in 7b model and 24 tokens/s in 13b apple 40/s Memory bandwidth cap is also the reason why llamas work so well on cpu (…) buying second gpu will increase memory capacity to 48gb but has no effect on bandwidth so 2x 4090 will have 48gb vram and 1008 gb/s bandwidth and 50% utilization”
- john_alan 2y agowhat are you talking about, it's literally the fastest single core retail CPU globally, and multicore is close too - https://browser.geekbench.com/processor-benchmarks https://browser.geekbench.com/processor-benchmarks
- edward28 2y agoI would take those benchmarks with a grain of salt, given that it shows a 96 core epyc loosing in multi thread to a 64 core epyc and a 32 core xeon.
- jwr 2y agoHmm. I'm running decent LLMs locally (deepseek-coder:33b-instruct-q8_0, mistral:7b-instruct-v0.2-q8_0, mixtral:8x7b-instruct-v0.1-q4_0) on my MacBook Pro and they respond pretty quickly. At least for interactive use they are fine and comparable to Anthropic Opus in speed. That MacBook has an M3 Max and 64GB RAM. I'd say it does live up to my expectations, perhaps even slightly exceeds them.
- atty 2y agoApples memory is on package, not on die.
- thsksbd 2y agoOld becomes new, the SGI O2 had (off chip) a unified memory model for performance reasons. Not a CS guy, but it seems to me that NUMA like architecture has to come back. Large RAM on chip (balancing a thermal budget between #ofcores vs ram), a much larger RAM off chip and even more RAM through a fast interconnect on a single kernel image. Like the Origin 300 had.
- Rinzler89 2y agoUMA in the SGI machines (and gaming consoles) made sense because all the memory chips at that time were equally slow, or fast, depending how you wanna look at it. PC HW split the video memory from system memory once GDDRAM become so much faster than system RAM, but GDDRAM has too high latency for CPUs and DDR has too low bandwidth for GPUs, so the separation made sense for each's strengths and still does to this day. Unifying it again, like with AMD's APUs, means either compromises for the CPU or for the GPU. There's no free lunch. Currently AMD APUs on the PC use unified DDRAM so CPU performance is top but GPU/NPU perforce is bottlenecked. If they were to use unified GDDRAM like in the PS5/Xbox then GPU/NPU performance would be top and CPU performance would be bottlenecked.
- Dalewyn 2y ago>Unifying it again, like with AMD's APUs, means either compromises for the CPU or for the GPU. There's no free lunch. I think the lunch here (it still ain't free) is that RAM speed means nothing if you don't have enough RAM in the first place, and this is a compromise solution to that practical problem.
- Rinzler89 2y ago>if you don't have enough RAM in the first place Enough RAM for what task exactly? System RAM is plentiful and cheap nowadays(unless you buy Apple). I got new a laptop with 32GB RAM for about 750 Euros. But the speeds are too low for high-end gamming or LLM training for the poor APU.
- 2y ago
- v1sea 2y agoedit: I was wrong.
- deleted 2y ago[deleted]
- oflordal 2y agoYou can do that on both HIP and cuda through e.g. hipHostMalloc and the cuda equivalent (Not officially supported on the AMD APUs but works in practice). With a discrete GPU the GPU will access memory across PCIe but on an APU it will go full speed to RAM as far as I can tell.
- smallmancontrov 2y ago> low latency results each frame What does that do to your utilization? I've been out of this space for a while, but in game dev any backwards information flow (GPU->CPU) completely murdered performance. "Whatever you do, don't stall the pipeline." Instant 50%-90% performance hit. Even if you had to awkwardly duplicate calculations on the CPU, it was almost always worth it, and not by a small amount. The caveat to "almost" was that if you were willing to wait 2-4 frames to get data back, you could do that without stalling the pipeline. I didn't think this was a memory architecture thing, I thought it was a data dependency thing. If you have to finish all calculations before readback and if you have to readback before starting new calculations, the time for all cores to empty out and fill back up is guaranteed to be dead time, regardless of whether the job was rendering polygons or finite element calculations or neural nets. Does shared memory actually change this somehow? Or does it just make it more convenient to shoot yourself in the foot? EDIT: or is the difference that HPC operates in a regime where long "frame time" dwarfs the pipeline empty/refill "dead time"?
- v1sea 2y agoIt was probably from my workloads being relatively small that I could get away with 90Hz read on the cpu side. I'll need to dig deeper into it. The metrics I was seeing were showing 200-300 microseconds of GPU time for physics calculations and within the same frame the cpu reading from that buffer. Maybe I'm wrong, need to test more.
- numpad0 2y agoNote that while UMA is great in the sense that they allow LLM models to be run at all, M-series chips aren't faster[1] when the model fits in VRAM. 1: screenshot from[2]: https://www.igorslab.de/wp-content/uploads/2023/06/Apple-M2-ULtra-SoC-Geekbench-5-OpenCL-Compute.jpg 2: https://wccftech.com/apple-m2-ultra-soc-isnt-faster-than-amd-intel-last-year-desktop-cpus-50-slower-than-nvidia-rtx-4080/
- cstejerean 2y agoThe problem is you're limited to 24 GB of VRAM unless you pay through the nose for datacenter GPUs, whereas you can get an M-series chip with 128 GB or 192 GB of unified memory.
- numpad0 2y agoSurely! The point is that they're not million times faster magic chips that makes NVIDIA bankrupt tomorrow. That's all. A laptop with up to 128GB "VRAM" is a great option, absolutely no doubt about that.
- john_alan 2y agoThey are powerful, but I agree with you, it's nice to be able to run Goliath locally, but it's a lot slower than my 4070.
- paulmd 2y agothat's openCL compute, LLM models ideally should be hitting the neural accelerator, not running on generalized gpu compute shaders.
- spamizbad 2y agoMy understanding is the unified RAM on the M-series die does not contribute significantly to their performance. You get a little bit better latency but not much. The real value to Apple is likely it greatly simplifies your design since you don't have to route out tons of DRAM signaling and power management on your logic board. Might make DRAM training easier too but that's way beyond my expertise.
- john_alan 2y agoalso provides the GPU with serious RAM allocation, a 64GB M3 chip comes with more ram for the GPU than a 4090
- rowanG077 2y agoSo why are you proposing Apple did that if not for performance as you claim? They waste a lot of silicon for those extra memory controllers. Basically a comparable amount to all the CPU cores in an M1 Max.
- AceJohnny2 2y ago> unified memory model (by having the memory on-package with CPU) That's not what "unified memory model" means. It means that the CPU and GPU (and ANE!) have access to the same banks of memory, unlike PC GPUs that have their own memory, separated from the CPU's by the PCIe bottleneck (as fast as that is, it's still smaller than direct shared DRAM access). It allows the hardware more flexibility in how the single pool of memory is allocated across devices, and faster sharing of data across devices. (throughput/latency depends on the internal system bus ports and how many each device have access to) The Apple M-Series chips also has the memory on-package with the CPU (technically SoC, "System-on-Chip"), but that provides different benefits.
- cmovq 2y agoHaving separated GPU memory also has its benefits. Once the data makes it through the PCIe bus, graphics memory typically has much higher bandwidth which also doesn’t need to split with the CPU.
- crawshaw 2y agoAn M2 Ultra has 800GB/s of memory bandwidth, an Nvidia 4090 has 1008GB/s. Apple have chosen to use relatively little system memory at unusually high bandwidth.
- mmaniac 2y agoThe benefit when having an ecosystem of discrete GPUs is that CPUs can get away with having low bandwidth memory. This is great if you want motherboards with socketed CPU and socketed RAM which are compatible with the whole range of product segments. CPUs don't really care about memory bandwidth until you get to extreme core counts (Threadripper/Xeon territory). Mainstream desktop and laptop CPUs are fine with just two channels of reasonably fast memory. This would bottleneck an iGP, but those are always weak anyway. The PC market has told users who need more to get a discrete GPU and to pay the extra costs involved with high bandwidth soldered memory only if they need it. The calculation Apple has made is different. You'll get exactly what you need as a complete package. You get CPU, GPU, and the bandwidth you need to feed both as a single integrated SoC all the way to the high end. Modularity is something PC users love but doing away with it does have advantages for integration.