10 ms·
The M4 Max goes up to 128GB RAM, and "over half a terabyte per second of unified memory bandwidth" - LLM users rejoice.
by BrentOzar 2y ago
The M4 Max goes up to 128GB RAM, and "over half a terabyte per second of unified memory bandwidth" - LLM users rejoice.
- manaskarekar 2y agoThe M3 Max was 400GBps, this is 540GBps. Truly an outstanding case for unified memory. DDR5 doesn't come anywhere near.
- vid 2y agoIt's not "DDR5" on its own, it's a few factors. Bandwidth (GB/s) = (Data Rate (MT/s) * Channel Width (bits) * Number of Channels) / 8 / 1000 (8800 MT/s * 64 bits * 8 channels) / 8 / 1000 = 563.2 GB/s This is still half the speed of a consumer NVidia card, but the large amounts of memory is great, if you don't mind running things more slowly and with fewer libraries.
- cjbprime 2y agoRight, the nvidia card maxes out at 24GB.
- vid 2y agoA 24gb model is fast and ranks 3. A 70b model is slow and 8. A top tier hosted model is fast and 100. Past what specialized models can do, it's about a mixture/agentic approach and next level, nuclear power scale. Having a computer with lots of relatively fast RAM is not magic.
- manaskarekar 2y agoThanks, but just to put things into perspective, this calculation has counted 8 channels which is 4 DIMMs and that's mostly desktops (not dismissing desktops, just highlighting that it's a different beast). Most laptops will be 2 DIMMS (probably soldered).
- wtallis 2y agoDesktops are two channels of 64 bits, or with DDR5 now four (sub)channels of 32 bits; either way, mainstream desktop platforms have had a total bus width of 128 bits for decades. 8x64 bit channels is only available from server platforms. (Some high-end GPUs have used 512-bit bus widths, and Apple's Max level of processors, but those are with memory types where the individual channels are typically 16 bits.)
- sliken 2y agoI think you are confusing channels and dimms. The vast majority of any x86 laptop or desktops are 128 bits wide. Often 2x64 bit channels up till last year or so, now 4x32 bit DDR5 in the last year or so. There are some benefits to 4 channels over 2, but generally you are still limited by 128 bits unless you buy a Xeon, Epyc, or Threadripper (or Intel equiv) that are expensive, hot, and don't fit in SFFs or laptops. So basically the PC world is crazy behind the 256, 512, and 1024 bit wide memory busses apple has offered since the M1 arrived.
- wtallis 2y ago> (8800 MT/s * 64 bits * 8 channels) / 8 / 1000 = 563.2 GB/s Was this example intended to describe any particular device? Because I'm not aware of anything that operates at 8800 MT/s, especially not with 64-bit channels.
- sliken 2y agoM4 max in the MBP (today) and in the Studio at some later date.
- wtallis 2y agoThat seems unlikely given the mismatched memory speed (see the parent comment) and the fact that Apple uses LPDDR which is typically 16 bits per channel. 8800MT/s seems to be a number pulled out of thin air or bad arithmetic.
- sliken 2y agoHeh, ok, maybe slightly different. But apple spec claims 546GB/sec which works out to 512 bits (64 bytes) * 8533. I didn't think the point was 8533 vs 8800. I believe I saw somewhere that the actual chips used are LPDDR5X-8533. Effectively the parents formula describes the M4 max, give or take 5%.
- sliken 2y agoFewer libraries? Any that a normal LLM user would care about? Pytorch, ollama, and others seem to have the normal use cases covered. Whenever I hear about a new LLM seems like the next post is some mac user reporting the token/sec. Often about 5 tokens/sec for 70B models which seems reasonable for a single user.
- vid 2y agoIs there a normal LLM user yet? Most people would want their options to be as wide as possible. The big ones usually get covered (eventually), and there are distinct good libraries emerging for Mac only (sigh), but last I checked the experience of running every kit (stable diffusion, server-class, etc) involved overhead for the Mac world.
- Y-bar 2y ago> This is still half the speed of a consumer NVidia card, but the large amounts of memory is great, if you don't mind running things more slowly and with fewer libraries. But it has more than 2x longer battery life and a better keyboard than a GPU card ;)
- jsheard 2y agoIt's not the memory being unified that makes it fast, it's the combination of the memory bus being extremely wide and the memory being extremely close to the processor. It's the same principle that discrete GPUs or server CPUs with onboard HBM memory use to make their non-unified memory go ultra fast.
- smith7018 2y agoI thought “unified memory” was just a marketing term for the memory being extremely close to the processor?
- hollerith 2y agoI thought it meant that both the GPU and the CPU can access it. In most systems, GPU memory cannot be accessed by the CPU (without going through the GPU); and vice versa.
- layer8 2y agoCPUs access GPU memory via MMIO (though usually only a small portion), and GPUs can in principle access main memory via DMA. Meaning, both can share an address space and access each other’s memory. However, that wouldn’t be called Unified Memory, because it’s still mediated by an external bus (PCIe) and thus relatively slower.
- bobmcnamara 2y agoAre they cache coherent these days? I feel like any unified memories should be.
- jsheard 2y agoNo, unified memory usually means the CPU and GPU (and miscellaneous things like the NPU) all use the same physical pool of RAM and moving data between them is essentially zero-cost. That's in contrast to the usual PC setup where the CPU has its own pool of RAM, which is unified with the iGPU if it has one, but the discrete GPU has its own independent pool of VRAM and moving data between the two pools is a relatively slow operation. An RTX4090 or H100 has memory extremely close to the processor but I don't think you would call it unified memory.
- metadat 2y agoI was curious so I looked it up: https://en.wikipedia.org/wiki/DDR5_SDRAM https://en.wikipedia.org/wiki/DDR5_SDRAM (info from the first section): > DDR5 is capable of 8GT/s which translates to 64 GB/s (8 gigatransfers/second * 64-bit width / 8 bits/byte = 64 GB/s) of bandwidth per DIMM. So for example if you have a server with 16 DDR5 DIMMs (sticks) it equates to 1,024 GB/s of total bandwidth. DDR4 clocks in at 3.2GT/s and the fastest DDR3 at 2.1GT/s. DDR5 is an impressive jump. HBM is totally bonkers at 128GB/s per DIMM (HBM is the memory used in the top end Nvidia datacenter cards). Cheers.
- sroussey 2y agoYes, and wouldn’t it be bonkers if the M4 Max supported HBM on desktops?
- reliabilityguy 2y ago> So for example if you have a server with 16 DDR5 DIMMs (sticks) it equates to 1,024 GB/s of total bandwidth. Not quite as it depends on number of channels and not on the number of DIMMs. An extreme example: put all 16 DIMMs on single channel, you will get performance of a single channel.
- metadat 2y agoThanks for your reply. Are you up for updating the Wikipedia page?, because as of now the canonical reference is wrong.
- angoragoats 2y agoIf you're referring to the line you quoted, then no, it's not wrong. Each DIMM is perfectly capable of 64GiB/s, just as the article says. Where it might be confusing is that this article seems to only be concerning itself with the DIMM itself and not with the memory controller on the other end. As the other reply said, the actual bandwidth available also depends on the number of memory channels provided by the CPU, where each channel provides one DIMM worth of bandwidth. This means that in practice, consumer x86 CPUs have only 128GiB/s of DDR5 memory bandwidth available (regardless of the number of DIMM slots in the system), because the vast majority of them only offer two memory channels. Server CPUs can offer 4, 8, 12, or even more channels, but you can't just install 16 DIMMs and expect to get 1024GiB/s of bandwidth, unless you've verified that your CPU has 16 memory channels.
- Rohansi 2y agoApple is using LPDDR5 for M3. The bandwidth doesn't come from unified memory - it comes from using many channels. You could get the same bandwidth or more with normal DDR5 modules if you could use 8 or more channels, but in the PC space you don't usually see more than 2 or 4 channels (only common for servers). Unrelated but unified memory is a strange buzzword being used by Apple. Their memory is no different than other computers. In fact, every computer without a discrete GPU uses a unified memory model these days!
- manaskarekar 2y agoYes, it's just easier to call it that without having to sprinkle asterisks at each mention of it :) And yes, the impressive part is that this kind of bandwidth is hard to get on laptops. I suppose I should have been a bit more specific in my remark.
- binary132 2y agoI read all that marketing stuff and my brain just sees APU. I guess at some level, that’s just marketing stuff too, but it’s not a new idea.
- sroussey 2y agoEh… not quite. Maybe on an Instinct. Unified memory means the CPU and CPU means they can do zero copy to use the same memory buffer. Many integrated graphics segregate the memory into CPU owned and GPU owned, so that even if data is on the same DIMM, a copy still needs to be performed for one side to use what the other side already has. This means that the drivers, etc, all have to understand the unified memory model, etc. it’s not just hardware sharing DIMMs.
- binary132 2y agoI was under the impression PS4’s APU implemented unified memory, and it was even referred to by that name[1]. APUs with shared everything are not a new concept, they are actually older than programmable graphics coprocessors… https://www.heise.de/news/Gamescom-Playstation-4-bietet-Unified-Memory-Xbox-One-nicht-1939716.html https://www.heise.de/news/Gamescom-Playstation-4-bietet-Unif...
- mort96 2y agoThis is a case for on-package memory, not for unified memory... Laptops have had unified memory forever EDIT: wtf what's so bad about this comment that it deserves being downvoted so much
- willseth 2y agoIntel typically calls their iGPU architecture "shared memory"
- deleted 2y ago[deleted]
- mort96 2y agoHm it seems like they call it unified memory too, at least in some places, have a look at 5.7.1 "Unified Memory Architecture" in this document: https://www.intel.com/content/dam/develop/external/us/en/documents/the-compute-architecture-of-intel-processor-graphics-gen9-v1d0.pdf https://www.intel.com/content/dam/develop/external/us/en/doc... Intel processor graphics architecture has long pioneered sharing DRAM physical memory with the CPU. This unified memory architecture offers [...] It more or less seems like they use "unified memory" and "shared memory" interchangeably in that section
- Detrytus 2y agoI think "Unified" vs "shared" is just something Apple marketing department came up with. Calling something "shared" makes you think: "there's not enough of it, so it has to be shared". Calling something "unified" makes you think: "they are good engineers, they managed to unify two previously separate things, for my benefit".
- mort96 2y agoI don't think so? That PDF I linked is from 2015, way before Apple put focus on it through their M-series chips... And the Wikipedia article on "Glossary of computer graphics" has had an entry for unified memory since 2016: https://en.wikipedia.org/w/index.php?title=Glossary_of_computer_graphics&oldid=733057727#U https://en.wikipedia.org/w/index.php?title=Glossary_of_compu... For Apple to have come up with using the term "unified memory" to describe this kind of architecture, they would've needed to come up with it at least before 2016, meaning A9 chip or earlier. I have paid some attention to Apple's SoC launches through the years and can't recall them touting it as a feature in marketing materials before the M1. Do you have something which shows them using the term before 2016? To be clear, it wouldn't surprise me if it has been used by others before Intel did in 2015 as well, but it's a starting point: if Apple hasn't used the term before then, we know for sure that they didn't come up with it, while if Apple did use it to describe A9 or earlier, we'll have to go digging for older documents to determine whether Apple came up with it
- jedisct1 2y agoIs it GBps or Gbps?
- convexstrictly 2y agoGB per second
- deleted 2y ago[deleted]
- garciasn 2y agoWe run our LLM workloads on a M2 Ultra because of this. 2x the VRAM; one-time cost at $5350 was the same as, at the time, 1 month of 80GB VRAM GPU in GCP. Works well for us.
- manaskarekar 2y agoIf the 2x multiplier holds up, the Ultra update should bring it up to 1080GBps. Amazing.
- SirMaster 2y agoThere isn't even an M3 Ultra. Will there be an M4 Ultra?
- tromp 2y agoThat would make the most sense for the next Mac Studio version.
- hmottestad 2y agoAt some point there should be an upgrade to the M2 Ultra. It might be an M4 Ultra, it might be this year or next year. It might even be after the M5 comes out. Or it could be skipped in favour of the M5 Ultra. If anyone here knows they are definitely under NDA.
- Inviz 2y agoI have M3 Max with 128GB of ram, it's really liberating.
- sfn42 2y agoI have 32gb and I've never felt like I needed more.
- moffkalast 2y agoObviously you're not a golfer.
- andy_ppp 2y agohttps://www.youtube.com/watch?v=YzhKEHDR_rc https://www.youtube.com/watch?v=YzhKEHDR_rc :-) Thanks for that, I think I will watch The Big Lebowski tonight!
- moffkalast 2y agoFar out, man :P
- School-Cotton 2y agoHaving 128GB is really nice if you want to regularly run different full OSes as VMs simultaneously (and if those OSes might in turn have memory-intensive workloads running on them). Somewhat niche case, I know.
- shiroiushi 2y agoNo one needs more than 640kB.
- doctoboggan 2y agoThis is definitely tempting me to upgrade my M1 macbook pro. I think I have 400GB/s of memory bandwidth. I am wondering what the specific number "over half a terabyte" means.
- rsanek 2y ago540
- moffkalast 2y agoWell it's more like pick your poison, cause all options have caveats: - Apple: all the capacity and bandwidth, but no compute to utilize it - AMD/Nvidia: all the compute and bandwidth, but no capacity to load anything - DDR5: all the capacity, but no compute or bandwidth (cheap tho)
- Dibby053 2y agoWhy was this downvoted?
- moffkalast 2y agoTo quote an old meme, "They hated Jesus because he told them the truth."
- thimabi 2y agoAt least in the recent past, a hindrance was that MacOS limited how much of that unified memory could be assigned as VRAM. Those who wanted to exceed the limits had to tinker with kernel settings. I wonder if that has changed or is about to change as Apple pivots their devices to better serve AI workflows as well.
- segmondy 2y agoNeed more memory, 256GB will be nice. MistralLarge is 123B. Can't even give a quantized Llama405B a drive. LLM users rejoice. LLM power users, weep.
- jjcm 2y agoFor context, the 4090 has 1,008 GB/s of bandwidth.
- spacedcowboy 2y ago... but only 1/4 of the actual memory, right ? The M4-Max I just ordered comes with 128GB of RAM.
- culi 2y agoyou'd probably save money just paying for a VPS. And you wouldn't cook your personal laptop as fast. Not that people nowadays keep their electronics for long enough for that to matter :/
- losvedir 2y agoI'm curious about getting one of these to run LLM models locally, but I don't understand the cost benefit very well. Even 128GB can't run, like, a state of the art Claude 3.5 or GPT 4o model right? Conversely, even 16GB can (I think?) run a smaller, quantized Llama model. What's the sweet spot for running a capable model locally (and likely future local-scale models)?
- brandall10 2y agoYou'll be able to run 72B models w/ large context, lightly quantized with decent'ish performance, like 20-25 tok/sec. The best of the bunch are maybe 90% of a Claude 3.5. If you need to do some work offline, or for some reason the place you work blocks access to cloud providers, it's not a bad way to go, really. Note that if you're on battery, heavy LLM use can kill your battery in an hour.
- SkyMarshal 2y agoLots of discussion and testing of that over on https://www.reddit.com/r/LocalLLaMA/ https://www.reddit.com/r/LocalLLaMA/, worth following if you're not already.
- bufferoverflow 2y agoClaude 3.5 and GPT 4o are huge models. They don't run on consumer hardware.
- joeevans1000 2y agoI am always wondering if one shouldn't be doing the resource intensive LLM stuff in the cloud. I don't know enough to know the advantages of doing it locally.
- alexchantavy 2y agoCurious, what are others using local LLMs on a MBP for? Hobby?