7 ms·
The M3 Max was 400GBps, this is 540GBps. Truly an outstanding case for unified memory. DDR5 doesn't come anywhere near.
by manaskarekar 2y ago
The M3 Max was 400GBps, this is 540GBps. Truly an outstanding case for unified memory. DDR5 doesn't come anywhere near.
- vid 2y agoIt's not "DDR5" on its own, it's a few factors. Bandwidth (GB/s) = (Data Rate (MT/s) * Channel Width (bits) * Number of Channels) / 8 / 1000 (8800 MT/s * 64 bits * 8 channels) / 8 / 1000 = 563.2 GB/s This is still half the speed of a consumer NVidia card, but the large amounts of memory is great, if you don't mind running things more slowly and with fewer libraries.
- cjbprime 2y agoRight, the nvidia card maxes out at 24GB.
- vid 2y agoA 24gb model is fast and ranks 3. A 70b model is slow and 8. A top tier hosted model is fast and 100. Past what specialized models can do, it's about a mixture/agentic approach and next level, nuclear power scale. Having a computer with lots of relatively fast RAM is not magic.
- manaskarekar 2y agoThanks, but just to put things into perspective, this calculation has counted 8 channels which is 4 DIMMs and that's mostly desktops (not dismissing desktops, just highlighting that it's a different beast). Most laptops will be 2 DIMMS (probably soldered).
- wtallis 2y agoDesktops are two channels of 64 bits, or with DDR5 now four (sub)channels of 32 bits; either way, mainstream desktop platforms have had a total bus width of 128 bits for decades. 8x64 bit channels is only available from server platforms. (Some high-end GPUs have used 512-bit bus widths, and Apple's Max level of processors, but those are with memory types where the individual channels are typically 16 bits.)
- sliken 2y agoI think you are confusing channels and dimms. The vast majority of any x86 laptop or desktops are 128 bits wide. Often 2x64 bit channels up till last year or so, now 4x32 bit DDR5 in the last year or so. There are some benefits to 4 channels over 2, but generally you are still limited by 128 bits unless you buy a Xeon, Epyc, or Threadripper (or Intel equiv) that are expensive, hot, and don't fit in SFFs or laptops. So basically the PC world is crazy behind the 256, 512, and 1024 bit wide memory busses apple has offered since the M1 arrived.
- wtallis 2y ago> (8800 MT/s * 64 bits * 8 channels) / 8 / 1000 = 563.2 GB/s Was this example intended to describe any particular device? Because I'm not aware of anything that operates at 8800 MT/s, especially not with 64-bit channels.
- sliken 2y agoM4 max in the MBP (today) and in the Studio at some later date.
- wtallis 2y agoThat seems unlikely given the mismatched memory speed (see the parent comment) and the fact that Apple uses LPDDR which is typically 16 bits per channel. 8800MT/s seems to be a number pulled out of thin air or bad arithmetic.
- sliken 2y agoHeh, ok, maybe slightly different. But apple spec claims 546GB/sec which works out to 512 bits (64 bytes) * 8533. I didn't think the point was 8533 vs 8800. I believe I saw somewhere that the actual chips used are LPDDR5X-8533. Effectively the parents formula describes the M4 max, give or take 5%.
- sliken 2y agoFewer libraries? Any that a normal LLM user would care about? Pytorch, ollama, and others seem to have the normal use cases covered. Whenever I hear about a new LLM seems like the next post is some mac user reporting the token/sec. Often about 5 tokens/sec for 70B models which seems reasonable for a single user.
- vid 2y agoIs there a normal LLM user yet? Most people would want their options to be as wide as possible. The big ones usually get covered (eventually), and there are distinct good libraries emerging for Mac only (sigh), but last I checked the experience of running every kit (stable diffusion, server-class, etc) involved overhead for the Mac world.
- Y-bar 2y ago> This is still half the speed of a consumer NVidia card, but the large amounts of memory is great, if you don't mind running things more slowly and with fewer libraries. But it has more than 2x longer battery life and a better keyboard than a GPU card ;)
- jsheard 2y agoIt's not the memory being unified that makes it fast, it's the combination of the memory bus being extremely wide and the memory being extremely close to the processor. It's the same principle that discrete GPUs or server CPUs with onboard HBM memory use to make their non-unified memory go ultra fast.
- smith7018 2y agoI thought “unified memory” was just a marketing term for the memory being extremely close to the processor?
- hollerith 2y agoI thought it meant that both the GPU and the CPU can access it. In most systems, GPU memory cannot be accessed by the CPU (without going through the GPU); and vice versa.
- layer8 2y agoCPUs access GPU memory via MMIO (though usually only a small portion), and GPUs can in principle access main memory via DMA. Meaning, both can share an address space and access each other’s memory. However, that wouldn’t be called Unified Memory, because it’s still mediated by an external bus (PCIe) and thus relatively slower.
- bobmcnamara 2y agoAre they cache coherent these days? I feel like any unified memories should be.
- jsheard 2y agoNo, unified memory usually means the CPU and GPU (and miscellaneous things like the NPU) all use the same physical pool of RAM and moving data between them is essentially zero-cost. That's in contrast to the usual PC setup where the CPU has its own pool of RAM, which is unified with the iGPU if it has one, but the discrete GPU has its own independent pool of VRAM and moving data between the two pools is a relatively slow operation. An RTX4090 or H100 has memory extremely close to the processor but I don't think you would call it unified memory.
- metadat 2y agoI was curious so I looked it up: https://en.wikipedia.org/wiki/DDR5_SDRAM https://en.wikipedia.org/wiki/DDR5_SDRAM (info from the first section): > DDR5 is capable of 8GT/s which translates to 64 GB/s (8 gigatransfers/second * 64-bit width / 8 bits/byte = 64 GB/s) of bandwidth per DIMM. So for example if you have a server with 16 DDR5 DIMMs (sticks) it equates to 1,024 GB/s of total bandwidth. DDR4 clocks in at 3.2GT/s and the fastest DDR3 at 2.1GT/s. DDR5 is an impressive jump. HBM is totally bonkers at 128GB/s per DIMM (HBM is the memory used in the top end Nvidia datacenter cards). Cheers.
- sroussey 2y agoYes, and wouldn’t it be bonkers if the M4 Max supported HBM on desktops?
- reliabilityguy 2y ago> So for example if you have a server with 16 DDR5 DIMMs (sticks) it equates to 1,024 GB/s of total bandwidth. Not quite as it depends on number of channels and not on the number of DIMMs. An extreme example: put all 16 DIMMs on single channel, you will get performance of a single channel.
- metadat 2y agoThanks for your reply. Are you up for updating the Wikipedia page?, because as of now the canonical reference is wrong.
- angoragoats 2y agoIf you're referring to the line you quoted, then no, it's not wrong. Each DIMM is perfectly capable of 64GiB/s, just as the article says. Where it might be confusing is that this article seems to only be concerning itself with the DIMM itself and not with the memory controller on the other end. As the other reply said, the actual bandwidth available also depends on the number of memory channels provided by the CPU, where each channel provides one DIMM worth of bandwidth. This means that in practice, consumer x86 CPUs have only 128GiB/s of DDR5 memory bandwidth available (regardless of the number of DIMM slots in the system), because the vast majority of them only offer two memory channels. Server CPUs can offer 4, 8, 12, or even more channels, but you can't just install 16 DIMMs and expect to get 1024GiB/s of bandwidth, unless you've verified that your CPU has 16 memory channels.
- Rohansi 2y agoApple is using LPDDR5 for M3. The bandwidth doesn't come from unified memory - it comes from using many channels. You could get the same bandwidth or more with normal DDR5 modules if you could use 8 or more channels, but in the PC space you don't usually see more than 2 or 4 channels (only common for servers). Unrelated but unified memory is a strange buzzword being used by Apple. Their memory is no different than other computers. In fact, every computer without a discrete GPU uses a unified memory model these days!
- manaskarekar 2y agoYes, it's just easier to call it that without having to sprinkle asterisks at each mention of it :) And yes, the impressive part is that this kind of bandwidth is hard to get on laptops. I suppose I should have been a bit more specific in my remark.
- binary132 2y agoI read all that marketing stuff and my brain just sees APU. I guess at some level, that’s just marketing stuff too, but it’s not a new idea.
- sroussey 2y agoEh… not quite. Maybe on an Instinct. Unified memory means the CPU and CPU means they can do zero copy to use the same memory buffer. Many integrated graphics segregate the memory into CPU owned and GPU owned, so that even if data is on the same DIMM, a copy still needs to be performed for one side to use what the other side already has. This means that the drivers, etc, all have to understand the unified memory model, etc. it’s not just hardware sharing DIMMs.
- binary132 2y agoI was under the impression PS4’s APU implemented unified memory, and it was even referred to by that name[1]. APUs with shared everything are not a new concept, they are actually older than programmable graphics coprocessors… https://www.heise.de/news/Gamescom-Playstation-4-bietet-Unified-Memory-Xbox-One-nicht-1939716.html https://www.heise.de/news/Gamescom-Playstation-4-bietet-Unif...
- mort96 2y agoThis is a case for on-package memory, not for unified memory... Laptops have had unified memory forever EDIT: wtf what's so bad about this comment that it deserves being downvoted so much
- willseth 2y agoIntel typically calls their iGPU architecture "shared memory"
- deleted 2y ago[deleted]
- mort96 2y agoHm it seems like they call it unified memory too, at least in some places, have a look at 5.7.1 "Unified Memory Architecture" in this document: https://www.intel.com/content/dam/develop/external/us/en/documents/the-compute-architecture-of-intel-processor-graphics-gen9-v1d0.pdf https://www.intel.com/content/dam/develop/external/us/en/doc... Intel processor graphics architecture has long pioneered sharing DRAM physical memory with the CPU. This unified memory architecture offers [...] It more or less seems like they use "unified memory" and "shared memory" interchangeably in that section
- Detrytus 2y agoI think "Unified" vs "shared" is just something Apple marketing department came up with. Calling something "shared" makes you think: "there's not enough of it, so it has to be shared". Calling something "unified" makes you think: "they are good engineers, they managed to unify two previously separate things, for my benefit".
- mort96 2y agoI don't think so? That PDF I linked is from 2015, way before Apple put focus on it through their M-series chips... And the Wikipedia article on "Glossary of computer graphics" has had an entry for unified memory since 2016: https://en.wikipedia.org/w/index.php?title=Glossary_of_computer_graphics&oldid=733057727#U https://en.wikipedia.org/w/index.php?title=Glossary_of_compu... For Apple to have come up with using the term "unified memory" to describe this kind of architecture, they would've needed to come up with it at least before 2016, meaning A9 chip or earlier. I have paid some attention to Apple's SoC launches through the years and can't recall them touting it as a feature in marketing materials before the M1. Do you have something which shows them using the term before 2016? To be clear, it wouldn't surprise me if it has been used by others before Intel did in 2015 as well, but it's a starting point: if Apple hasn't used the term before then, we know for sure that they didn't come up with it, while if Apple did use it to describe A9 or earlier, we'll have to go digging for older documents to determine whether Apple came up with it
- jedisct1 2y agoIs it GBps or Gbps?
- convexstrictly 2y agoGB per second