13 ms·
The Unified Memory pool is what will continue to be the “game changer” in systems architecture, especially outside of data centers. The reality is even cutting
by stego-tech 4mo ago
The Unified Memory pool is what will continue to be the “game changer” in systems architecture, especially outside of data centers.
The reality is even cutting edge games and consumer workloads don’t actually take full use of the PCIe bandwidth of the GPU or the bandwidth of its GDDR memory. Even local AI use cases don’t substantially or meaningfully benefit from faster memory, at least to average consumers.
A unified memory pool does two things:
1) Lets systems optimize utilization based on need, rather than be confined to specific pools
2) Reduce overall memory cost, by letting system builders purchase a single type of memory in bulk instead of having to figure out GDDR vs DDR memory placement (important for SFF/portable machines)
So at a time when memory is expensive, unified pools make more sense. Even when memory becomes cheap and plentiful again, it’s just practical at this point to allocate a larger overall pool instead of managing discrete sets.
The one big drawback is security. A shared memory pool means side-channel attacks against memory from the GPU or CPU could potentially compromise the other as well, meaning memory-safe designs are going to be critical to security going forward (which is good for Rust adherents, I figure).
- Asmod4n 4mo agoyeah, you only see double digits in performance degradation from going from pcie 5 to 3 with a 5090 (at x16 speed), with everything else its like in the single digits area.
- stego-tech 4mo agoAnd the thing we gamers forget is that we’re the outlier. We’re the edge case. Most consumers will never really care about, let alone see, the difference in PCIe or memory bandwidth impacts from such a shift to unified memory pools. We might (being, at least in my case, a huge nerd), but I’m increasingly of the opinion that if modern blockbuster games are built for upscaling/reconstruction anyhow, then suddenly such sacrifices to performance seem acceptable relative to the gains in efficiency.
- jayd16 4mo agoWell I mean, the idea with games is it all fits in vram. You really don't want to be thrashing. It's that things are still so slow that they must be avoided entirely, no? No copy unified memory will help with that but you do pay the read speed costs.
- BoredPositron 4mo agogen3 is 16 years old.
- vlovich123 4mo ago> (which is good for Rust adherents, I figure). As a Rust adherent, please do not put words in our mouths or set up unrealistic expectations for other people by linking together concepts at a very shallow level. Language level memory safety has no answer for hardware security flaws which is what side channel attacks are. No programming language can provide memory privacy if another chip in your machine can read your memory. Just like no programming language can protect your application from a kernel vulnerability of the kernel it’s running on.
- stego-tech 4mo agoDamn. That wasn’t my intention at all, I was just pointing out that Rust has another reason to see wider adoption vis a vis the usual Valley advertising bullshit of deliberately conflating hardware security with software security. I personally give no fucks what something is written in, only that it’s written well enough that I don’t have to twist arms or babysit yet another sloppy piece of code in my enterprise.
- b112 4mo agoBut... it's rust.
- jmyeet 4mo agoUnified memory is only a feature because NVidia so aggressively uses VRAM for market segmentation. The 5090 ($2k MSRP but realistically $3-3.5k) is almost the same as the RTX 6000 Pro (~$10k). Same memory bandwidth (1800GB/s). Slightly different CUDA cores (21k vs 24k). Big difference? VRAM (32GB vs 96GB). NVidia ultimately doesn't want to upset this segmentation so the RTX Spark will never undermine their other offerings. This is why I think Apple has a real market opportunity if they choose to embrace it.
- zozbot234 4mo agoEven low-VRAM cards are actually very useful for running the comparatively smaller dense layers in large local MoE models. This only requires transfering very small amounts of data across the PCIe bus (similar to pipeline parallelism) so it fits nicely around the existing bottlenecks on that hardware.
- woodson 4mo ago> 5090 ($2k MSRP but realistically $3-3.5k) These days, more like >$4.1K (at least in the US).
- dahart 4mo agoI have so many questions… Since Apple already sells unified memory systems, what is the market opportunity you envision? Do you see Nvidia and Apple as competitors, and how? (And I’m not suggesting they’re not, necessarily, but I want to hear where you’re coming from, and they do have very different markets.) Hasn’t Apple used storage size (RAM & disk) for market segmentation for decades? And how does a machine with 128GB unified mem not potentially cut into some people’s reasons for wanting a 96GB GPU?
- JohnBooty 4mo agoI'm not the person you're replying to, but I wholeheartedly agree with them... Quick background: doing AI inference requires three things. Lots of memory, lots of memory bandwidth, and of course plenty of compute that has access to that memory. Quick reference: nVidia 5090 has 1,792 GB/sec bandwidth. 3090 gets about 1000 GB/sec. DGX Spark and AMD 395 whatever get about 275 GB/sec. Apple M1 Max gets 400GB/sec, M5 Max gets 614GB/sec. Ultra variants get 2x that bandwidth, base variants get 1/2 that bandwidth. However... their compute is rather weak. Right now, Apple's offerings are juuuuuust fast enough to run dense 27B models at usable speeds at like, 10% of the performance/watt of nVidia. They're world-leading general purpose CPUs but not killer GPUs. By all accounts, these Windows PCs nVidia is touting seem to have DGX Spark like performance, which is less than impressive. Same with the upcoming AMD AI-oriented consumer stuff. The other context here is that running your own AI at home is just starting to become feasible in terms of open model availability and the ability to run it at usable speeds. Many are interested in it for reasons of privacy, security, and cost certainty vs. buying tokens. Since Apple already sells unified memory systems, what is the market opportunity you envision? nVidia and AMD can't make their consumer offerings too good at AI, because that risks interfering with their higher-margin data center sales. (And, let's face it. Even if nVidia did release a 6090 with 64-128GB of memory for an affordable price, consumers wouldn't get their hands on them anyway because people would just start filling data centers with them) So. Now you see Apple's opportunity, right? No data center sales to interfere with. No relationship with nVidia or AMD to worry about. They could choose to make an absolute beast of a home AI machine. The M5 Ultra, if announced, might be that. It's admittedly a niche market, but people are already buying 64GB+ Macs faster than Apple can make them and they're fetching high prices on the used market as well. The only real questions are if this market is even something Apple would find time to care about, and if they could secure enough DRAM to make a go at it. They are enormous obviously but they're feeling the RAM pinch just like everybody.
- Retr0id 4mo agoMemory safety is orthogonal to side-channels, and hardware-enforced isolation (e.g. IOMMU) is more powerful than compiler-enforced isolation (but both are good!)
- bobmcnamara 4mo agoOh no now I have to worry about shaders row hammering my OS ram /s
- Retr0id 4mo agoYou really do have to worry about that!
- RiverCrochet 4mo agoIsn't this how the Xbox 360 got hacked? Not necessarily rowhammer but other methods.. IIRC some shader code in King Kong was able to affect CPU execution or something like that.
- supertroop 4mo agoIntel was doing UMA with their i740 graphics in the late 90s. Codename TIMNA was cancelled, but they pioneered it and used it on their you/cpu chips as well as their breakthrough 810 chipset that dominated graphics market for a decade. It was despised because it wa ubiquitous and a low performing graphics engine but games had to accommodate it. Funny that it is getting credit only now.
- p_l 4mo agoSGI O2 was the famous "unified memory architecture" graphics system, two years before i740 that didn't really do UMA. O2 was popular in systems where large textures or textures generated dynamically (like mapping external video input to texture) was important
- seemaze 4mo agoAnd here I am with 128GB Strix Halo longingly eyeing the Blackwell cards that spit tokens 10-20x the speed. The question is ultimate shape of knowledge compression and bandwidth optimization at which we arrive I suppose.
- canyp 4mo agoIf you haven't already, check/increase the GPU memory carve-out on your UEFI. More details: https://rocm.docs.amd.com/en/docs-7.2.0/how-to/system-optimization/strixhalo.html#memory-settings https://rocm.docs.amd.com/en/docs-7.2.0/how-to/system-optimi...
- electroglyph 4mo agothat link actually recommends not doing it from UEFI and doing it via software
- seemaze 4mo agoCurrently utilizing 126GB GTT on a headless host
- aabdi 4mo agoIf this thing only has as much gpu bandwidth as the spark, it’s kinda pointles
- cthalupa 4mo agoNot true. This is aimed squarely at the Strix Halo and Mac markets. It's basically just strictly better than the Strix, and it's not clear cut vs that Macs in any sort of blanket statement. My M5 Max 128gb MBP decodes faster than one of my Sparks, but the Spark's prefill is so much faster it can often answer the same query before the mac's prefill is finished. If you have large prompts, low cacheability, etc., a spark might be a very good options. Not to mention you get can get two sparks and the MBP will be 85%+ of the cost at half the RAM. I'm kind of tempted to pick one up. Leave running big models to my dual dgx setup, and all the misc. random stuff on an rtx.
- zozbot234 4mo agoPrefill will be a huge deal if batched unattended inference of SOTA models (on consumer platforms) becomes viable, because at that point it's the main remaining bottleneck. If running 30 inferences together boosts your decode throughput to 3x (that's consistent with some very rough experiments, though these haven't even looked at trying to mask SSD offload latency just yet), that's a 10x in total decode time but a 30x in total prefill time, because prefill workloads are fully compute bound already on consumer platforms and don't benefit from batching much at all.
- aabdi 4mo agoFair, but I don’t see what case you have w this. Mind sharing? Seems niche to be both uncacheable and long context?
- cthalupa 4mo agoAnything where you're dealing with a large volume of records/documents. Lots of people are using these for large-scale digitization of documents - scanned stuff being OCR'ed and summarized, generating embeddings, etc. Large scale translation. Anywhere where you might have a large backlog of data to work with can end up in this sort of situation.
- pbalau 4mo agoWhat is the difference between unified memory and shared memory? Shared memory existed since the first CPU with an embedded GPU came to market and you could set in BIOS how much memory goes to what component. I do have an opinion about how unified memory could be different, but I want a proper explanation.
- Gareth321 4mo agoSystem RAM has much lower bandwidth and less predictable access. Notably, the transfer from system to GPU is very slow. About 30x slower. LLMs aren’t designed to queue or parallelise operations to account for this. They just become much slower.
- ImprobableTruth 4mo agoShared memory of the past meant reserving a part of the memory for the GPU, which could then not be used or accessed by the CPU. If the CPU wanted to access something, it had to copy it from the GPU's section of the memory to its own. Unified memory means both just fully share the same memory.
- saltcured 4mo agoI'm not sure everyone uses the terms consistently, but the difference is that the old "shared" memory was reserving a section to act as VRAM under the control of the GPU, ignored by the OS. The CPU ran the same kind of code pretending there is a "bus transfer" between host memory and graphics memory. In unified memory, all the memory is host memory and data can go from program to GPU with zero copy movements. The addresses of buffers can be shared via appropriate MMU translation support, so that the application and graphics subsystem are communicating effectively through the basic RAM cache coherency protocols over the same buffers. Edit to add: Aside from the zero copy transfer potential, it also means dynamic allocation strategies can shift the balance between host and graphics allocations on the fly. Individual image and message buffers can be allocated on the fly instead of setting a static split between the two worlds.
- pbalau 4mo agoThat's my understanding, or, maybe a better word would be "guess". The CPU telling the GPU: this is your memory now.
- Salgat 4mo agoThat was the main reason for the big hype around Memristors 15 years ago. High density, high speed persistent memory to completely remove the need for hdd/ssds, potentially even removing the need for external memory altogether. So frustrating that it still seems like we're a long ways from that becoming reality. There's some renewed interest in Memristors as they can simulate neural network connections in models, so maybe the funding will return for it.
- zozbot234 4mo agoThe one example of persistent memory that managed to reach the mass market was Intel Optane/3dXPoint (still popular today among people looking to save on RAM costs) and that used a kind of phase-change memory, which is but tangentially related to memristors. ReRAM is somewhat closer, but it's also been less successful so far.
- ForOldHack 4mo agoWell, back in the day... The MacIIfx had video memory, ( dual ported ram ) that could be read and written to out of different ports. Wicked fast. It 486DX2s more than a year to catch up.
- Melatonic 4mo agoOptane was still much slower than Ram. And not that much faster than NVME (theoretically)
- testing22321 4mo ago> The Unified Memory pool is the “game changer” M1 knocking from 2020. Gamed changed, past tense, six years ago. This is catch-up.
- bombcar 4mo agoI want unified but not uniform - everything can address anything, but you can add slower RAM to the system without requiring an entirely new chip. NUMA is cool.
- jandrese 4mo agoHell, SGI O2s from 1996 had this. For all of the hype the performance gains were pretty modest.
- wmf 4mo agoUMA was never about performance and it still isn't. Spark is slower than a 5090.
- JMiao 4mo agodid they learn why? were there other gains?
- p_l 4mo agoO2 GPU was slower than other SGI options at the time, however it could use hilariously larger pool of memory without copying, which meant that O2 could use approaches that were punishingly hard (very tight transfer loops) or impossible (huge textures that couldn't be virtualized due to needing whole texture). That was because unlike other GPUs at the time, O2's didn't have dedicated memory but shared the memory with CPU - way slower, but zero copies and bigger. Arguably early home computers and workstations also used "unified memory" :D
- zdw 4mo agoFWIW, the O2's UMA let it handle far more textures than almost any other contemporary system with reasonable performance. Most other SGIs had single or low double-digit megabytes of texture memory, whereas the O2 could host one gigabyte of unified memory and use a huge chunk of that for textures.
- nalekberov 4mo agoIt’s also the reason, why you will never be able to repair or upgrade your computer in the future. From technological point of view these are indeed big advancements. However, I couldn’t care less about faster CPU when: 1. It limits my ability to upgrade my system 2. Windows gets increasingly bloated and slower
- merb 4mo agoLPCAMM2
- cm2187 4mo agoAnd conveniently, by making your machine non upgradeable, it allows the manufacturer to enforce market segmentation / charge a huge premium for small RAM upgrade (a la Apple)
- MBCook 4mo agoIs that required or just a choice Apple made?
- cm2187 4mo agoWhat do you mean by required? Apple's prices are notoriously disconnected from the cost of manufacturing.
- MBCook 4mo agoI mean is it possible to make unified memory systems with good performance or is it not really feasible due to memory timing/trace length issues? It’s possible if you’re willing to go with much slower RAM than GPUs like but CPUs often use. Thats what integrated graphics laptops have done for a long time right? But can you get high end CPU and GPU performance with unified memory and maintain user upgradable memory in a reasonable way? Thats what I don’t know.
- wtallis 4mo ago> I mean is it possible to make unified memory systems with good performance or is it not really feasible due to memory timing/trace length issues? LPCAMM and similar solutions exist, but have never been demonstrated running at speeds that match what the leading soldered memory systems are using; there's always been some speed penalty. I'm not sure we've ever seen a system demonstrated using LPCAMM or similar for a 512-bit bus to match Apple's Max tier SoCs, so it's somewhat of an open question whether those solutions can offer upgradability at the high end of the market for unified memory systems.
- 4mo ago
- tjoff 4mo agoDon't really buy the economic argument. For 99% pf all workloads you need at least an order of magnitude more system memory than gpu memory. Most systems barely need more gpu memory than what is required for video, browsing etc. Just because we found a new usecase doesn't flip that on its head. Besides, I want to keep doing what I'm doing today. So if I need 128GB today and my local AI needs 128 GB then I'd need 256 GB to keep doing the same work. The argument rather seems to be that we shouldn't use such expensive memory on the GPU. Which might be true if you only want to do inference on it.
- Joel_Mckay 4mo agoJensen Huang has publicly stated he wants a future where "AI" agents use more PC computers than people. It is ambitious, and absurd... like all CEOs that eventually go loopy. =3
- david-gpu 4mo agoDRAM optimized for CPU usage looks very different from DRAM optimized for GPU usage. You are leaving a lot performance on the table when you have a unified memory architecture. It makes sense in some situations, but it is not a silver bullet.
- Izikiel43 4mo ago> The Unified Memory pool is what will continue to be the “game changer” in systems architecture, especially outside of data centers. The ps4 was the prime example of this, and how it could run so many great games.
- NikolaNovak 4mo agoThe "one big drawback" is the lack of consumer upgrades, and the seemingly arbitrary prices charged by vendors for memory upgrades at time of system purchase. I'm not saying it has to be that way, but seems like it has been so far :-(
- maccard 4mo ago> The reality is even cutting edge games and consumer workloads don’t actually take full use of the PCIe bandwidth of the GPU or the bandwidth of its GDDR memory Game dev here. For anyone reading this - it’s not because we’re lazy, it’s because _it’s really hard to do_. One of the biggest differences between the current generation consoles and the current gen PCs is unified memory.
- BuyMyBitcoins 4mo agoHow much of that difficulty comes from the chosen game engine? I assume the engine is the primary factor in how resources are allocated.
- keyringlight 4mo agoOne related question that you need to follow that with is the associated costs of switching the whole studio to another engine that's technically better, or if proposing teach studio tailor-make their own engine the costs of that engineering, if presumably they have or learn the expertise to surpass whatever they're using currently. I'm not a game developer, but it would also seem to be a link between resource usage by the engine, and whatever content the production side are making. For all the commentary about how brilliant the id software engines are, if you examine the levels you pass through they're also very efficient with what they demand out of the engine - it's like an orchestra playing well together, not one instrument that means you can do anything.
- maccard 4mo agoBoth lots and none at the same time. The engines definitely make decisions for you but with unreal (for example) you can modify the RDG any way you see fit. The problem is that when you need something in gpu you have to go through RAM first (unless you have DMA which is a more recent addition). That doesn’t just add latency it also adds an extra step of cache invalidation, so you have to plan for that from the highest level of gameplay. If you need to prepare for a GPU memory miss _and_ a CPU memory miss as a worst case all the time, it’s very hard to make good use of the bandwidth in the best case
- 4mo ago
- up2isomorphism 4mo agoThis kind of post shows you have little idea why cpu and gpu are not sharing memory in the first place.
- GTP 4mo agoWhile I'm a supporter of Rust, I have to point out that Rust's memory safety doesn't help against side-channel attacks.
- jorvi 4mo agoYeah, no. GDDR is functionally very different than SDRAM. GDDR tries to push out as much bandwidth as possible, because that really matters for (traditional) GPU workloads. A constant but insignificant (= correctable) error rate is considered completely fine for GDDR, because that sacrifice allows the memory to be pushed much farther. Meanwhile most (traditional) SDRAM workloads don't give a hoot about bandwidth but really care about latency. And ideally you want no errors, hence ECC RAM being so venerated. If you unify memory, you're gonna have to choose to sacrifice one of those workloads or go suboptimal for both. Weirdly enough this mostly matters for non-gaming workloads. The Apple M-series are absolute monsters in gaming, completely crushing the RTX XX90 editions in performance-per-watt, but as soon as memory bandwidth becomes paramount the M-series falls heavily behind.
- jayd16 4mo ago>[..] take full use of the PCIe bandwidth of the GPU or the bandwidth of its GDDR memory. I'm honestly a little confused by what you mean here. Why would we want to maximize those things? Games are about consistent output under the frame deadline, not full saturation of the hardware. Why would anyone try to saturate a 5090 with their game? The addressable market is tiny and you'd have to hope their full spec runs as well as or better than your test rig or they'll still not hit framerate.
- simonbw 4mo agoYou could do some sort of adaptive quality where you spend time incrementally improving fidelity until your frame budget is up. In practice I think that might be trickier than it sounds, but I feel like theoretically there's something there that could get you the best graphics your rig can handle without dropping frames. I've been considering doing something like this when I've been building a game/engine lately.
- Rohansi 4mo agoThere's only so high you can go because the game assets have a maximum quality. Maybe you'll be able to max out the 5090 but what about the next flagship GPU? You're also likely not going to maximize all of bandwidth, compute, etc. because one of them will likely be your bottleneck. And it might be different depending on the GPU, too.
- rustystump 4mo agoMost games are strictly scaled on resolution due to how deferred pipelines run. This is exactly the slider to max or not max everything on a gpu for games. The more pixels the more memory and the more compute.
- Rohansi 4mo agoIf you're rendering at native resolution, which many PC gamers do, going higher isn't significantly better because it just helps with antialiasing via supersampling. There's no point rendering so much more pixels just because you can, that's just a waste of electricity.
- AnthonyMouse 4mo ago> Lets systems optimize utilization based on need, rather than be confined to specific pools The trouble with this is that the different types of memory have different characteristics. Latency for ordinary system memory is actually better than it is for GDDR, because GDDR is optimized for bandwidth. RTX 5090 has 1.8TB/s of memory bandwidth with a 512-bit memory bus. The same bus width for DDR5-9600 would have better latency but only a third of the bandwidth. CPU workloads are generally bounded by latency and GPU workloads are generally bounded by bandwidth, which is why they use two different types. > Reduce overall memory cost, by letting system builders purchase a single type of memory in bulk instead of having to figure out GDDR vs DDR memory placement (important for SFF/portable machines) The trouble with this is cost. In principle you could get the same 1.8TB/s of memory bandwidth as the RTX 5090 has, with the better latency of DDR5, by using DDR5 with a 1536-bit bus. This is indeed with multi-socket servers do, two sockets with 768-bit in memory channels per socket, but now check how much those system boards cost. But the remaining alternatives are both worse. If you use GDDR for the unified memory then GDDR costs more than DDR and you're going to have significantly worse latency for the CPU. If you use DDR without a 3-4 times wider bus than the already-wide GPU then the GPU gets starved for bandwidth.
- fc417fc802 4mo agoThese are all good points that I agree with but rather than seeing an intractable problem I predict we'll see the role that GDDR would otherwise fill in this scenario replaced by a small block of HBM on the APU die. I don't know if it will ultimately end up unified or not but either way I don't think memory segmentation is the core problem here. Simply not needing to send transfers across the narrow and slow PCIe bus would fix most of the practical problems (at least AFAIK but I'm not an expert). Transitioning over to wild speculation here, I think that most likely this will be treated as part of an absurdly large L3 (ala 3D V-Cache) or as an additional L4. In either case I expect the latency and power tradeoffs introduced to be tolerated as "good enough" even for the highest end consumer gear. (Actually I wonder if some sort of special case cache would be feasible, with memory addresses flagged by the graphics driver and regular CPU related stuff skipping over it entirely. But by then we've squarely entered the territory of vaguely unhinged rambling on my part.) Alternatively if the performance caveats are deemed to be important enough to justify the added complexity it wouldn't surprise me to see the HBM treated as an independent memory pool analogous to that of a dGPU. That wouldn't change the current status quo with respect to the GPU APIs but it would significantly ameliorate the memory bandwidth bottleneck for inference workloads and from a software perspective is a drop in replacement. You'd still write the code targeting the dGPU with explicit swapping to RAM but when run on an appropriate APU it would get a massive speedup for free instead of suddenly being starved for bandwidth while also performing unnecessary copy operations.
- wren6991 4mo ago> Even local AI use cases don’t substantially or meaningfully benefit from faster memory, at least to average consumers. I'm not sure what you mean by this. Memory bandwidth is the main bottleneck for single-user decode. The bottleneck is actually more severe for end-user inference than cloud inference, because end users don't have the option to increase arithmetic intensity by computing tokens for multiple clients in the same pass. One thing we've learned from Apple is the viability of spamming more LPDDR5X channels (up to 1024-bit total bus width on M3U) as a means of achieving high bandwidth while keeping the cost/capacity reasonable.
- pjjpo 4mo agoIsn't the big drawback not having a swappable GPU? Perhaps that's not as important anymore but I'm not sure we've confirmed the market demand for that.
- Lplololopo 4mo ago" Even local AI use cases don’t substantially or meaningfully benefit from faster memory, at least to average consumers." What do you mean by this? Memory bandwidth is fundamental to the speed of an local AI model