6 ms·
What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?
by fooker 12d ago
What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?
- refulgentis 12d agoIn one sense, nothing, in another, everything. It is DRAM, but the bandwidth requirements mean it’s paired to a processor, i.e. no DIMMs. Not 100% sure but things like MacBooks and the Framework tower, where you have fixed RAM for the device lifetime, have ~0 tradeoff.
- nutjob2 12d agoNothing except CPU manufacturer choices. Mac laptops use it and they're consumer products. People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.
- chessgecko 12d agopretty sure its lpddr not hbm.
- addaon 12d ago> Mac laptops use it and they're consumer products No, Mac laptops use LPDDR, currently LPDDR5X.
- nomorewords 12d agoThe normal non-tech-savvy person already does this. They simply don't know that their ram is upgradeable or something else breaks first, before having to touch ram.
- dylan604 12d agowhat is this too much RAM thing you mention? I thought the only valid RAM situation you could find yourself is not enough RAM. Too much? That's just fantasy
- KeplerBoy 12d agoMacBooks use regular soldered lpddr5(x) RAM. Same RAM as every other laptop manufacturer, they just use more lanes to achieve a higher bandwidth.
- bunderbunder 12d agoPerhaps more noteworthy for general home and business computing, doesn’t it also allow for lower latency?
- Rohansi 12d agoMore channels or soldered memory? Channels are basically RAID 0 so it depends what you're measuring. Soldering memory down was the only way to use LPDDR5X so if you wanted the best memory you had to solder it down. LPCAMM2 exists though so newer devices can use that instead of soldering them down, but not all devices would be able to fit the required LPCAMM2 slots.
- throwaway85825 12d agoSOCAMM2 allows for removable ram in nearly 0 added space.
- Rohansi 12d agoThat's tiny! But it still depends a lot on form factor. LPDDR5X is used in phones and making memory removable would have its compromises. You may even have compromises in laptops. Look at how tiny a MacBook Air's mainboard is and you'll see the RAM modules on the same package as the SoC. SOCAMM2 is too large for that but a variant with only two modules could possibly work.
- Melatonic 12d agoI bet it could fit. It just do a combo of soldered ram and non soldered like many laptops used to.
- fooker 12d agoApple's "unified memory" marketing is so strong that even tech literate people seem to have this misconception! They have managed to pull this sort of thing off many many times. https://en.wikipedia.org/wiki/Reality_distortion_field https://en.wikipedia.org/wiki/Reality_distortion_field
- Rohansi 12d agoYup, the only reason Macs have higher memory bandwidth is because they use more memory channels, which gives them a wider bus. Both Intel and AMD only allow more than dual channel memory on server class processors these days.
- kstrauser 12d agoThe “Apple only does X better because they do Y” thing has been a meme for ages. I remember dismissals like “PowerPC is only faster at math because it has more integer units” or something along those lines, and thinking, uh, isn’t that a good thing?
- JohnBooty 12d ago"You're only better than me at sports because you practice more and try harder!"
- pixl97 12d agoDepends on the expense trade off. If I get 10% more performance for 50% more cost it really depends on one's needs, for example.
- Rohansi 12d agoI am just saying it's not magic and x86 is capable of doing the same. Quad channel memory used to be more common in consumer hardware but now it looks like you don't even have the option anymore for desktops. AMD's Strix Halo was the first sign to reversing that (it has quad channel memory!) and hopefully we see more of that in the future.
- bob1029 12d agoIt's not that there's a blocker. It's that it takes roughly 3x the manufacturing capacity to produce an HBM package at the same storage capacity as DRAM. We are sacrificing total bytes for bandwidth.
- tliltocatl 12d agoHow so? FEOL is pretty much the same, BEOL is almost the same save TSVs, the packaging tech is different and more advanced, but not exactly 1:1 comparable. Do TSVs really occupy 3x the area of DDR IO's?
- monster_truck 12d ago3x is a reasonable figure. They are not literally that large, though.
- Const-me 12d agoSee remark on the slide 11: https://www.servethehome.com/micron-evolving-memory-architectures-for-ai-at-hot-chips-2026/ https://www.servethehome.com/micron-evolving-memory-architec... That presentation is by Micron.
- skavi 12d agodirect link: https://www.servethehome.com/micron-evolving-memory-architectures-for-ai-slide-11/ https://www.servethehome.com/micron-evolving-memory-architec...
- tliltocatl 12d agoYuck. Time to build an xSPI/HyperRAM workstation (if only these had multi-bank chips).
- buildbot 12d agoSadly the $ per byte of xSPI and HyperRAM quite high
- chessgecko 12d agoI think people might prefer the lower idle power consumption from lpddr over the better bandwidth in hbm in battery powered stuff. That said right now the price is definitely preventing us from finding out.
- fooker 12d agoIdle yes, but HBM energy consumption / memory operations seems to be a bit better than DRAM.
- vlovich123 12d agoOnly if you’re running at 100%. Consumers generally do not.
- monster_truck 12d agoThat hasn't been true since early HBM2 days, before the controllers standardized on power/voltage management and did things like leave them in P0 to ship on time
- vlovich123 12d agoLPDDR/DDR/GDDR generally still win over HBM when there’s no data being transferred. HBM is primarily better in watts/byte transferred. Consumer electronics spend most of their time idle.
- Zagitta 12d agoRacing to idle is a very common power optimization technique
- vlovich123 12d agoRight, but HBM idle is significantly worse than LPDDR idle or even DDR idle for that matter. That matters a lot precisely because the device is idle most of the time. Your idle power draw dominates.
- phkahler 12d ago>> What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics? HBM is meant to be integrated into the same package as the CPU, so no more DIMM sockets. It also has higher latency apparently.
- MarleTangible 12d agoOne of the comments mentioned that they have a bus size of 2048, which may be why the latency is higher. > These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5.
- reliabilityguy 12d agoHBM is a stack of DRAMs, so there is no “instead”.
- buckle8017 12d agoThe vias to enable stacking is a significant amount of the die area.
- saltcured 12d agoIf they're talking about production capacity, that is some product of die area and process steps, right? It doesn't have to be 3x die area, just 3x lower factory throughput for the same number of functioning memory bits.
- buckle8017 12d agoHBM is less dense at a water scale than DDR because of all the vias, but each memory but is as dense or denser. You make HBM instead of DDR and the number of bits you're making goes down. It's really that simple.
- HarHarVeryFunny 12d agoWhy would you want/need to? The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU. So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second! These are not normal requirements, other than for a GPU.
- fooker 12d agoSIMD (and especially the modern matrix extensions) can use as much bandwidth you can throw at it.
- b112 12d agoWithout further clarification, that statement seems impossible. "As much" being unbounded and all. You should expand what you mean.
- fooker 12d agoThis will blow your mind, but it actually is pretty close to being unbounded. :) Consider the 'MMA N matrices' primitive modern CPUs are starting to support. For the current generation of CPUs, N is a constant like 16 or 32, but there's nothing preventing it from being 1024 or larger if we have more memory bandwidth. All this with a single instruction.
- articulatepang 12d agoSurely something prevents it being 1 quadrillion bits per instruction? Since that’s well within “unbounded”.
- pixl97 12d agoMost likely the chip running at the core temperature of the sun. We'll have to figure out how to read and right to the surface of a black hole to get speeds that high.
- torginus 12d agoMainly bus width. Afaik HBM is like 1024 bits vs DDRs 64 so you need lots of transfers in parallel to saturate the bus, and CPUs kinda want 64 bytes of data as that's the size of a cache line ASAP. So you need a ton of in flight transfers which isn't a thing CPUs provide, maybe multicore workloads. Buy the way you win with CPUs is with latency, and not bandwidth, which is why Apple M series actually uses DDR with lower latency because of the stacking.
- tjwebbnorfolk 12d agoEven if the cost were the same of the RAM itself, you'd need much bigger and more expensive CPU to deal with it. Running 17 chrome tabs doesn't benefit at all from that HBM and all the additional hardware+software complexities that come with it. You want a specialized coprocessor to handle specialized workloads. The GPU exists separately from the CPU for a reason.
- KurSix 12d agoHBM isn't dramatically better in every dimension. You get huge bandwidth and good energy efficiency per bit transferred, but not necessarily a meaningful latency improvement and capacity expansion becomes tied to the package