5 ms·
What's the limitation that keeps memory limited to 96GB? Could one put 512GB of memory on a card? I'm curious about what is the limiting factor.
by 1024core 2y ago
What's the limitation that keeps memory limited to 96GB? Could one put 512GB of memory on a card? I'm curious about what is the limiting factor.
- jsheard 2y agoGDDR memory buses are so fast that the RAM chips have to be packed tightly around the GPU core to maintain signal integrity, so the limit is more or less how many chips they can physically fit multiplied by the biggest chip capacity their suppliers can provide.
- lazide 2y agoAlso limited by heat dissipation.
- codedokode 2y agoBut theoretically RAM chips do not need to be synchronous with each other. Even more, the data lanes on the chip do not need to be synchronous - you can treat each lane as an independent serial channel. And GDDR latency is high enough that longer lanes won't change anything.
- wtallis 2y agoYou can't quite treat each lane as an independent serial channel; DRAM chips are at least 8 bits wide, and this GPU needs to either connect 16 bits to each die or connect groups of 32 bits to two dies, and the DRAM die does want to get the whole word on the same clock cycle. There's no SERDES at either end like you get with PCIe or Ethernet. Just 512 PHYs doing 32+Gbps PAM3. If you want those to be long-reach PHYs, you're not going to have much die space or power budget left for compute.
- etiam 2y agoDoes that mean it's perfectly feasible to have more if one accepts a higher latency? Seems like there could be plenty of use cases where that's preferable.
- codedokode 2y agoGDDR chips already have very high latency.
- Lramseyer 2y agoNot exactly. The name of the game with GDDR memory is "speed on the cheap." To do this, it uses a parallel bus with data rates pushed to the max. Not much headroom for things that could compromise signal integrity like socketed parts, or even board traces longer than they absolutely need to be. That's why the DRAM modules are close to the GPU and they're always soldered down. Also, the latency with GDDR7 is pretty terrible. It uses PAM3 signaling with a cursed packet encoding scheme. At least they were nice enough to add in a static data scrambler this time around! The lack of RLL was kind of a pain in GDDR6.
- 0x073 2y agoLike the gtx 970 with 3,5 + 0,4 memory.
- GuuD 2y agoBandwidth/GPU real estate in terms of area. Biggest GDDR7 chips are 3GB 32-bit, it has 512bit wide bus. And even this is going to moonlight as a space heater
- codedokode 2y agoWhy Apple packs something like 64 Gb on the CPU chip and NVIDIA cannot?
- bryanlarsen 2y agoApple uses LPDDR4 which comes in densities of up to 16Gb AFAICT. So they can have ~5X as much memory using the same bus width.
- wmf 2y agoLPDDR dies can stack but GDDR cannot?
- pavlov 2y agoApple sells a Mac Studio with the M3 Ultra chip and 512GB VRAM (unified memory between CPU and GPU). It costs $9,500. Their secret is that the memory is manufactured within the chip package.
- codedokode 2y agoWhy NVIDIA cannot manufacture 512 Gb chips and put 16 of them on the board?
- jsheard 2y agoApples architecture comes with its own trade-offs, it gives them huge capacity and pretty good bandwidth, but not nearly as much as Nvidia's architectures have. The M3 Ultra is 800GB/sec, the RTX 5090 is 1.8TB/sec, and the H200 is 4.8TB/s(!). Huge capacity with middling bandwidth is in vogue because it's a good fit for AI inference, but AI training and most other applications of GPUs need as much bandwidth as they can get.
- deleted 2y ago[deleted]
- codedokode 2y agoWell, if you have 16 M3-equivalent chips you can multiply the bandwidth by 16, right? Also, as I understand, ML is basically matrix multiplication and it has O(N³) operations on O(N²) numbers, so bandwidth might be not as important as number of ALUs.
- ein0p 2y agoIt's actually not within the chip's package. It's soldered to the board. It's just regular, fairly high spec LPDDR5X IIRC, there are just a TON of memory channels.
- pavlov 2y agoIt's not in the package? TIL... My misunderstanding seems to be common across the interwebs. I remember Apple used to show slides depicting the M1 SoC as one unit containing a CPU, GPU, Neural Engine, cache, and DRAM all together. But slides shown at an Apple event definitely qualify for artistic license.
- YetAnotherNick 2y agoWhat's the usecase of 512 GB memory that you cannot achieve through multi GPUs? Maybe you can make it bit cheaper as you don't require multiple chip, but I would say it is just a maybe because the chip is not the costliest part for Nvidia to manufacture, it is the memory[1]. [1]: https://www.nextplatform.com/2024/02/27/he-who-can-pay-top-dollar-for-hbm-memory-controls-ai-training/ https://www.nextplatform.com/2024/02/27/he-who-can-pay-top-d...
- immibis 2y agoOne card with twice the memory would let you run the model in half as many cards, at half the speed.
- utf_8x 2y agoWhile there are definitely physical limits, the core limitation here is greed. They would sell less cards. Same reason why their consumer cards are limited to ridiculously low amounts of VRAM (16GB on an RTX5080, only 8 on the RTX4060, etc) so if you want to do any serious AI you have to buy their overpriced enterprise cards.
- therealpygon 2y agoNot just less cards; by making their cards have less ram, they are helping prop up cloud-based inference, which in turn generates revenue for their most expensive line. It is the reason you used to be able to get 12GB ram on a 3060, and that now you have to move up 3 increasingly expensive models to get the same, restricted only by drivers and not capability. They made it clear that this was all intentional because they didn’t want consumer hardware in data centers as it costs profits.
- blitzar 2y agoPlanned obsolescence, got to sell 7xxx cards somehow.
- octacat 2y agoWiring also, chips for ddr7 look like they have a lot of pads. And you need pads on the graphics chip for them too.