3 ms·
There was a time when video card memory was sort of an anti-segmentation thing. It used to be so unimportant that manufacturers would design cards with more RA
by hakfoo 2y ago
There was a time when video card memory was sort of an anti-segmentation thing. It used to be so unimportant that manufacturers would design cards with more RAM on them than the GPU could use effectively.
For example, there was no technical justification for the 4GB GT710-- any software that could use 4GB of video memory would crap out for the anemic core performance. But it was a great market differentiator-- we'll give you a larger number for the same basic price.
- paulmd 2y agoyep, this really wound to a close once the Titan series started. There were some models of 780 6GB - but there were never any models of 780 Ti 6GB, because that would have been basically a GTX Titan Black without the professional drivers. And then you had the actual workstation cards with double that much again (Quadro K6000 is 12GB). Another one people don't realize - the reason AMD put 8GB on the 390 series is because it was cheaper. The Hawaii die uses 512b, which is 16x GDDR modules... in 2014/2015 the price for 4 gigabit (=512 megabyte) modules was actually cheaper than the price for 2 gigabit (=256 megabyte) modules. Production naturally tends to roll over to the newer, higher-capacity modules over time, and the 2Gb modules were falling off the tail end of the production curve. The reason I'm bringing this up is that it's kind of illustrative of an underlying difference between then and now. Memory density was actually outstripping the actual needs at that point - and today GPU VRAM capacity is actually heavily limited by the density of the memory modules. By the time you hit Pascal in 2016, the GTX 1080 was using 1GB GDDR5X modules, and that was actually a very new tech at the time. Quadro P4000 and P6000 actually used clamshell because the 2GB modules weren't available yet. 2GB modules first really hit the market with the turing quadro series (RTX 6000 used a non-clamshell 12x2GB, RTX A8000 used a clamshell 2x12x2GB) but GDDR6X lagged behind again and didn't hit the market until midway through Ampere. The 3090 was a clamshell of 1GB modules (2x12x1GB GDDR6X) and the 3090 Ti moved to 2GB modules in a non-clamshell configuration. And there still isn't anything bigger than 2GB (=16 gbit) modules available today. There probably won't be until at least middle of next year (super refresh?). The density increases in DRAM started stalling out about the same time as everything else, and you can't double the module size and then clamshell anymore, so there simply is a lot less flexibility. Ironically the MCD idea was quite timely imo. For an Infinity Link (not infinity fabric!) that's smaller than a single GDDR6 PHY, they move two entire memory controllers and PHYs to an external die. That's part of the reason "AMD is less stingy with VRAM" in the internet discourse... they built the technology that disaggregates the memory config from the memory bus config. They paid some huge design penalties (in total area, idle power, data movement, physical package dimensions, higher BOM costs for that memory, etc) and then... just didn't bother to exploit any of the advantages it gave them, and immediately abandoned it in RDNA4. There is no reason AMD couldn't put 4 memory PHYs on a new MCD instead of 2. It's "disaggregated" for a reason beyond just yields, right? But AMD never bothered to make any other MCDs, so there is only the 2x PHY MCD die. Obviously you'd need probably some new packages especially for the big one - you could do similar "reverse 7900GRE" interposers that let you put a N32 into a N31 package while still driving all the memory, just wouldn't increase actual bandwidth, but the big die would need a bigger package to route twice the memory. But actually the larger physical size becomes an advantage at that point - bigger die is more shoreline. ctrl-f shoreline for some discussion on how that's affecting current designs: https://www.semianalysis.com/p/cxl-is-dead-in-the-ai-era https://www.semianalysis.com/p/cxl-is-dead-in-the-ai-era I personally think there is a similar advantage with PHY pinout, the PHY pins need to be on the edge and AMD simply has a bigger package with more edge area (without driving up costs too much). And the thing is, that bigger package area and higher idle power are also big downsides in other scenarios. AMD had big big hopes of making inroads into the laptop market with RDNA3/7900M, and they got almost zero uptake, because they were offering a bigger package with higher idle power, at a moment in time when laptop OEMs are actually looking at ditching dGPUs entirely to cram in more battery life (most are not against the FAA limits yet) so they can compete with Apple. AMD literally made exactly the diametrically opposite/wrong product for that market, and yet also failed to do any of the interesting high-VRAM configurations that disaggregation could have allowed too. Typical red-team moment right there - and this is the whole reason 7900GRE exists, to move the sorts of yields they would have shunted to 7900M. That's sort of the unfortunate reality, is that memory capacity has just slowed to a crawl, and the nature of die shrinks (PHY area doesn't shrink, which makes it preferable to shift from more PHYs to fewer+more cache instead, or use faster memory like 5X or 6X) has pushed actual bus width down over time to contain costs. And this is probably not going to change very quickly either - maybe we get 24gbit modules in 2025, but when is 32gbit coming, like 2026 or 2027? You can't just keep increasing bus width forever either, especially with wafer costs rising significantly every gen. And AMD actually had an answer to that, and just chose not to do anything with it, or use it ever again! The server market is a lot less afflicted because they've got CoWoS stacking and HBM. Like if you have 8-hi or 16-hi HBM3E then realistically you've got more than the gaming market will need for a long time... but it's too expensive to use on gaming cards for now (maybe if AI crashes it will lead to gaming HBM cards to soak up the capacity). Like yeah, it's kinda not just that NVIDIA doesn't care about the gaming market because AI is more profitable... it's that these are the solutions they can offer at this point in time, based on the tech that is available to them. And AMD had a workaround for that. -- anyway: yeah I was the sucker who bought the double VRAM cards, but I actually did get good usage out of it. ;) I had a Radeon 7850 2GB I got late in 2012 for $150 on black friday, and that was actually a great little card for 3 years or so. I think that's the appeal, that not only do you not ever have to worry about VRAM, but that it will have a useful second life as a downmarket card in your spare PC or kid's PC. I also did have a GT 640 4GB (GDDR5, at least!)... and that was pretty nifty back in 2014 for this new-fangled CUDA stuff! GK208 actually had all the capabilities of a Tesla K40 wedged into a $60 package. I do totally understand where people are coming from on the cost, I see it too... but right now the cost increases mean that if you're not willing to spend more, you're stepping down product tiers, which means things like smaller memory buses and lower total capacity etc. And actually it's when things slow down that "futureproofing" does become possible - just like the 2600K became a legendary cpu because cpu progress slowed to a crawl after that. Something like a 3090 24GB is both cheap and still doesn't really have any major downsides, other than 40-series having framegen (and AMD implemented non-DLSS framegen anyway) and 40-series generally being much better at path-tracing. Buy the higher card and keep it longer, it makes more sense than stepping down product tiers because you're upgrading every year. -- And my overall point with the first section is that people really tend to discount those technical factors when they analyze the market. It's not just that "NVIDIA is being stingy" or whatever. Sure they are reserving clamshell for workstation cards (because besides drivers, that's really the only segmentation they have left), but actually this is just what the tech can deliver right now without costs going mad. Turing and Ampere used older/cheaper nodes with much lower density - and they were massively large, inefficient dies by historical standards. GTX 1070 was a cut-down 300mm2 die just like RTX 4070... and people forget the accusations of "fake MSRP" at pascal launch too. NVIDIA tried to position FE as an above-MSRP premium card, and nobody else followed the MSRP either. It was a $449 card at launch, in 2016 dollars. Similarly GTX 670 was a $399 card at launch, for a cutdown ~300mm2 die. It's just that Turing and Ampere had obscenely large dies by historical standards. RTX 2060 had a 443 mm2 die or somesuch, nearly as big as GTX 1080! Both 2080 Ti and 3090 used a die that was about 70% larger than the 1080 Ti! People tend to make simplistic analyses about "the die isn't big enough [compared to the last 2 gens with atypically large low-density dies on old/bad/cheap nodes]" and whine the memory bus isn't big enough and the capacity isn't high enough... but there's also all these technical forces pushing the engineering in that direction, too. https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_processing_units https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces... I often wish that gamers and gamer-adjacent media had a little more of a sense of shou ga nai/shikata ga nai. Like if it doesn't make sense to upgrade then don't - I don't think anyone expects 1-gen upgrade cycles to be worthwhile anymore. It's just somewhat silly the amount of tears cried over things that ultimately can't be helped, and nobody's directing the hate at, say, Micron or Hynix or Samsung for not keeping up with their own cadences. This is just how this market is in this post-post-post moore's law era - even just the basic shrinks and basic improvements in other components are falling through badly, and obviously that has implications for the downstream products that want to use those technologies. Shikata ga nai. DDR5 isn't great either, 2DPC is basically almost not worthwhile anymore due to the performance hit and honestly RAM clocks are not increasing all that much either, unless you are running in 1DPC mode, at which point 6000 or 6400 is doable. And DDR6 is supposed to not be that great either. Things are just falling apart, the tech is not progressing very well anymore. Hitting the physical limits.