3 ms·
Yeah, I agree fully on the "pandemic brain" thing. One interesting observation I made is that usually mainstream midrange cards seem to hover just over 200mm².
by ColonelPhantom 3y ago
Yeah, I agree fully on the "pandemic brain" thing. One interesting observation I made is that usually mainstream midrange cards seem to hover just over 200mm². That includes for example cards like the HD 7850, but extends all the way to the likes of the RX 580, RX 5700 XT, but also the GTX 1060 or GTX 660. Those cards are all 256 bit if AMD and 192 bit otherwise, funnily enough.
It seems die sizes are on the rise though! The 6600 XT is a similar size to the 5700 XT despite having much less hardware (except cache), although I guess RDNA2's better clocks also ate some die space. The RTX 3060 is also reasonable large at 276 mm², which seems more like what's usually 60 Ti territory. And the 3070 is even bigger than the 1080!
Even in that context though, the 4060 Ti looks fairly anemic at 188 mm². The 4060 is even smaller at 159 mm², which almost equals the RX 5500 XT that debuted at $170, just over half the cost of the 4060.
Intel's current GPUs are meanwhile comically large seeing as the 400 mm² or so ACM-G10 competes with N23/N33 and GA106 which are much smaller, and even the tiny 4060. I think that's mostly because of the way too fine grained EUs which also makes the architecture very complex with slices and subslices. But at least for Battlemage they seem to be planning to grow the EUs, rather than throwing on more.
It's also worth noting that if Nvidia decided to "double up" the memory on the 3070 in a similar way to the 4060 Ti, they could have made a 32 GB version of it. Intel could also decide to do something like that for Battlemage, or of course AMD could do it with the 7900 XTX or similar. Despite AMDs weaker AI performance (no real matrix multiplication acceleration, unlike Nvidia or Intel) it would at least get AI developers to likely care about their cards as it means they have the best flagship in one important metric.
- paulmd 3y agoWith the way wafer costs keep increasing, physically smaller dies are inevitable. 50% higher density at 30% higher cost or whatever implies that if you keep the same die size, then cost goes up 30%, and in market terms that die has moved up a product tier. That number I was referring to is total transistor count and transistors-per-$ not die area, the cost of producing the same sized (eg 200mm2) die across nodes has skyrocketed, a 4060 Ti is far far more expensive than a 1060 to produce. Clamshell also adds PCB/assembly cost which people never account for. And all of this has to have partner margins rolled into it too etc. Eg at one point it was like $3.50 per 8gbit module, assuming 16gbit modules are comparable then 8GB x $4 = $32 of memory, but probably there's another $10-15 in the PCB and assembly, plus partner margin, etc. $50 would be partners losing margin, $75 probably they break even, $100 definitely is pushing up margin a bit, but that's also what people wanted, if you remember back to the bandwagon around EVGA. I think in the rush to bandwagon everyone just forgot who ultimately pays that margin ;) (and of course in launch terms, once cards are manufactured it's painful to open them up and remanufacture them all, are you going to desolder the BGA to put it on a new PCB, one by one?) I've commented the same thing before, that 3060 Ti 32GB would also be a very appealing product at the moment. I guess they have their reasons but it seems like a slam-dunk cashgrab at this moment of AI mania? RDNA3 does have WMMA which should at least get the foot in the door for things like AI/ML(-weighted TAAU) upscaling. I expect them to take a crack at that with FSR4, I strongly discount rumors that they're not working on that, especially with the console refresh (PS5 Pro at least) having RDNA3 and some kind of AI core. Even if sony builds one in-house, AMD totally have to be working on one internally too, just there's various reasons not to talk about it externally etc (people wanted them to not announce overly early etc, this is what that looks like). They'd want something for devs who want to validate once across multiple platforms anyway, they definitely will launch something within a year if not alongside the PS5P itself. WMMA may not be the answer for pure training horsepower though, NVIDIA having full-fat tensor units (like CDNA's) in their gaming GPUs was a good move (I won't even say "lucky", they've worked for it). But RDNA3 WMMA is better than nothing. I really think Intel's architecture is aimed at a future-node (plus an attempt to get the foot in the door on game/driver support). wave-8 is indeed way too finely grained for most workloads right now, but the pendulum is swinging back from cache to logic, and that means you have more transistors to spend on such frivolities. Growing the EUs definitely makes sense to me, like they were proofing the control/scheduling side at a small scale first, now they build the rest of it, etc. It just feels obviously wrong to do wave-8 right of the gate, on a 3060/3070 tier product at best, so what was the real goal? Some kind of incremental/iterative development, presumably. But yeah intel is a can of worms. The last time I checked (probably 6mo ago) they were operating the graphics division at -200% margin, which is needless to say stunning. I just also think they really have no choice, the writing is on the wall with NVIDIA Grace and MI300X and Apple Silicon Max/Ultra that big APUs are the future, both for a number of consumer segments but especially for a lot of HPC and even enterprise segments. Can't be competitive in HPC or enterprise if you don't have it, and even for laptops, you can't not have an iGPU. So if you cancel it do you go to Imagination or someone and license PowerVR like the bad old days of Atom? That's not a compelling product when AMD is turning out chips like 7840HS and soon the Strix Point/Strix Halo line. Other people tend to read that as "-200% margin, they'll cancel it any day now" but I actually read it as the opposite, it's "-200% margin and it's so important they are barreling through anyway". They are canceling stuff left and right but they don't have a choice on Xe/Arc regardless of the cost. And specific dies being canceled based on the current stuff doesn't mean anything in that context - that's just staying agile rather than over-committing to a schedule before you know what product you're going to build. But in general Intel can't seem to execute much of anything very well these days. Every good release feels like a fluke and they're back to the woodshed in no time. Alder Lake came with the loss of AVX-512, the 2.5gbe NICs are stuck in a state of permanent re-spins and steppings, Sapphire Rapids has some fatal power flaw that needs a 1200W psu (and they're not kidding) due to silly 700W+ transients and is getting a respin just in time for emerald rapids, meteor lake was basically 6mo late and still had a messed-up BIOS and still undershot performance targets, on and on, and they're cutting pay and firing staff. Death-spiral territory right there.
- ColonelPhantom 3y ago> RDNA3 does have WMMA which should at least get the foot in the door for things like AI/ML(-weighted TAAU) upscaling. Do keep in mind that RDNA3 WMMA is very slow, running at the same theoretical TFLOPS as shader (which are doubled from RDNA2 due to very limited dual issue support that can be used in WMMA). Nvidia tensor cores and Intel XMX can run closer to 4:1 or 8:1 or so compared to vector workloads. > It just feels obviously wrong to do wave-8 right of the gate That's because Alchemist isn't a a true first generation product, it's a scaled-up version of Intel's (relatively mature at this point) integrated graphics product. This means that Alchemist has suffered large growing pains (Gen12 was not really designed to be used in products bigger than maybe 128 EUs, A770 is 512 EUs). It also has some form of separation anxiety, seeing its need for ReBAR. I also recall very early in Alchemist's life (pre-release) some driver optimization that had a huge benefit, which was just the wrong memory region being used, as all memory is the same on an iGPU but not a dGPU.
- paulmd 3y ago> It also has some form of separation anxiety, seeing its need for ReBAR. I also recall very early in Alchemist's life (pre-release) some driver optimization that had a huge benefit, which was just the wrong memory region being used, as all memory is the same on an iGPU but not a dGPU. yes, the iGPU is actually a client of the ringbus on intel so some of these bugs seem to have been overlooked (I'd guess the fix probably improved performance on the iGPUs too lol). AMD has always interfaced the iGPU via PCIe[0] which probably helped modularity. https://www.techpowerup.com/review/amd-ryzen-7-5700g/3.html https://www.techpowerup.com/review/amd-ryzen-7-5700g/3.html Plus in general AMD just seems to be better at modularity and re-use period. I think that's the biggest headwind for Intel in general. Every. single. product. is completely one-off and custom and has its own set of bugs. Just define an interface and get used to it. But I think that's a Conway's Law situation of the hardware design resembling the org structure. Intel is a mess inside and so are their products. [0] Infinity Fabric is de-facto coherent PCIe fabric, the intel equivalent would be putting it over DMI. And AMD explicitly offers IF as a CXL competitor too. I do love that in the modern era everything is PCIe, the many-faced god. And it all just works - plug your OCP 2.0 card or M.2 card into an adapter and away you go, or tunnel pcie over Oculink or MCIO, etc. The greatest tech success story of the last 30 years.
- paulmd 3y ago> One interesting observation I made is that usually mainstream midrange cards seem to hover just over 200mm². That includes for example cards like the HD 7850, but extends all the way to the likes of the RX 580, RX 5700 XT, but also the GTX 1060 or GTX 660. Those cards are all 256 bit if AMD and 192 bit otherwise, funnily enough. this is indeed an interesting observation btw, and I have mused before that it's interesting how NVIDIA leans towards narrower buses with more advanced tech while AMD tends to lean towards plainer wider buses. Which is sort of the same observation but from the other direction. by that I mean, they were first to lean into (lossless) delta compression, and continued to retain an advantage in compression ratio for most subsequent generations. They were first to lean into quad-data rate with GDDR5X, first to lean into PAM4 with GDDR6X, etc. They very clearly favor narrower buses with higher "intensity" for their high-end stuff. Supposedly this reduces the power-per-bit-transfered but man looking at GDDR6X I really don't know, 3070 Ti is massively worse at efficiency than the 3070. Evidently they don't clock down the memory very well and at full transfer rate the total power is significantly higher (even if the per-bit is less). I wonder when we'll see the return of consumer HBM cards. It feels like the time has to be approaching soon, especially with high-NA reticle limit hitting in a gen or two. Memory PHYs are an obvious thing to cut back on for NVIDIA in particular, since they don't have MCM yet. the analogy also applies to high-end where NVIDIA is usually topping out at 384 and AMD has gone as high as 512b before. I think this is no longer possible with GDDR6/6X due to signal integrity/routing and the pcie card dimensions, everyone is topping out at 384b now, even AMD.