4 ms·
> (and a fair number of people (including reviewers!) actively propagandize about these cards, like complaining the 4060 or 4060 Ti is actually slower than its
by ColonelPhantom 3y ago
> (and a fair number of people (including reviewers!) actively propagandize about these cards, like complaining the 4060 or 4060 Ti is actually slower than its predecessor - they're absolutely not, even at 4K (tough for a 6600XT tier card!) the 4060 Ti is still faster on average. And as a general statement, AMD made the same switch to 128b memory bus last gen already without catastrophic issues, like the aforementioned 6600XT. A card can still be a mediocre step without making shit up about it, and the 3060 Ti was the absolute peak of the 30-series value so it's understandable the 40-series struggles to drastically outpace it. But it's not slower, either.)
Yeah, the 3060 Ti was probably the best 30-series card, but the 4060 Ti being 128 bit is a much bigger issue than the 6600 XT or 4060 being 128 bit. You see, the 6600 XT succeeded the 192-bit 5600 XT, and the 4060 succeeded the also 192-bit 3060. But the 3060 Ti was a 256 bit card.
I'm also pretty sure that the 3070, which is only a small step up from the 3060 Ti, often outpacing the 4060 Ti and it becomes a much worse thing, also considering that the 3060 Ti and 4060 Ti both launched for $400. (The 6600 XT also launched for an MSRP of $380, but this was amidst the huge pricing bubble and AMD didn't want to hand over all of that to scalpers.)
To add insult to injury, the 4060 Ti is available in a 16 GB variant, which has so much memory relative to its bit width that Nvidia has to put the chips both front and back. GA104 cards which wouldn't need such tricks did not get one.
The 4070 and 7800 XT are fine from a value perspective, but it's above the price range that most people are in. But the value proposition drops off quite hard there. GPUs like the 7600 or 4060 Ti offer almost nothing over their predecessor in terms of gaming performance, so it makes sense why nobody buys them. But those are the segments that are supposed to sell well.
- paulmd 3y agoI agree, 128b is too small for the 4060 Ti's price point. Both AD106 and AD103 really needed another 2 memory controllers per die, that would have given you 4060 Ti 12GB/24GB and 4080 16GB/4080 super 20gb (with quadros using clamshell). They missed the market expectations on VRAM at the price (and performance) levels they wanted to target, and it probably would only have increased the die size by 10% or so to have the extra memory controllers. My suspicion is that these are pandemic-brain decisions. I think Ada entered mass production at the start of 2022 absolute latest (remember the "TSMC refuses to cancel 5nm wafer order" headlines?) so it would have been specced early/mid 2021 most likely. It's about a year from tapeout to launch typically, and if rumors were true they were ready to launch in june 2022 then they would have taped out in mid-2021 and been in volume late 2021/early 2022. Everyone in early 2021 was trying to absolutely maximize the number of units shipped, and if you need 4GB more memory per card then you ship 33% less units for your GDDR supply, you get more than 10% less dies per wafer, etc. It would have seemed like a good decision at the time especially with mesh shaders eventually coming in and reducing a bit of the VRAM pressure etc. As it stands, short of taping out some new dies, NVIDIA really cannot do anything about some of the gaps in the market because their only option is insane cutdowns on AD102 and AD104 to fit the gap. Cutting AD102 down to (rumored) 4080 Super shader count would be a 40% cut. Cutting AD104 down to 4060 Ti shader count would be a 50% cut. That is not yield harvesting, that's throwing away half your die. And that's why a 4080 Super 20GB was absolutely never in the cards once AI started to really take off - NVIDIA can already sell those as Quadro A5000 or A6000 anyway, why would they throw away half of a working chip to make it cheaper for gamers? 4090D is absolutely the final nail in that coffin, it will not happen at this point, every shitty yield they have will go there. AMD themselves have gone a different route and disambiguated memory from compute. They already put 2 memory PHYs per MCD attached to a single infinity link (so they are "doubling up") which relieves the pressure of GDDR density stalling out at 16gbit for a prolonged (I think unexpectedly) long period of time. It also provides a degree of physical fanout which relieves the routing pressure etc (look at the footprint of an A6000 or RTX 3090, it's crazy how tight they rammed the chips in). It surprises me that AMD haven't done a 4-PHY MCD. Even if you can't route that many GDDR lanes, you could also do sort of the opposite of the 7900GRE/7900M and put a smaller chip in a larger package but it commands the full memory bus width of a 7900XTX, so you have Radeon Pro 7800 48GB or whatever. It's not without its downsides (RDNA3 idle and low-load power consumption remains atrocious and probably always will due to the link power and due to the caches being on the other side of the link) but they're accidentally sitting on the key to big-VRAM cards at the moment of big-LLM models. What could be more of a push for ROCm then "hey here's a Radeon Pro W7800 48GB for $1499 and we can actually ship them today"?
- ColonelPhantom 3y agoYeah, I agree fully on the "pandemic brain" thing. One interesting observation I made is that usually mainstream midrange cards seem to hover just over 200mm². That includes for example cards like the HD 7850, but extends all the way to the likes of the RX 580, RX 5700 XT, but also the GTX 1060 or GTX 660. Those cards are all 256 bit if AMD and 192 bit otherwise, funnily enough. It seems die sizes are on the rise though! The 6600 XT is a similar size to the 5700 XT despite having much less hardware (except cache), although I guess RDNA2's better clocks also ate some die space. The RTX 3060 is also reasonable large at 276 mm², which seems more like what's usually 60 Ti territory. And the 3070 is even bigger than the 1080! Even in that context though, the 4060 Ti looks fairly anemic at 188 mm². The 4060 is even smaller at 159 mm², which almost equals the RX 5500 XT that debuted at $170, just over half the cost of the 4060. Intel's current GPUs are meanwhile comically large seeing as the 400 mm² or so ACM-G10 competes with N23/N33 and GA106 which are much smaller, and even the tiny 4060. I think that's mostly because of the way too fine grained EUs which also makes the architecture very complex with slices and subslices. But at least for Battlemage they seem to be planning to grow the EUs, rather than throwing on more. It's also worth noting that if Nvidia decided to "double up" the memory on the 3070 in a similar way to the 4060 Ti, they could have made a 32 GB version of it. Intel could also decide to do something like that for Battlemage, or of course AMD could do it with the 7900 XTX or similar. Despite AMDs weaker AI performance (no real matrix multiplication acceleration, unlike Nvidia or Intel) it would at least get AI developers to likely care about their cards as it means they have the best flagship in one important metric.
- paulmd 3y agoWith the way wafer costs keep increasing, physically smaller dies are inevitable. 50% higher density at 30% higher cost or whatever implies that if you keep the same die size, then cost goes up 30%, and in market terms that die has moved up a product tier. That number I was referring to is total transistor count and transistors-per-$ not die area, the cost of producing the same sized (eg 200mm2) die across nodes has skyrocketed, a 4060 Ti is far far more expensive than a 1060 to produce. Clamshell also adds PCB/assembly cost which people never account for. And all of this has to have partner margins rolled into it too etc. Eg at one point it was like $3.50 per 8gbit module, assuming 16gbit modules are comparable then 8GB x $4 = $32 of memory, but probably there's another $10-15 in the PCB and assembly, plus partner margin, etc. $50 would be partners losing margin, $75 probably they break even, $100 definitely is pushing up margin a bit, but that's also what people wanted, if you remember back to the bandwagon around EVGA. I think in the rush to bandwagon everyone just forgot who ultimately pays that margin ;) (and of course in launch terms, once cards are manufactured it's painful to open them up and remanufacture them all, are you going to desolder the BGA to put it on a new PCB, one by one?) I've commented the same thing before, that 3060 Ti 32GB would also be a very appealing product at the moment. I guess they have their reasons but it seems like a slam-dunk cashgrab at this moment of AI mania? RDNA3 does have WMMA which should at least get the foot in the door for things like AI/ML(-weighted TAAU) upscaling. I expect them to take a crack at that with FSR4, I strongly discount rumors that they're not working on that, especially with the console refresh (PS5 Pro at least) having RDNA3 and some kind of AI core. Even if sony builds one in-house, AMD totally have to be working on one internally too, just there's various reasons not to talk about it externally etc (people wanted them to not announce overly early etc, this is what that looks like). They'd want something for devs who want to validate once across multiple platforms anyway, they definitely will launch something within a year if not alongside the PS5P itself. WMMA may not be the answer for pure training horsepower though, NVIDIA having full-fat tensor units (like CDNA's) in their gaming GPUs was a good move (I won't even say "lucky", they've worked for it). But RDNA3 WMMA is better than nothing. I really think Intel's architecture is aimed at a future-node (plus an attempt to get the foot in the door on game/driver support). wave-8 is indeed way too finely grained for most workloads right now, but the pendulum is swinging back from cache to logic, and that means you have more transistors to spend on such frivolities. Growing the EUs definitely makes sense to me, like they were proofing the control/scheduling side at a small scale first, now they build the rest of it, etc. It just feels obviously wrong to do wave-8 right of the gate, on a 3060/3070 tier product at best, so what was the real goal? Some kind of incremental/iterative development, presumably. But yeah intel is a can of worms. The last time I checked (probably 6mo ago) they were operating the graphics division at -200% margin, which is needless to say stunning. I just also think they really have no choice, the writing is on the wall with NVIDIA Grace and MI300X and Apple Silicon Max/Ultra that big APUs are the future, both for a number of consumer segments but especially for a lot of HPC and even enterprise segments. Can't be competitive in HPC or enterprise if you don't have it, and even for laptops, you can't not have an iGPU. So if you cancel it do you go to Imagination or someone and license PowerVR like the bad old days of Atom? That's not a compelling product when AMD is turning out chips like 7840HS and soon the Strix Point/Strix Halo line. Other people tend to read that as "-200% margin, they'll cancel it any day now" but I actually read it as the opposite, it's "-200% margin and it's so important they are barreling through anyway". They are canceling stuff left and right but they don't have a choice on Xe/Arc regardless of the cost. And specific dies being canceled based on the current stuff doesn't mean anything in that context - that's just staying agile rather than over-committing to a schedule before you know what product you're going to build. But in general Intel can't seem to execute much of anything very well these days. Every good release feels like a fluke and they're back to the woodshed in no time. Alder Lake came with the loss of AVX-512, the 2.5gbe NICs are stuck in a state of permanent re-spins and steppings, Sapphire Rapids has some fatal power flaw that needs a 1200W psu (and they're not kidding) due to silly 700W+ transients and is getting a respin just in time for emerald rapids, meteor lake was basically 6mo late and still had a messed-up BIOS and still undershot performance targets, on and on, and they're cutting pay and firing staff. Death-spiral territory right there.