4 ms·
> RDNA3 does have WMMA which should at least get the foot in the door for things like AI/ML(-weighted TAAU) upscaling. Do keep in mind that RDNA3 WMMA is very
by ColonelPhantom 3y ago
> RDNA3 does have WMMA which should at least get the foot in the door for things like AI/ML(-weighted TAAU) upscaling.
Do keep in mind that RDNA3 WMMA is very slow, running at the same theoretical TFLOPS as shader (which are doubled from RDNA2 due to very limited dual issue support that can be used in WMMA). Nvidia tensor cores and Intel XMX can run closer to 4:1 or 8:1 or so compared to vector workloads.
> It just feels obviously wrong to do wave-8 right of the gate
That's because Alchemist isn't a a true first generation product, it's a scaled-up version of Intel's (relatively mature at this point) integrated graphics product. This means that Alchemist has suffered large growing pains (Gen12 was not really designed to be used in products bigger than maybe 128 EUs, A770 is 512 EUs). It also has some form of separation anxiety, seeing its need for ReBAR. I also recall very early in Alchemist's life (pre-release) some driver optimization that had a huge benefit, which was just the wrong memory region being used, as all memory is the same on an iGPU but not a dGPU.
- paulmd 3y ago> It also has some form of separation anxiety, seeing its need for ReBAR. I also recall very early in Alchemist's life (pre-release) some driver optimization that had a huge benefit, which was just the wrong memory region being used, as all memory is the same on an iGPU but not a dGPU. yes, the iGPU is actually a client of the ringbus on intel so some of these bugs seem to have been overlooked (I'd guess the fix probably improved performance on the iGPUs too lol). AMD has always interfaced the iGPU via PCIe[0] which probably helped modularity. https://www.techpowerup.com/review/amd-ryzen-7-5700g/3.html https://www.techpowerup.com/review/amd-ryzen-7-5700g/3.html Plus in general AMD just seems to be better at modularity and re-use period. I think that's the biggest headwind for Intel in general. Every. single. product. is completely one-off and custom and has its own set of bugs. Just define an interface and get used to it. But I think that's a Conway's Law situation of the hardware design resembling the org structure. Intel is a mess inside and so are their products. [0] Infinity Fabric is de-facto coherent PCIe fabric, the intel equivalent would be putting it over DMI. And AMD explicitly offers IF as a CXL competitor too. I do love that in the modern era everything is PCIe, the many-faced god. And it all just works - plug your OCP 2.0 card or M.2 card into an adapter and away you go, or tunnel pcie over Oculink or MCIO, etc. The greatest tech success story of the last 30 years.