3 ms·
I mean actually executing the operations at that precision with improved performance. Apple GPUs support both FP16 and FP32 as data types, but the ALU throughpu
by ribit 4y ago
I mean actually executing the operations at that precision with improved performance. Apple GPUs support both FP16 and FP32 as data types, but the ALU throughput for both is identical (my personal speculation is that ALUs are 32-bit only and rest is data type conversion). From the operational standpoint, Apple G13 SIMD can only do 32 flops per cycle, not more and not less.
But other GPUs support doing operations on limited-precision data types faster. And Nvidia has dedicated matrix multiplication units that can perform very wide limited precision operations per cycle (Apple has similar units but they are part of the CPU clusters).
Apple since A15/M2 offers SIMD matrix multiplication intrinsic (very similar to VK_NV_cooperative_matrix). But the performance is limited by the fact that each SIMD only offers 32 ALUs. If they add the ability to reconfigure these as 64 FP16 ALUs (or 128 FP8 ALUs) and then maybe even doubled the ALUs like Nvidia/AMD recently did with their architectures, they could achieve much higher matmul performance for ML.