3 ms·
Their “CUDA core” is not warp-wide, it’s a single lane. If you’re talking about FP32 rates, they double it because of FMA (floating-point multiply-accumulate).
by subharmonicon 3y ago
Their “CUDA core” is not warp-wide, it’s a single lane.
If you’re talking about FP32 rates, they double it because of FMA (floating-point multiply-accumulate). Everyone does that.