3 ms·
Your table doesn't include the new "tensor core" TFLOPS of V100. That's a core that does 4x4 FP16 matrix multiplication + 4x4 FP32 accumulation in one go. Tha
by Tom1971 9y ago
Your table doesn't include the new "tensor core" TFLOPS of V100.
That's a core that does 4x4 FP16 matrix multiplication + 4x4 FP32 accumulation in one go.
That's where V100 gets its boost, up to 120 TFLOPS.