4 ms·
V100 GPUs have non tensor core fp16 operations too I think
by deepnotderp 5y ago
V100 GPUs have non tensor core fp16 operations too I think
- ml_hardware 5y agoThat's true.. in fact, seeing V100 FP16 < T4 FP16 makes me believe you're right, the V100 should be much faster if the tensor cores were being used.
- woadwarrior01 5y agoYes. Non tensor core fp16 ops are the default. Tensor cores are essentially 4x4 fp16 mac units and there's a requirement that matrix dimensions are multiples of 8[1] that needs to be met for them to be used. [1]: https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html#tensorop https://docs.nvidia.com/deeplearning/performance/mixed-preci...