4 ms·
The "NVIDIA Tesla Family Specification Comparison" table indicates 112 TFLOPS "Tensor Performance (Deep Learning)" for Tesla V100 (PCIe). Is that double precis
by visionscaper 9y ago
The "NVIDIA Tesla Family Specification Comparison" table indicates 112 TFLOPS "Tensor Performance
(Deep Learning)" for Tesla V100 (PCIe).
Is that double precision?
The Nvidia 1080 Ti has a double precision performance of 332 GFLOPS [1].
If the above number is for double precision computing, the Tesla V100 (PCIe) would be about 337 times as fast (!!)
Does anyone have more insight into these numbers?
[1] https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_processing_units https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces...
UPDATE : I should have read the article more carefully, it seems to be a mix of FP16 (half precision) and FP32 (single precision). That would likely mean a factor ~10 in computation performance (specifically for deep learning)
- wmf 9y ago"...matrix operations with FP16 inputs and FP32 accumulation..." http://images.nvidia.com/content/volta-architecture/pdf/Volta-Architecture-Whitepaper-v1.0.pdf http://images.nvidia.com/content/volta-architecture/pdf/Volt...
- visionscaper 9y agoThanks, I also read the article better now. I updated my comment.
- davesque 9y agoOn this topic, hasn't there been some research that has shown that lesser precision values can be used in neural network algorithms without much loss in performance?
- visionscaper 9y agoThere had been a lot of research on this topic (just Google it), with good results.