4 ms·
It is actually 45 teraflops per TPU2 chip and 180 teraflops for a module (4x chips): https://arstechnica.com/information-technology/2017/05/google-brings-45-ter
by asdf_ 9y ago
It is actually 45 teraflops per TPU2 chip and 180 teraflops for a module (4x chips):
https://arstechnica.com/information-technology/2017/05/google-brings-45-teraflops-tensor-flow-processors-to-its-compute-cloud/ https://arstechnica.com/information-technology/2017/05/googl...
V100 is 120 teraflops of tensor ops per chip:
https://arstechnica.com/gadgets/2017/05/nvidia-tesla-v100-gpu-details/ https://arstechnica.com/gadgets/2017/05/nvidia-tesla-v100-gp...
- jlebar 9y agoA GPU is also comprised of multiple chips (RAM, etc). I don't think "performance per discrete piece of silicon" is an interesting metric.
- asdf_ 9y agoA single V100 board is vastly smaller than a TPU module with 4 TPU chips. The closest comparison in terms of size would be 1 Volta DGX-1 (8x V100s) compared to 2 TPU modules (8x TPU2 chips). Volta DGX-1: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/dgx-1/data-center-products-dgx-1-components-843-u1.jpg https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent... TPU Module: https://storage.googleapis.com/gweb-uniblog-publish-prod/images/tpu-V2-hero.width-1000.png https://storage.googleapis.com/gweb-uniblog-publish-prod/ima... And for completeness, this is the size of a single V100: https://cdn.arstechnica.net/wp-content/uploads/2017/05/34446705711_76ab786243_o.jpg https://cdn.arstechnica.net/wp-content/uploads/2017/05/34446... You can see that 8x V100s are still more computationally dense than 8x TPU2 chips. Density is a very important factor in datacenter design.