4 ms·
> the largest Cloud TPU Pod configuration tested (256 chips) delivers a 200X speedup over an individual V100 GPU So a Cloud TPU chip is slower than a V100 chip
by ianhowson 8y ago
> the largest Cloud TPU Pod configuration tested (256 chips) delivers a 200X speedup over an individual V100 GPU
So a Cloud TPU chip is slower than a V100 chip? Seems like that's the answer.
- KenoFischer 8y agoWell, you lose some performance to distributed memory overhead of course, plus the TPU measure for that number is a previous generation TPU. The new ones are a decent bit faster. Plus you get four chips per "Cloud TPU". Putting all that together, TPUs are certainly a very competitive offering.
- wskinner 8y agoThey seem roughly comprable. If you can get 85% scale out efficiency, based on the numbers in this post, 256 V100s would be at parity with 256 TPUs. In practice, you can get 90% scale out on Resnet with last year’s GPU and networking hardware. Network bandwidth is the bottleneck, and with better network hardware that could be improved. Maybe that’s why Google didn’t publish a benchmark for 256 V100s. I’m also curious what the is the performance of the network interconnects in the TPU pod. I couldn’t find it documented, though I didn’t look super hard. [1] https://github.com/uber/horovod/blob/master/docs/benchmarks.md https://github.com/uber/horovod/blob/master/docs/benchmarks....