5 ms·
Back of the envelope, a TPU costs a little more than 2x as much as a Volta on AWS P3, and delivers a little less than 2x the performance (180 TOPs for the TPU,
by borramakot 9y ago
Back of the envelope, a TPU costs a little more than 2x as much as a Volta on AWS P3, and delivers a little less than 2x the performance (180 TOPs for the TPU, 100 for Volta). On a raw performance/$ metric, I'm not sure the TPU is that interesting.
It might be worth it if I were willing to pay a huge amount to get back results from an experiment faster, by using lots of TPUs- distributed learning on GPUs doesn't seem easy yet.
- surajrmal 9y agoIt'll make a lot more sense when the TPU pods they alluded to come out.
- jlebar 9y agoI know people don't know what to expect from tpu performance, but does anyone actually get 100tops out of Volta? I thought you'd have to spin the tensorcores and never touch memory, which is...not realistic. I know you hedged by saying "back of the envelope", but I'd much rather compare on real benchmarks than based on cited peak performance numbers, which are kind of meaningless.
- deepnotderp 9y agoThis is true of the TPU as well, check out their paper's utilization numbers. If you ignore one outlier at ~90% utilization, their utilization plummets. I'm glad people are finally looking past the b.s. "peak" numbers for once though.
- jlebar 9y agoHas Google published data on the memory bandwidth of TPU v2 (aka "cloud TPU")? I'm having trouble finding it. In any case I agree, we shouldn't be looking at the stated peak compute of either of the chips. (Disclaimer: I work at Google on XLA, and have in the past worked on TPUs.)
- boulos 9y agoFrom the blog post is the link to the fairly recent NIPS presentation: https://supercomputersfordl2017.github.io/Presentations/ImageNetNewMNIST.pdf https://supercomputersfordl2017.github.io/Presentations/Imag... which claims 2400 GB/s for the board and 600 GB/s per “chip”.
- jlebar 9y agoThis is in comparison to 900gb/s for V100.
- boulos 9y agoDisclosure: I work on Google Cloud. Peak ops/second isn’t the only thing that matters though. You have to be able to feed the units. The V100 does lots of finer-grained matrix multiplies which can make it harder to keep up. Don’t get me wrong, the V100 is a great chip. And we’re all looking forward to more (preferably third-party) benchmark results, to tease out when one is the better choice for a workload. But don’t just compare ops/second or any other architectural number.
- deepnotderp 9y agoThis makes no sense, the V100 has more memory bandwidth than both the TPU and TPUv2
- jamesblonde 9y agoYes, when training DNNs memory bandwidth is the only figure you need to look at. That's why the 1080Ti is by far and away the best bang for buck right now (ignore the EULA nonsense). It has about 55% of the memory b/w of the V100 for 10% of the price.
- boulos 9y agoWe mostly focus on the “whole board” numbers. So it’s not only units <=> “local” HBM, but NVLINK versus TPU to TPU. Sorry for the confusion. Edit for this part of the thread: the best public numbers are in the linked presentation [1]. [1] https://supercomputersfordl2017.github.io/Presentations/ImageNetNewMNIST.pdf https://supercomputersfordl2017.github.io/Presentations/Imag...
- deepnotderp 9y agoThat's... a skewed ... comparison, NVLINK is a board to board connection whereas you're talking about TPU to TPU on board communication if I understand correctly?
- boulos 9y agoThat's sort of the point though! We're actually selling these as the "board". So the right way to compare things is sort of DGX-1 style "deep learning rig" versus a board of four TPU units (or several connected). The on-chip network is a big part of its overall efficiency. I don't recall what (if anything) we've said about how we link up the boards across racks, but the folks at Next Platform looked pretty carefully at the pictures: https://www.nextplatform.com/2017/05/22/hood-googles-tpu2-machine-learning-clusters/ https://www.nextplatform.com/2017/05/22/hood-googles-tpu2-ma...