9 ms·
Benchmarking Google’s new TPUv2
- angrygoat 9y agoGoogle claim 29x better performance-per-Watt with TPUs than contemporary GPUs[0]. Interesting to contrast that to the images-per-$ figure in this post, which is more like 2x. I assume there's a high capital cost for this new hardware, but when they scale it up I wonder if the ratio of cost TPU to GPU will trend towards the ratio of power-per-Watt between the platforms? Seems like a natural limit, even if it never quite gets there. [0] https://cloud.google.com/blog/big-data/2017/05/an-in-depth-look-at-googles-first-tensor-processing-unit-tpu https://cloud.google.com/blog/big-data/2017/05/an-in-depth-l...
- zitterbewegung 9y agoI think you might want to factor that they are adding their own fees? Also, I think the market may have changed since then (AWS lowered prices and new GPUs have come out). Google's workload is also different than the benchmark that is given in this.
- dgacmu 9y agoThat was TPUv1 (inference only). This article is about the new Cloud TPU (or TPUv2 as they call it), which handles both inference and training. The competitive landscape also changed a lot in the interim - NVidia added tensor cores to Volta to accelerate deep learning computation.
- gumby 9y ago> Google claim 29x better performance-per-Watt with TPUs than contemporary GPUs[0]. Interesting to contrast that to the images-per-$ figure in this post, which is more like 2x. But you aren't paying for the electricity, you're paying for processing, which is an unconnected parameter. They only "sell" these chips per use, not on the open market Presuming power is a major cost input (which I assume) their profit/op is much higher. So they could sell for less than an equivalent GPU and make more money. But they think they can get away with value pricing it (processing more/unit time is presumably worth it to many customers) and more power to them. (that last bit was not an intentional pun; only noticed after typing it)
- aje403 9y agoMaybe Google will pivot from high tech into crypto mining
- deleted 9y ago[deleted]
- chapill 9y agoI wonder if Chinese companies will use (or be allowed to use) TPUs. It seems like a pretty obvious way to have the NSA scoop up any Chinese AI advancements China may want to keep secret.
- NelsonMinar 9y agoI wonder which Chinese companies are developing their own processors like TPUs.
- chapill 9y agoWell, they do have the fastest supercomputer in the world currently and it's made with homegrown chips. No Intel ME backdoors there. Smaller chinese companies could, for a little more money, get similar performance buying 8x V100 machines from NVidia. I don't think they want to share their advancements in AI fighter pilots with USA. They have a big lead.
- olfactory 9y agoWhat is the hardest thing to accomplish with something like a TPU? Is it the IP or the fabrication? How does the TPU design offer improved performance? By leveraging IP or fabrication improvements?
- deepnotderp 9y agoNeither, it's a matrix multiply systolic array ASIC, that's been done decades ago. There are host of Chinese companies developing similar processors.
- olfactory 9y agoWhy is Google investing in its own?
- 9y ago
- bhouston 9y agoIt is hard for Google to make money on these TPUs as the whole engineering cost has to be made back from its pricing on Google Cloud, where as with NVIDIA it can pay back its engineering costs via multiple mature channels (games, super computers, and multiple cloud providers.) I wonder which is higher, the cost for creating the TPUs in terms of engineering and manufacturing or the cost differential in terms of usage as compared to NVIDIA's latest? I worry about Google long term here. I am surprised the TPU doesn't kick the ass of the NVIDIA chips.
- puzzle 9y agoGoogle probably got back a lot of the engineering costs before it even rented out the first TPU, simply by virtue of running its own workloads, without having to buy tons of CPUs or GPUs. They're also very, very good at reducing computing resource waste (I know this firsthand). I wouldn't be surprised if public TPUs are to some degree a way to print money: at least for a while, Google can probably just rent out its unused capacity that it had already planned and paid for. :-)
- arcanus 9y agoFurthermore, computer hardware is not static. Is this a real long term investment by Google? If they do not continue to improve on process, they will fall behind in just a few years.
- Mononokay 9y agoThere's no real need to worry about Google in the long term - nVidia can make back their money solely with their GPUs; Google probably made their expenses back this weekend with searches around the Olympics. It'd be pointless for them to not use their TPUs themselves, and their main product, Adsense, uses ML.
- bitL 9y agoAre you sure about Adsense? Talked to ad pros recently and they all complained Adsense is ancient (still mySQL?) and often broken; doesn't look like Google emphasizes it despite being their cash cow, more like deep state of neglect.
- dkobran 9y agoJust to clarify, is this benchmark leveraging mixed-precision mode on the Volta V100? The major innovation of the Volta generation is mixed-precision which NVIDIA claims is a huge performance increase over the Pascal generation (P100 in the case of your benchmark). Link to NVIDIA documentation on mixed-precision TensorCores: https://devblogs.nvidia.com/inside-volta/ https://devblogs.nvidia.com/inside-volta/
- elmarhaussmann 9y agoWhere specified "fp16", the V100 benchmarks use the code from https://github.com/tensorflow/benchmarks/tree/master/scripts/tf_cnn_benchmarks https://github.com/tensorflow/benchmarks/tree/master/scripts... with the flag --use_fp16=true which enables fp16 for some but not all Tensors.
- dkobran 9y agoIt's my understanding that fp16 (available on the previous generation P100) and mixed-precision (major innovation of V100) are different things and the speedup of TensorCores is entirely missing from this benchmark. Unlike the general purpose P100, the TPU is a heavily optimized chip built for Deep Learning, hence it's performance increase. However, the V100 is also heavily optimized for Deep Learning (arguably the first non-GPU chip) from NVIDIA. I'm in no position to defend NVIDIA here haha but it seems like the benchmark misses the point if this is indeed the case.
- elmarhaussmann 9y agoIt was my understanding that the TensorFlow benchmarks do make use of TensorCores on the V100. We'll verify and update accordingly.
- amelius 9y ago> In order to efficiently use TPUs, your code should build on the high-level Estimator abstraction. Does this mean it's inference-only? (I only quickly scanned the article)
- jlebar 9y agoNo, this whole blog post is about training models.
- twtw 9y agoIIRC, TPUv2 uses 16 bit floating point in some format with higher dynamic range and lower precision than standard fp16. Can someone confirm? If that is right, is the "Tensorflow-optimized" Resnet-50 using 16bit floats when running on TPUv2?
- deepnotderp 9y agoRe: fp16 dynamic range: yes.
- slashcom 9y agoWait but, the batch size is 8x bigger for the TPU? That's not a fair comparison; increasing batch size always speeds things up...
- fooker 9y agoBut typically does not have an impact on power usage. They are claiming a 29x improvement in that area.
- elmarhaussmann 9y agoAuthor here. Note that the TPU supports larger batch sizes because it has more RAM. We tested multiple batch sizes for GPUs and reported the fastest one. We'll try increasing the batch sizes as far as possible and report. The overall comparison will likely not change by much - we saw speed increases of around 5% doubling the batch size from 64 to 128. (https://www.tensorflow.org/performance/benchmarks https://www.tensorflow.org/performance/benchmarks also reports numbers for batch sizes of 32 and 64 on the P100)
- boulos 9y agoDisclosure: I work on Google Cloud. Oh! You should definitely say that. It's semi-reasonable then to choose the batch size that is optimal for the part. It'd be good to make sure this isn't why your LSTM didn't converge though...
- elmarhaussmann 9y agoI tested many different batch sizes for the LSTM, so I am pretty confident it's not the reason.
- jrk 9y ago[Edited] The top line results focus on comparing four TPUs in a rack node (which marketing cleverly named “one cloud TPU”), running ~16 bit mixed precision, to one GPU (out of 8 in a rack node), also capable of 16 bit or mixed precision, but handicapped to 32 bit IEEE 754. That is a misleading comparison. Images/$ are obviously more directly comparable, but again the emphasized comparisons are at different precision. Very different batch sizes make this significantly more misleading, still. Images/$ also only tells us that Google has chosen to look at the competition and set a competitive price; the per-die or per-package comparison is much more relevant to understand any intrinsic architectural advantage, since these are all large dies on roughly comparable process nodes.
- dgacmu 9y agoThat's why you scroll down the page to the cost comparison, which places it on a more even keel. They do also compare float16 on Volta. Physical packaging is irrelevant -- what matters is dollars to convergence and time to convergence. (I'm obviously biased - I helped with parts of the cloud-side of cloud TPU - but I presume this comment stands on its own. :-)
- jrk 9y agoTo be clear, I had read the whole post, I was just being terse since the emphasis seemed to be so heavily on an apples to bananas comparison (I believe 100% of the results cited in the prose, many in bold, are with mismatched precision and batch size), with minimal articulation of the many axes of nuance here. Precision isn't defined at all in the LSTM case, and could easily be the cause of the failure of the TPU run to converge where the GPU runs do. To a non-expert audience I think the end result is confusing and misleading. Also, while I certainly agree that the performance/dollar comparison is highly relevant to customers at a given instant, that may only tell us that Google is subsidizing this hardware now that they've deployed it, and/or that, lacking serious competition, NVIDIA has been building crazy margins into their P100/V100 prices. In understanding fundamental technological tradeoffs, and even the limits of what the pricing in a more competitive market could be, it is relevant to compare performance per unit of hardware resources (mm^2, die/package, watt, GB of HBM, etc.) In short, these comparisons are hard, and there is no one which tells a complete story. I pushed back because, while the post includes some nuance, it brushes a great deal under the rug and focuses primarily on problematic comparison (Further disclosure: I'm at least the third person in this sub-thread with some Google Brain/Cloud affiliation. I am speaking in my independent academic voice. I also think TPUs are great, having them publicly available now is great, and competition and diversity of architectural approach in accelerators is great. I appreciate the effort of the authors, but think the subtlety of these comparisons requires serious care.)
- boulos 9y agoDisclosure: I work on Google Cloud. While not perfect, I want to commend the RiseML folks for doing not only an “just out of the box” run in both regular and fp16 mode (for V100), but also adding their own LSTM experiment to the mix. We need third-party benchmarks whenever new hardware or software are being sold by vendors (reminder: I benefit from you buying Google Cloud!). I hope the authors are able to collect some of the feedback here and update their benchmark and blog post. The question about batch size comparisons is probably the most direct, but like others, I’d encourage a run on 1, 2, 4 and 8 V100s as well.
- joe_the_user 9y agoSo this is a chip that no one outside of Google is going to be able to get a physical copy of ever? It makes any benchmarks become Google-cloud benchmarks, right? Edit: I am complaining a bit about the lack of availability but there's also a real point here. If there's no source for TPUs outside of Google, Google Cloud competes only with other cloud providers and with owning physical GPUs - long term, it has no incentive to be anything but little bit more efficient than these however much it's price for producing TPUs declines.
- azinman2 9y agoI thought they were going to provide to other cloud providers? I’m also guessing if you’re willing to purchase a lot of them then they’re willing to talk...
- joe_the_user 9y agoI'd be interested if anyone has details. It may be that the other cloud providers would then sell them to those individuals. Indeed, the job of entities called "distributors" to buy big lots from manufacturers and break them up. And of course, I don't know what the point of (apparently) keeping them out of the average person's hand would be.
- newbuser 9y ago
- gok 9y agoThe bar graph seems a little whacky. It groups the TPU (which can only do FP16) with the FP32 results from the GPUs, then puts the FP16 GPU results off to the side even though that's much closer to what the TPU is doing. Impressive results regardless though; quite a bit faster than V100 than the paper specs would suggest.
- ekelsen 9y agoIt also seems like the price comparison should compare with the fp16 numbers on both platforms, not the fp32 numbers.
- elmarhaussmann 9y agoAuthor here. Good point, I agree that the FP16 GPU results should be closer or grouped with the TPU results. We'll try to update accordingly.
- alexnewman 9y agoThe entire idea that people are going to gain some huge advantage over nvidia with hardware softmax seems dubious. I do think it will buy them some time but eventually it seems as though nvidia will win this one.
- Nokinside 9y agoSpecialization brings speedups. TPUv2 is specially optimized for deep learning. Nvidia's Volta microarchitecture is graphics processor with additional tensor units. It's a General-purpose (GPGPU) chip designed with graphics and other scientific computing tasks in mind. Nvidia has enjoyed monopoly power in the market and single microarchitecture has been enough in every high performance category. Next logical step for Nvidia is to develop specialized deep learning TPU to compete with TPUv2 and others.
- deepnotderp 9y agoVolta V100 already has "tensor cores" which are basically little matrix multiplication ASICs.
- Nokinside 9y agoThat's what I said. The microarchctiecture has many unnecessary things and it's not optimized as a whole for deep learning.
- deepnotderp 9y agoI believe it was either the last MICRO* or the one before that when Dally addressed this point. The specialized hardware for graphics ends up comprising such a small portion of the overall chip that it wasn't worth it to remove it. The "GPUs were made for graphics thus aren't good for DL" argument really doesn't hold a lot of water IMO. * It might've been a difference conference now that I think of it.
- twtw 9y ago> Next logical step for Nvidia is to develop specialized deep learning TPU to compete with TPUv2 and others I don't know, this benchmark seems to show V100 doing pretty well against a specialized ASIC. It may well be that all NVIDIA has to do is cut costs on V100 to make a two V100s about as expensive as the cloud TPUv2. With increased batch size, it looks like two V100s would have performance comparable to TPUv2.
- PaulHoule 9y agoDoes this take into account the fact that you might need fewer epochs if you reduce the batch size? (as is done for the CPU?)
- elmarhaussmann 9y agoAuthor here. No, this is really only comparing the throughput on the devices. A thorough comparison should really focus on time to reach a certain quality - including all of the tricks available for a certain architecture.
- neves 9y agoWould I be able to buy one of these for my home? Or just in the cloud? If I can buy, how much would it cost?
- elmarhaussmann 9y agoAuthor here. These are only available on the Google Cloud right now. I don't think there are plans to sell them anytime soon.
- ysleepy 9y agoI'd be interested how the superior perf/watt claims holds in googles practical setup. The additional Networking gear and power supply losses and so on might make the difference less. I'm also not sure how we can take googles word for the numbers, since they might as well be eating a less-than-ideal power cost to promote their platform. Any upfront cost will probably offset by locked-in customers later on. I might just be a bit cynical though.