6 ms·
Google TPU Performance Analysis
- alexnewman 9y agoSo many details that people gloss over. I have used tensorflow (TF) and it is true that GPUs suck at interference at it. But it's not always the GPUs fault - TF can't do anything quantized on GPUs. It just switches back to to the CPU/TPU. - TF gets relatively poor utilization of the GPU and tends to not be careful with memory use. - I was able to do certain types of classification hundreds of times faster by seeing what TF was doing it and hand writing it in OCL. Using https://docs.rs/ocl/0.14.1/ocl/ https://docs.rs/ocl/0.14.1/ocl/. It's a super cool library for rust. Also users should checkout tensorRT https://github.com/NVIDIA/gpu-rest-engine/tree/master/tensorrt https://github.com/NVIDIA/gpu-rest-engine/tree/master/tensor.... It's not super well supported and may go away, but it is fast
- gcp 9y agoQuestion about TensorRT: It takes in Caffe trained models. Is this BVLC Caffe or NV-Caffe? In my experience the models aren't compatible between the two.
- alexnewman 9y agoEither but I've never tried
- wcrichton 9y agoDo you know why TF is getting poor GPU util? Is it the pipeline feeding data to the core compute ops, inefficiencies in the core compute ops, or something else?
- flx42_ 9y agoJust wanted to chime in on TensorRT, it's a well supported product and it's different than gpu-rest-engine. This GitHub repo is simply an example of how to use TensorRT in a specific situation.
- baybal2 9y agoBack in early-noughties, I remember that there were a company that was developing an accelerator chip for seismic data analysis for oil exploration companies. I can't remember the name now. Can anybody remember? They were proposing a chip that did nothing but a limited set of linear algebra operations at gigabit rates. They were former Transmeta people
- moonbug22 9y agoClearspeed? The HPC history books are littered with the bankrupt corpses of special purpose hardware.
- semi-extrinsic 9y agoI remember an ASIC that was supposed to accelerate multigrid preconditioners, out of some big German university. They were never able to get stuff to market fast enough to beat Intel and Moore's law. Perhaps the biggest recent success story in this field is Anton. https://en.wikipedia.org/wiki/Anton_(computer) https://en.wikipedia.org/wiki/Anton_(computer)
- nhaehnle 9y agoI really don't get how they came up with those numbers comparing CPUs to GPUs. They claim to have 3.5x as much on-chip memory as a GPU, but the R9 Fury X has 16.7 MiB of register memory compared to their 28MiB. And then of course there's caches on top of that (which funnily add up to less than the register memory, I believe). I also don't get how they come up with those MAC numbers. An RX Vega 64 can do 27 TFlop/s of half-precision arithmetic, which is way more than 1/25x the 92 TOp/s they claim for the TPU. In fact, it makes the GPU look pretty damn good, considering the TPU only does 8-bit ops. Of course I'd expect the TPU to beat a GPU in terms of perf/watt, but that's not what they're comparing on that particular slide. There's the whole question of how you manage latency in inference, but then I'd expect them to talk about the utilization of the GPU resources relative to the theoretical peak.
- aub3bhat 9y agoFloating point operations (TFlop) != Tensor Operations (TOp)
- nhaehnle 9y agoSure, but read the slide. They have 64k multiply-accumulate units, running at 700 MHz. That means ~46T/s multiply-accumulates, which means ~92T/s individual arithmetic ops. It's a standard way to measure this. I think it's fair to say that 92T/s 8-bit arithmetic ops is much less than 25x the 27T/s half-float operations of a GPU.
- deleted 9y ago[deleted]
- Symmetry 9y agoIf an 8-bit integer is sufficient for a problem there isn't anything to be gained by moving to 32 bit floats. You can't just use a 32-bit floating point operation to emulate 4 8-bit integer operations for free (or vice versa) so you just can't compare the two the way you're trying to. Especially since moving to larger precision values would balloon memory and bandwidth requirements. For an honest comparison find out what the 8-bit integer performance of the GPU is.
- shaklee3 9y agoThis article just seems odd. They're still quoting numbers from how they compared 2 years ago to Kepler GPUs. Unless they have a new TPU out, these are worse than the V100 GPU out today, so it's strange that in a field moving so fast they're constantly quoting old data. It doesn't matter anymore that you had the fastest chip in 2015. If you haven't iterated since then, you are probably losing.
- jorgemf 9y agoThe link is about TPUv1, but Google is already using TPUv2 (or maybe TPUv3, they don't talk too much about this things).
- digitalzombie 9y agoBottom of article the author missed out on the TPU v2 talk. TPU v2 is in alpha stage right now but if you're a research you can apply to use it over at google cloud service.
- mooneater 9y agoLooks to be all about TPU1? Which is inference-only. Afaik TPU2 allows for training as well, Im much more interested in that. Last line: "There was a TPU2 talk earlier that I missed that I need to look through the slides of and write up later"
- Symmetry 9y agoThe Hot Chips talks will eventually make their way onto YouTube...
- jcbeard 9y agoSeems very much "back to the future." Systolic array processors were used to accelerate neural networks in the 1980's. Great for matrix math too. (ref: http://repository.cmu.edu/cgi/viewcontent.cgi?article=2939&context=compsci http://repository.cmu.edu/cgi/viewcontent.cgi?article=2939&c...). These aren't quite the systolic array processor of old, but too close to be considered new arch/micro-arch. The formula is simple, have low precision MM to accelerate, drop in a matrix multiply unit that can be blocked for and high bandwidth memory to feed it and let it go. I'm waiting for more new takes on old arch....as fabbing chips becomes more economical, I hope to see more retro chips. Especially things that didn't quite make the jump from research to production b/c of scaling (or other reason), might now make sense.