5 ms·
Interesting points I took from the paper[1]: * They actually started deploying them in 2015, they're probably already hard at work on a new version! * The TPU
by struct 10y ago
Interesting points I took from the paper[1]:
* They actually started deploying them in 2015, they're probably already hard at work on a new version!
* The TPU only operates on 8-bit integers (and 16-bit at half speed), whereas CPU/GPUs are 32-bit floating point. They point out in the discussion section that they did have an 8-bit CPU version of one of the benchmarks, and the TPU was ~3.5x faster.
* Used via TensorFlow.
* They don't really break out hardware vs hardware for each model, it seems like the TPU suffers a lot whenever there's a really large number of weights and layers that it must handle - but they don't break out the performance on each model individually, so it's hard to see whether the TPU offers an advantage over the GPU for arbitrary networks.
[1] https://drive.google.com/file/d/0Bx4hafXDDq2EMzRNcy1vSUxtcEk/view https://drive.google.com/file/d/0Bx4hafXDDq2EMzRNcy1vSUxtcEk...
- make3 10y ago"The TPU only operates on 8-bit integers" The 8 bit part is fine, but integers? what the hell. that's a new one for me
- mtgx 10y agoNvidia has shipped a couple of INT8 inference cards as well: http://www.anandtech.com/show/10675/nvidia-announces-tesla-p40-tesla-p4 http://www.anandtech.com/show/10675/nvidia-announces-tesla-p...
- hamilyon2 10y agoYes, how do they perform? Linked article cites 47 TOPS on 8 bit integer arithmetic and they are useful for training too.
- struct 10y agoI imagine that once they've trained the floating-point models, they'll then quantize them into integers to make inference faster. It's not something I've done, but I imagine that the limited range of the of integers may cause problems (though they say in the paper that the 16-bit product can be accumulated to something that's 32-bit). The features to do this will be coming fairly soon to regular TensorFlow too.[1] [1] https://youtu.be/0r9w3V923rk?list=PLOU2XLYxmsIKGc_NBoIhTn2Qhraji53cv&t=1617 https://youtu.be/0r9w3V923rk?list=PLOU2XLYxmsIKGc_NBoIhTn2Qh...
- dlubarov 10y agoAnecdotally, it seems most models can be quantized to 8 bits without much loss of accuracy, and fixed point arithmetic requires much less hardware. Training is still done with floating point though.
- ChuckMcM 10y agoThis, when you get right down to it a lot of models do fine with only 256 unique weights.
- gcr 10y agoYeah...you could implement 16-bit multiplication/addition as two 8-bit multiplications plus carry... so worst case, if you want 16-bit multiplies, you implement it yourself
- stephencanon 10y agoFour multiplications, not two (three if you only need the low 16b of the result).
- yorwba 9y agoYou can get the full 32-bit result in only three multiplications using https://en.wikipedia.org/wiki/Karatsuba_algorithm https://en.wikipedia.org/wiki/Karatsuba_algorithm , but it requires more additions.
- stephencanon 9y agoYeah, in practice Karatsuba isn't a performant option for small operands unless your multiplier is catastrophically slow. (And it still doesn't get you to two multiplications.)
- alttab 10y agoAgreed - however as we progress I expect a comment like this to be akin to Bill Gate's 64K comment.
- woodson 10y agoTake a look at Google's gemmlowp low-precision GEMM library: https://github.com/google/gemmlowp https://github.com/google/gemmlowp (Used in tensorflow)
- throwaway71958 9y agoNote however, that on Intel it's actually slower than run off the mill float32 linear algebra library like Eigen or OpenBLAS. Its main forte seems to be ARM.
- mtgx 10y agoAnd what's amazing is that it was built on 28nm. So TPU 2.0 could increase by another 2x in perf/W just by going to 14nm (most likely) - even more if it's built on newer processes than that. Intel's latest chips will be even further behind compared to the next-generation TPU than Haswell was compared to TPU 1.0.
- struct 10y ago28nm was quite a cheap fabrication technology even in 2015, but it costs a lot to have a completely custom production run. My guess it that it approximately works out in savings of power and space over the lifetime of the chip. It probably doesn't make sense for them to move to something smaller (and thus more expensive) whilst the performance benefit remains so substantial. If I were Intel, I probably wouldn't lose too much sleep over it either, because you still need something to attach the highly-specialised TPU to, and that'll be a Xeon for the forseeable future.
- monk_e_boy 10y agoAgreed, the market for these chips is pretty small at the moment. What other platforms would need these other than cloud? Cars? Phones .... maybe?
- pavanky 10y agoThose limitations sounds awfully similar to that of an FPGA..
- vvanders 10y agoI was going to say as well. It seems like if caches are the bane of sequential processing(CPU) then routing has to be the counterpart on the parallel(FPGA/ASIC) side of the equation.
- shepardrtc 10y agoHere's a paper they published a little while ago about limited numerical precision and deep learning: https://arxiv.org/abs/1502.02551 https://arxiv.org/abs/1502.02551
- tarlinian 10y agoThe lack of real 8-bit comparison data makes the whole paper a little suspect IMO...it's sort of like the early GPU papers that claimed 100x improvement over the CPU while running x87 scalar CPU instructions...the benefits are definitely still there but handicapping one architecture when it has features that are specifically capable of doing this is a bit stupid. It's not like they didn't have to do a lot of work on TF to make it output TPU instructions. When you're down to a 4x improvement...the benefits of specialized accelerators start to become somewhat questionable. I do like that they highlighted the importance of low latency output though...that's even more critical for future non "Web" applications which have to run in real time.
- bobdole1234 10y agoThe difference is the power numbers. 3.5x faster than CPU doesn't sound special, but when you're building inference capacity by the megawatt, you get a lot more of that 3.5x faster TPU inside that hard power constraint.
- nickpsecurity 10y agoRegarding 8-bit numbers, here's a thread on why 8 bits are enough and an old product that used that to good effect: https://news.ycombinator.com/item?id=10244398 https://news.ycombinator.com/item?id=10244398 http://www.eetimes.com/document.asp?doc_id=1140287 http://www.eetimes.com/document.asp?doc_id=1140287 It's something that keeps getting rediscovered. I know embedded industry shoehorns all kinds of problems into 8- and 16-bitters. Some even use 4-bit MCU's. Might be worthwhile if someone does a survey of all the things you can handle easily or without too much work in 8-16-bit cores. This might help for people building systems out of existing parts or people trying to design heterogenous SOC's.