3 ms·
The aforementioned google translate paper suggests so. My hunch is that TPUs work with a mix of 8 bit and 16 bit integer arithmetic on quantized networks with
by ogrisel 10y ago
The aforementioned google translate paper suggests so.
My hunch is that TPUs work with a mix of 8 bit and 16 bit integer arithmetic on quantized networks with quantized data [1].
As far as I know nobody has managed to make SGD work properly with integer weights and that could be a reason why TPUs are not used for training (yet).
[1] https://petewarden.com/2016/05/03/how-to-quantize-neural-networks-with-tensorflow/ https://petewarden.com/2016/05/03/how-to-quantize-neural-net...
I could be completely wrong though.