3 ms·
Probably not. The noise in the training data itself, noise from dropout and various other sources are going to trump any quantization noise anyway, so you can u
by benanne 12y ago
Probably not. The noise in the training data itself, noise from dropout and various other sources are going to trump any quantization noise anyway, so you can use fairly inaccurate representations. This is why everyone's using gamer GPUs for this stuff, and not Tesla cards: single precision is enough, and much cheaper to come by.
Although, I guess the cutoff might have to be a bit higher than 8.23. Maybe the neuron activations would never exceed that range, but some intermediate computations could.
Supposedly the new Maxwell GPUs have some new instructions for working with half-precision floats ( https://developer.nvidia.com/sites/default/files/akamai/opengl/specs/GL_NV_shader_atomic_fp16_vector.txt https://developer.nvidia.com/sites/default/files/akamai/open... ). I wonder how complete this implementation is, because half-precision might be sufficient to train a convnet, and this would result in a significant speedup.
- smhx 12y agoIn these deep relu networks which are not renormalized in between, (like overfeat which has no normalization), some of the activations become pretty big in size! (in the order of 1e3). You also cant clip them and get away with it, you have to either renormalize the layers to do half-precision (and live with the extra cost) or stick to full-precision. I was doing fun stuff early this year that did fixed-precision nets (8-bit/16-bit). Things get very interesting :)
- varelse 12y ago16-bit float has dynamic range from + or - 6.5e04 to 6.1e-05. http://en.wikipedia.org/wiki/Half-precision_floating-point_format http://en.wikipedia.org/wiki/Half-precision_floating-point_f... That's plenty IMO for most inputs and weights. Where it gets tricky is in accumulation. You could constrain the weights for each unit I guess, but this is the sort of work best done under the hood rather than by the data scientist IMO. I'd personally choose 32-bit accumulation just because it would drastically simplify code development. I've also worked with fixed precision elsewhere. It's awesome if you understand the dynamic range of your application. It's a migraine headache if you don't.