4 ms·
ML might benefit a lot from 10bit bytes. Accelerators have a separate memory space from the CPU after all, and have their own hbm dram as close as possible to t
by kleton 4y ago
ML might benefit a lot from 10bit bytes. Accelerators have a separate memory space from the CPU after all, and have their own hbm dram as close as possible to the dies. In exchange, you could get decent exponent size on a float10 that might not kill your gradients when training a model
- londons_explore 4y agoThere seems to be as-yet no consensus on the best math primitives for ML. People have invented new ones for ML (eg the Brain Float16), but even then some people have demonstrated training on int8 or even int4. There isn't even consensus on how to map the state space onto the numberline - is linear (as in ints) or exponential (as in floats) better? Perhaps some entirely new mapping? And obviously there could be different optimal numbersystems for different ML applications or different phases of training or inference.