5 ms·
You may be conflating distillation, which may be a fault of other people getting it wrong. Going from floating to fixed point doesn't change precision or size.
by Neywiny 12d ago
You may be conflating distillation, which may be a fault of other people getting it wrong. Going from floating to fixed point doesn't change precision or size. It changes representation. For example, if all your values are between 0 and 1 or -1 and 1, you're not using the entire range of a standard IEEE float. If you instead use Q31 format, you get basically 31 instead of 23 bits of mantissa. So it's an increase in precision if and only if you can basically stretch your sub-range that you were using over a larger range of bit representations. You have to think of the total number of values. 32 bit is 4 billion. 80 bit extended precision is a lot more than that. But he didn't say I think what his fixed point bit width was. That's what matters.
Now maybe you don't need to represent 1/2^32 in precision. Maybe you just need to know to 1/8th. That's where you can quantize to a lower precision to save space.
But if you start with 32 bit float and quantize to 32 bit fixed, I struggle to think how that saves storage.
- 4RealFreedom 12d agoI wasn't conflating distillation. You are basically repeating what I said - you're describing quantization. The OP was talking about changing the computation. The article says the 8- and 16-bit integer approach actually had higher resolution than the floating-point approach. You said we already had quantization to do this but quantization isn't what the OP was describing.