2 ms·
Turns out using floats is a feature and not a bug?
by shomp 1y ago
Turns out using floats is a feature and not a bug?
- Dylan16807 1y agoNo, I don't think so, in that I don't think anyone has ever called that a bug.
- shomp 1y agoIn the paper summary they did not call it a bug explicitly, but they do say there are 32x improvements in using single bits instead.
- reactordev 1y agoTo memory, sure. At the cost of 32x slower speeds.
- Dylan16807 1y agoThat's an obvious exaggeration. The competition is using smaller weights already, some of which are floating point and some of which aren't. And they use full size floats for training.
- imtringued 1y agoThat means their paper is actually worse than SOTA, which is concerned with training in fp4 natively without full precision [0] for QAT. [0] "full precision" in ML usually means 16 bit floats like bfloat16
- Dylan16807 1y agoI wouldn't say "worse". It's focusing on inference cost and leaving training at a default for now.