5 ms·
For LLM, INT8 is old news but still exciting. FP8 would definitely be an improvement. However the new coolness is INT4. > Excitingly, we manage to reach the
by fswd 4y ago
For LLM, INT8 is old news but still exciting. FP8 would definitely be an improvement. However the new coolness is INT4.
> Excitingly, we manage to reach the INT4 weight quantization for GLM-130B while existing successes have thus far only come to the INT8 level. Memory-wise, by comparing to INT8, the INT4
version helps additionally save half of the required GPU memory to 70GB, thus allowing GLM130B inference on 4 × RTX 3090 Ti (24G) or 8 × RTX 2080 Ti (11G). Performance-wise, Table 2
left indicates that without post-training at all, the INT4-version GLM-130B experiences almost no
performance degradation, thus maintaining the advantages over GPT-3 on common benchmarks.
Page 7 https://arxiv.org/pdf/2210.02414.pdf https://arxiv.org/pdf/2210.02414.pdf
- cypress66 4y agoHopper seems to drop int4 support so maybe it's old news now? https://en.m.wikipedia.org/wiki/Hopper_(microarchitecture) https://en.m.wikipedia.org/wiki/Hopper_(microarchitecture)
- dragontamer 4y agoAt this rate, we're going to end up with FP1 (1-bit floating point) numbers... I guess that's nonsensical. 1-bit in all FP is the sign bit. So I guess the minimum size is 2-bit FP (1-bit sign + 1-bit exponent + 0-bit implicit 1 mantissa)
- kortex 4y agoAt FP2, you are probably better off with {-1, 0, 1, NaN} (sign+mantissa) rather than sign/exponent. You basically bit pack. FP3 gives you sign, 1x "exponent", 1x mantissa, so still kinda bit packing. I could see FP4 with sign, 1x exponent, 2x mantissa. Exponent would really just be a 4x multiplier, giving +/-0,1,2,3,4,8,12 Or invert all those, so you are expressing common fractions on 0..1
- dragontamer 4y agoReal life has the E3 series: 1, 2.2, 4.7, and then 10, 22, 47, 100, 220, 470, 1000, etc. Etc. EEs would recognize these values to be the Preferred Resistors values for projects (though more commonly the E6 series is used in projects, the E3 and E1 values are preferred) That's 3 values per decade, which is slightly more dispersed than a FP4 that consists of 1 sign + 3 exponent + 0 (implicit mantissa 1 bit). Or the values -128, -64, -32, -16, -8, -4, -2, -1, 1, 2, ... 128. Maybe we can take -128 and call that zero instead, cause zero is useful. -------- Given how even E3 is still useful in real world electrical engineering problems, I'm more inclined to allocate more bits to the exponent than the mantissa.
- ben_w 4y ago> Real life has the E3 series: 1, 2.2, 4.7, and then 10, 22, 47, 100, 220, 470, 1000, etc. Etc. Took me until today to realise that sequence is a rounded version of 10^(n/3) for integer n.
- simcop2387 4y agoAlmost so, there are some manual adjustments to help with overlap from tolerances at various places, but that pure math layout of it would be where you'd probably want to do it for ml
- stkdump 4y agoE3 is only relevant to decimal numbers. I don't see how computing can benefit from it, unless base conversion for human interaction is a particularly large part of the problem. And there is a reason why stuff like BCD isn't used anywhere near anything resembling high performance computing: it's practically never worth it.
- Dylan16807 4y agoIf you're going to bother doing floats you should probably make them balanced around 1. And exponent seems to be much more important for these small sizes. The first paper that shows up for FP4 almost has negative mantissa bits. Their encoding has 0, 1/64, 1/16, 1/4, 1, 4, 16, 64.
- SideQuark 4y ago1 bit would work fine - make the values represent +-1 or so
- varispeed 4y agoThat once we get into asymmetrical number coding so that you could use numbers that take fraction of bits.
- visarga 4y agoI think I read somewhere it only goes as low as int4. Can't find the reference.
- dimatura 4y agoBinary neural networks, where weights and/or activations are just 0/1s, are an active research area. In theory they could be implemented very efficiently in hardware. But in contrast to FP16 (or to some extent, int8), just quantizing FP32 to 1 bit doesn't work very well. There have been successful methods in practice. There was a company called Xnor.ai that was built partially around this technology, but it was sold to Apple a couple years ago. I don't know what's the current SOTA in this area, though.
- nutanc 4y agoWe have created a binary embedding to help us with the embedding sizes. I think in the future we will see a lot more research in reducing the model and embedding sizes. https://medium.com/ozonetel-ai/compressing-bert-sentence-embeddings-6120c84f5f4c https://medium.com/ozonetel-ai/compressing-bert-sentence-emb...
- dyno12345 4y ago*binarized
- mike_hock 4y agoSo +0, +infinity, -0, -infinity, no NaNs?
- NikkiA 4y ago1-bit unsigned float would probably be useful to some
- maizek 4y agoThe GLM work seems exciting but has there been any other research groups/LMs that have achieved similar performance even after such drastic quantisation?
- Straw 4y agoFor inference, sure. Not yet for training. Weight quantization is much easier than weight + activation.