4 ms·
When you have so few bits, does it really make sense to invent a meaning for the bit positions? Just use an index into a "palette" of pre-determined numbers.
by teo_zero 6mo ago
When you have so few bits, does it really make sense to invent a meaning for the bit positions? Just use an index into a "palette" of pre-determined numbers.
As a bonus, any operation can be replaced with a lookup into a nxn table.
- childintime 6mo agoExactly. And pick them on the e^x curve.
- adrian_b 6mo agoAs explained in an article linked at the bottom of TFA, the weights of a LLM have a normal (Gaussian) distribution. Because of that, the best compromise when the weights are quantized to few levels is to place the points encoded by the numeric format used for the weights using a Gaussian function, instead of placing them uniformly on a logarithmic scale, like the usual floating-point formats attempt.
- 0-_-0 6mo agoYou want to make multiplication cheap, it's not just about compression
- mysterydip 6mo agoWouldn’t multiplication just be an 8 bit lookup table? a*b is just lut[a<<4+b]
- 0-_-0 6mo agoA 256 element lookup table is much bigger than a simple multiplier
- kevmo314 6mo agoMultiplication at this resolution is already implemented via lookup tables.
- ineedasername 6mo agoFor FP4, yes... sometimes... it depends. But newer Nvidia architecture eg Blackwell w/ NVFP4 does not, they perform micro block scaling in the core. For older architectures, low quants like FP4 are also often not done native, and instead inflated back to BF16, eg with BnB.
- londons_explore 6mo agoSpecifically, you want to choose 16 values, all of which you can multiply an activation value by using circuitry which is as small as possible.
- petters 6mo agoThat's a good idea and it exists: https://www.johndcook.com/blog/2026/04/18/qlora/ https://www.johndcook.com/blog/2026/04/18/qlora/ It seems quite wastful to have two zeros when you only have 4 bits it total
- saulpw 6mo agoOTOH, it seems quite plausible that the most important numbers to represent are: +0 -0 +1 -1 +inf -inf
- Dwedit 6mo agoWhy waste a slot on -0?
- saulpw 6mo agoBecause it means "infinitesimal negative" which is distinct from "infinitesimal positive".
- Dylan16807 6mo agoThat sounds pretty niche. What's a use case where you have less than 8 bits and that distinction is more important than having an extra finite value? I don't think AI is one.
- jlokier 6mo agoFor neural net gradient descent, automatic differentiation etc, the widely used ReLU function has infornation carrying derivatives at +0 and –0 if those are infinitesimals.
- Dylan16807 6mo agoBarely any information. After surviving RELU that signed zero is probably getting added to another value and then oops the information is gone. It sounds a lot worse than properly spaced values.