4 ms·
Huh, I thought this is what was usually meant by quantising
by mattsan 3y ago
Huh, I thought this is what was usually meant by quantising
- refulgentis 3y agoI believe the distinction they're drawing is they're also not "cast" to float32 during processing (which I also didn't know)
- mattsan 3y agoYep I didn't know other models cast during processing
- andy99 3y agoIt's a nontrivial problem to do faster arithmetic operations with quantized values (which are not just rounded to fit in 4 bits but are usually block quantized with one or more scaling parameters stored at full precision) faster than to convert them to floats that the cpu/GPU is already optimized to multiply/add. That is what this project is proposing a solution to.
- phkahler 3y agoIf you have a 4 bit signal multiplied by a 4 bit weigh, you can just use a lookup table for any operation with any output type you want. A 256 entry table of 32bit floats fits in 1K for example, which will fit in cache with plenty left over. This is independent of the quantization scheme so long as it's not dynamic.
- ajtulloch 3y agoIt's quite unfavorable on modern hardware. A Sapphire Rapids core can do 2 separate 32 half-precision FMAs (vfmadd132ph, [1]) per clock, which is 128 FLOPs/cycle. It is not possible to achieve that kind of throughput with an 8-bit LUT and accumulation, even just a shuffle with vpshufb is too slow. [1]: https://www.intel.com/content/www/us/en/docs/intrinsics-guide/index.html#text=_mm512_fmadd_ph&ig_expand=3117,3117 https://www.intel.com/content/www/us/en/docs/intrinsics-guid...
- anonymoushn 3y agoThat's absolutely wild. If you really only needed vpshufb, the throughput is the same in terms of values, because there are twice as many values per register and you get to retire half as many instructions, but it takes a bunch more instructions to combine the two inputs and apply a LUT of 256 values :(
- phkahler 3y ago>> It's quite unfavorable on modern hardware. Fair point. It might help if the system is DRAM bandwidth limited, so reducing the data size helps even though individual operations take multiple instructions. But that is not the situation with todays hardware.
- mattsan 3y agoHmmm is there anything preventing dedicated AI chips to have this LUT built in and to vectorise it too?