4 ms·
This is quantized down to 2.5 bits per weight, whereas single precision accuracy is 32 bits.
by ftxbro 3y ago
This is quantized down to 2.5 bits per weight, whereas single precision accuracy is 32 bits.
- tutfbhuf 3y agoHow is 2.5 bits possible? Is it an average between 2 and 3 over all weights?
- ftxbro 3y ago> Is it an average between 2 and 3 over all weights? Yes I think it's an average where different quantization levels are used for different layers or weights. Here are more details about the quantization scheme: https://github.com/turboderp/exllamav2#exl2-quantization https://github.com/turboderp/exllamav2#exl2-quantization
- 3abiton 3y agoAny benchmarks on performance?
- redox99 3y agoThese models aren't even trained on FP32. The original format is FP16. And quantizing to INT8/FP8 is almost lossless. But yes, 2.5 bits per weight is pretty insane.
- version_five 3y agoThis should be in the headline (the 2.5 bit part), without the qualifier the result means nothing.