3 ms·
Standard models are designed to quantize down to 4-bits relatively well. Anything below that, and especially 1.58b - is typically complete garbage, and you're
by onlyrealcuzzo 2mo ago
Standard models are designed to quantize down to 4-bits relatively well.
Anything below that, and especially 1.58b - is typically complete garbage, and you're much better off running a model 100x smaller at regular precision (compared to one 7x smaller quantized into complete garbage).
If the model was designed specifically to quantize down to 1.58b, then it's different.
AFAIK, there's no large models designed for this yet.
- richardfey 2mo ago> If the model was designed specifically to quantize down to 1.58b, then it's different. > AFAIK, there's no large models designed for this yet. Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)
- onlyrealcuzzo 2mo ago2B is pretty small... No 100B+ param (certainly no 2T+ param) models have been trained natively to quantize down to 1.58b.