3 ms·
> If the model was designed specifically to quantize down to 1.58b, then it's different. > AFAIK, there's no large models designed for this yet. Isn't BitNet
by richardfey 2mo ago
> If the model was designed specifically to quantize down to 1.58b, then it's different.
> AFAIK, there's no large models designed for this yet.
Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)
- onlyrealcuzzo 2mo ago2B is pretty small... No 100B+ param (certainly no 2T+ param) models have been trained natively to quantize down to 1.58b.