3 ms·
With 1.58-bit ternary quantization, you may think you're running a big model but really you're just running a "mini" version of it
by lazerlapin 7mo ago
With 1.58-bit ternary quantization, you may think you're running a big model but really you're just running a "mini" version of it