4 ms·
It takes a 2/1.5bit model, groups parameters together then exploits a lack of entropy in the parameters to compress it a bit like text compression. It was only
by chessgecko 3y ago
It takes a 2/1.5bit model, groups parameters together then exploits a lack of entropy in the parameters to compress it a bit like text compression. It was only below 1bit for the ultra large model, guess the smaller ones weren’t quite as random.
It’ll be interesting to see if it works on the new mistral moe model, which is less sparse and probably trained more per param than these.