2 ms·
Can someone explain the math to me? Why is 1-bit only ten percent less memory than 2-bit?
by c7b 4mo ago
Can someone explain the math to me? Why is 1-bit only ten percent less memory than 2-bit?
- incognito124 4mo agoKeyword dynamic, the parameters are quantized on a case by case basis
- idonotknowwhy 4mo ago2 reasons. First, it's not really "1 bit", actually much closer to 2-bit. IQ1_M is actually 1.75bit and IQ2_XXS is 2.06bit This is from the ./llama-quantize --help with most of the quant types and their size in bpw: https://pastebin.com/bCUqGfeE https://pastebin.com/bCUqGfeE And to elaborate on the "dynamic" aspect inconito said in the other comment, if you click on one of the .gguf files in huggingface: https://huggingface.co/unsloth/GLM-5.2-GGUF/blob/main/UD-IQ1_M/GLM-5.2-UD-IQ1_M-00002-of-00006.gguf https://huggingface.co/unsloth/GLM-5.2-GGUF/blob/main/UD-IQ1... There are a lot of Q5_K, Q6_K, etc tensors. Only the routed experts (ffn_gate_exps.weight, ffn_up_exps.weight, ffn_down_exps.weight) are heavily quantized, and it looks like the down_proj is actually iq3_xxs for this model.