4 ms·EfficientQAT: LLM Quantization, gets a 2-bit llama2-70B outperform regular 13B21 points by jackbravo 2y ago