4 ms·
Is this what is used in e.g. ggml/llama.cpp which is the most prominent quantized inference program I'm aware of? If not, would it result in a material differen
by version_five 3y ago
Is this what is used in e.g. ggml/llama.cpp which is the most prominent quantized inference program I'm aware of? If not, would it result in a material difference vs whatever it's using now (which I thought was 4-8 bit float numbers). Anybody know?
- brucethemoose2 3y agoIn the wild, people tend to use GTPQ quantization for pure GPU inference: https://github.com/PanQiWei/AutoGPTQ https://github.com/PanQiWei/AutoGPTQ And ggml's quant for CPU inference with some offload, which just got updated to a more GPTQ-like method days ago: https://github.com/ggerganov/llama.cpp/pull/1684 https://github.com/ggerganov/llama.cpp/pull/1684 Some other runtimes like Apache TVM also have their own quant implementations: https://github.com/mlc-ai/mlc-llm https://github.com/mlc-ai/mlc-llm For training, 4-bit bitsandbytes is SOTA, as far as I know: https://huggingface.co/blog/4bit-transformers-bitsandbytes https://huggingface.co/blog/4bit-transformers-bitsandbytes TBH I'm not sure why this November paper on 8 bit is being linked now. Few are running 8 bit model inference when they could fit a better 3-5 bit model in the same memory pool.