4 ms·
Some results here: https://github.com/ggerganov/llama.cpp/discussions/406 https://github.com/ggerganov/llama.cpp/discussions/406 tl;dr quantizing the 13B model
by bakkoting 4y ago
Some results here: https://github.com/ggerganov/llama.cpp/discussions/406 https://github.com/ggerganov/llama.cpp/discussions/406
tl;dr quantizing the 13B model gives up about 30% of the improvement you get from moving from 7B to 13B - so quantized 13B is still much better than unquantized 7B. Similar results for the larger models.
- terafo 4y agoI wonder where such difference between llama.cpp and [1] repo comes from. F16 difference in perplexity is .3 on 7B model, which is not insignificant. ggml quirks are definitely need to be fixed. [1] https://github.com/qwopqwop200/GPTQ-for-LLaMa https://github.com/qwopqwop200/GPTQ-for-LLaMa
- bakkoting 4y agoI'd guess the GPTQ-for-LLaMa repo is using a larger context size. Poking around it looks like GPTQ-for-llama is specifying 2048 [1] vs the default 512 for llama.cpp [2]. You can just specify a longer size on the CLI for llama.cpp if you are OK with the extra memory. [1] https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/934034c8ef4024e52a3c0c5893d86db84f1b52b6/llama.py#L20 https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/934034c8e... [2] https://github.com/ggerganov/llama.cpp/tree/3525899277d2e2bdc8ec3f0e6e40c47251608700#latest-measurements https://github.com/ggerganov/llama.cpp/tree/3525899277d2e2bd...
- gliptic 4y agoGPTQ-for-LLaMa recently implemented some quantization tricks suggested by the GPTQ authors that improved 7B especially. Maybe llama.cpp hasn't been evaluated with those in place?