3 ms·
I'd guess the GPTQ-for-LLaMa repo is using a larger context size. Poking around it looks like GPTQ-for-llama is specifying 2048 [1] vs the default 512 for llama
by bakkoting 4y ago
I'd guess the GPTQ-for-LLaMa repo is using a larger context size. Poking around it looks like GPTQ-for-llama is specifying 2048 [1] vs the default 512 for llama.cpp [2]. You can just specify a longer size on the CLI for llama.cpp if you are OK with the extra memory.
[1] https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/934034c8ef4024e52a3c0c5893d86db84f1b52b6/llama.py#L20 https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/934034c8e...
[2] https://github.com/ggerganov/llama.cpp/tree/3525899277d2e2bdc8ec3f0e6e40c47251608700#latest-measurements https://github.com/ggerganov/llama.cpp/tree/3525899277d2e2bd...