2 ms·
Thanks for the headstart, I saw hf also has PQ2_0 and able to finetune the command and in a MBA M4 24GB averages around 10t/s with the command. ./llama-prism-b
by ithkai92 14d ago
Thanks for the headstart, I saw hf also has PQ2_0 and able to finetune the command and in a MBA M4 24GB averages around 10t/s with the command.
./llama-prism-b10685-7dffb15/llama-server \
-m Ternary-Bonsai-2-27B-PQ2_0.gguf \
--port 8331 -ngl 99 -fa on -c 65536 --jinja \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0