3 ms·
UPDATE: I did more experiments - this time with llama-server and pi harness. i'll leave the results here for posteriority: ./llama.cpp/build/bin/llama-server
by wombat23 14d ago
UPDATE: I did more experiments - this time with llama-server and pi harness. i'll leave the results here for posteriority:
./llama.cpp/build/bin/llama-server \
-c "$context_size" \
-ctk q4_0 \
-ctv q4_0 \
--no-models-autoload \
--models-dir ~/bonsai \
--reasoning-preserve \
-ngl 99
the max context size i could serve is 64K on GPU only (the -ngl 99 setting). tested with pi harness and it is very fast. the /thinking level always gets reset to off though and it is not very smart like this. haven't figure out a way to fix that.