3 ms·
I think it's a vllm vs llama_cpp performance thing, will pay more into it. One note I had between the two is that gemma has a much higher prefix cache hit rate
by trouve_search 2mo ago
I think it's a vllm vs llama_cpp performance thing, will pay more into it.
One note I had between the two is that gemma has a much higher prefix cache hit rate in general.