4 ms·
Fortunately llama.cpp works well with 4 concurrent requests now. This is the default, and you can increase or decrease it with -np N
by idonotknowwhy 2mo ago
Fortunately llama.cpp works well with 4 concurrent requests now. This is the default, and you can increase or decrease it with -np N