4 ms·
It divides the context into smaller "slots", so it can process requests concurrently with continuous batching. See also: https://github.com/ggerganov/llama.cpp/
by mcharytoniuk 2y ago
It divides the context into smaller "slots", so it can process requests concurrently with continuous batching. See also: https://github.com/ggerganov/llama.cpp/tree/master/examples/server https://github.com/ggerganov/llama.cpp/tree/master/examples/...