3 ms·
Looks great! Are there other benchmarks? How does the speed compare to other LLM engines like llama.cpp / vllm (on GPUs)? Is it able to do continuous batching o
by nailk 3y ago
Looks great!
Are there other benchmarks? How does the speed compare to other LLM engines like llama.cpp / vllm (on GPUs)?
Is it able to do continuous batching of incoming requests like vllm?