3 ms·
> I would never want to use something like ollama in a production setting. We benchmarked vLLM and Ollama on both startup time and tokens per seconds. Ollama c
by steren 1y ago
> I would never want to use something like ollama in a production setting.
We benchmarked vLLM and Ollama on both startup time and tokens per seconds. Ollama comes at the top. We hope to be able to publish these results soon.
- ekianjo 1y agoyou need to benchmark against llama.cpp as well.
- apitman 1y agoDid you test multi-user cases?
- jasonjmcghee 1y agoAssuming this is equivalent to parallel sessions, I would hope so, this is like the entire point of vLLM
- sbinnee 1y agovllm and ollama assume different settings and hardware. Vllm backed by the paged attention expect a lot of requests from multiple users whereas ollama is usually for single user on a local machine.