3 ms·
This really depends on what GPUs you use. If you GPUs has very small amount of memory, vLLM will help more. vLLM addresses the memory bottleneck for saving KV
by zhisbug 3y ago
This really depends on what GPUs you use. If you GPUs has very small amount of memory, vLLM will help more.
vLLM addresses the memory bottleneck for saving KV caches and hence increases the throughput.