3 ms·
Not really. vLLM optimizes the throughput of your LLM, but does not reduce the minimum required amount of resource to run your model.
by wskwon 3y ago
Not really. vLLM optimizes the throughput of your LLM, but does not reduce the minimum required amount of resource to run your model.
- e12e 3y agoBut (in theory) - llama.cpp could implement similar approach to paging/memory and see a speedup for 4bit models on cpu?