Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wskwon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Kimi K3 on vLLM: Up to 370 Tokens/sec
(vllm.ai)
7 points
by
wskwon
2mo ago
|
0 comments
2.
▲
by
wskwon
3y ago
Thanks! Please try it out and share any feedback you might have.
3.
▲
by
wskwon
3y ago
Thanks for the explanation! I believe the two ideas are basically orthogonal. FlashAttention reduces memory read/writes, while PagedAttention reduces memory waste.
4.
▲
by
wskwon
3y ago
Yes, vLLM focuses on maximizing throughput when the VRAM is fully utilized. Nevertheless, I believe users can still benefit from vLLM even if they don't utilize the memory to its full capacity, because vLLM also includes other optimiza
5.
▲
by
wskwon
3y ago
We used matplotlib for the performance charts, and used a free website to convert google slides to the animation gifs.
6.
▲
by
wskwon
3y ago
Not really. vLLM optimizes the throughput of your LLM, but does not reduce the minimum required amount of resource to run your model.
7.
▲
by
wskwon
3y ago
You can think of LMSYS Vicuna: https://chat.lmsys.org as our hosted demo, as it actually uses vLLM as the backend.
8.
▲
vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention
(vllm.ai)
295 points
by
wskwon
3y ago
|
42 comments
9.
▲
by
wskwon
3y ago
vLLM has been adopted by LMSYS for serving Vicuna and Chatbot Arena.