3 ms·
Show HN: Proxima serves 4x more requests with no hardware change on vLLM
hey everyone, i decided to make a vLLM plugin that implements the Star-KV paper. the results are quiet promising with a decode kernel thats faster than FA2 in higher batch sizes. would love any thoughts and recommendations
- pythongiant 2mo agoTenosra is all about making AI more efficient and resourceful you can read more abour it at: https://www.tenosra.com/proxima https://www.tenosra.com/proxima