4 ms·
Our team has built this open source project, LMCache, to reduce repetitive computation in LLM inference and make systems serve more people (3x more throughput i
by lihanc111 1y ago
Our team has built this open source project, LMCache, to reduce repetitive computation in LLM inference and make systems serve more people (3x more throughput in chat applications) and it has been used in IBM's open source LLM inference stack.
In LLM serving, the input is computed into intermediate states called KV cache to further provide answers. These data are relatively large (~1-2GB for long context) and are often evicted when GPU memory is not enough. In these cases, when users ask a follow up question, the software needs to recompute for the same KV Cache. LMCache is designed to combat that by efficiently offloading and loading these KV cache to and from DRAM and disk.
Ask us anything!
- dist-epoch 1y agoHow is it possible to do non-prefix KV cache? I was under the impression that the V for one token potentially depends on the V of all previous ones.
- da-x 1y agoYes, there's KV cache 'Blending' see [1]. Future versions of LMCache are aiming to support this. [1] CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion- https://arxiv.org/abs/2405.16444 https://arxiv.org/abs/2405.16444
- pama 1y agoIs your aim targetting the inference at scale or specialized/new/simpler inference pipelines? Sglang and vllm have disaggregated prefix and decoding serving (eg https://docs.vllm.ai/examples/online_serving/disaggregated_serving.html https://docs.vllm.ai/examples/online_serving/disaggregated_s... or https://github.com/sgl-project/sglang/issues/3554 https://github.com/sgl-project/sglang/issues/3554 and https://github.com/sgl-project/sglang/issues/4655 https://github.com/sgl-project/sglang/issues/4655) — could your solution enable a model-agnostic cache store/server or is that orthogonal to what you are trying to achieve?
- nativeit 1y agoHas it been used in IBM's inference stack, or used with IBM's inference stack? In other words, has this been merged into IBM's own repositories, or has someone just tested it using them?
- lihanc111 1y agoIt is in IBM's llm-d open source stack
- behnamoh 1y ago> Our team So this is something that might in the future turning to a commercial product? something like Langchain and thousands of open source projects that started as "open source" but then ended up implementing proprietary features for a cost.
- Tokumei-no-hito 1y agoi don't see anything wrong with that approach, do you?
- behnamoh 1y agoGive it time and you'll come to my conclusion.