4 ms·
Apply video compression on KV cache to 10,000x less error at Q4 quant
- mungoman2 6mo agoThis is cool. It makes storage of the KV cache much smaller, making it possible to keep more of it in fast memory. Bandwidth-wise it is worse (more bytes accessed) to generate and do random recall on than the vanilla approach, and significantly worse than a quantized approach. That’s because the reference needs to be accessed. I guess implied is that since the KV cache is smaller, the probability is higher that the parts it that are needed are in fast memory, and that bandwidth requirements of slow links is reduced, and performance goes up. Would be interesting with a discussion about benefits/drawbacks of the approach. Ideally backed by data.
- Reubend 6mo agoThis is really cool research, but I'm wondering how much it slows down inference. The readme says that it's "...distinguished by zero overhead (no learned components, no entropy coding)" but does that really mean that this is a "free win"?
- mike_hearn 6mo agoNice, although perhaps slightly academic given that good KV cache compression algorithms already exist. Probably the frontier labs were using them for a long time already. Nice to have it in llama.cpp though. I'm curious who "we" refers to. I can't see any authorship information or a paper and this is the user's only repository. Maybe it doesn't need one. Also interesting that it was developed and tested on AMD hardware. The main utility of this beyond just saving money for model servers would be deliberately prefilling very long contexts and then saving them to fast flash so you can then later quickly load and query them. I think only Anthropic's API would give enough control to do this today, maybe Google's, OpenAI's makes caching fully implicit. Like one or two prompts per codebase or something like that, so you can then query the entire codebase in parallel with questions without needing grepping or RAG. Modern serving pipelines all use disaggregated prefill as far as I know so there are inter-machine transfers anyway, and it directly saves on GPU cost.
- tveita 6mo ago"video compression" by analogy only, what this claims to actually do is delta encode the values in each token from the previous token. Interesting idea, but the results seem almost suspicious? even accounting for the extra bits used to store the 16-bit start value for each block - ~5% for k=64 The code does funky things, like the encoder updates the reference value for each encoded token, using the non-quantized value! [1] But the decoder just ignored all that. [2] how can this work? [1] https://github.com/cenconq25/delta-compress-llm/commit/f185ffcd4cf5ee8041a95d93008bbcc0914d04e4#diff-7974ac143ef46eaf6e413b2aa0aa7bfe1e81958925597f6b922f14886ee53883R111 https://github.com/cenconq25/delta-compress-llm/commit/f185f... [2] https://github.com/cenconq25/delta-compress-llm/commit/f185ffcd4cf5ee8041a95d93008bbcc0914d04e4#diff-7974ac143ef46eaf6e413b2aa0aa7bfe1e81958925597f6b922f14886ee53883R160 https://github.com/cenconq25/delta-compress-llm/commit/f185f...
- free_bip 6mo agoIt looks largely LLM-generated. I'm somewhat skeptical about the results as well, hopefully someone can independently confirm that the technique works.
- cenconq25 6mo agoHey everyone, I’m the owner of the repo. This started as one of my daily crazy ideas for optimizing LLMs using my software engineering background and I’m just trying to think outside the box when it comes to LLM optimization, even though I’m not an AI engineer or researcher. There’s no academic paper behind this, it’s really just me coming up with the idea and Claude helping prove it out. And tveita is right, there is a real bug, although it doesn’t affect the published benchmark results, I will get that fixed in the next release.
- naasking 6mo agoInteresting idea, but I hope people just start switching to ParoQuant and eliminate basically all quantization errors relative to fp16/bf16 even going down to 4-bits: https://github.com/z-lab/paroquant https://github.com/z-lab/paroquant