4 ms·
Cartridges: Storing long contexts in tiny caches with self-study
- dvrp 1y agoFrom their repo: tl;dr When we put lots of text (e.g. a whole code repo) into a language model's context, generation cost soars because of the KV cache's size. What if we trained a smaller KV cache for our documents offline? Using a test-time training recipe called self-study, we show that this simple idea can improve throughput by 26× while maintaining quality. Link to their blog post: https://hazyresearch.stanford.edu/blog/2025-06-08-cartridges https://hazyresearch.stanford.edu/blog/2025-06-08-cartridges