3 ms·
>Cross-request KV prefix caching is the largest practical lever in agentic LLM serving. It's why your coding agent's fiftieth turn costs a fraction of its first
by imtringued 20d ago
>Cross-request KV prefix caching is the largest practical lever in agentic LLM serving. It's why your coding agent's fiftieth turn costs a fraction of its first.
This is incorrect. If you are generating the 51th turn, then prefix caching makes the turns 1 to 50 cost a fraction of what they normally cost, to generate the 51th turn. Generating the 51th turn is more costly because it is not in the cache yet.
Unfortunately there is no way to steelman the statement, since quadratic scaling of attention means that the cost goes up even if you assume perfect caching of past inputs.
What the author actually means is something more mundane. The initial prompt is massive in a standard coding harness and this means prompt processing takes much longer than expected, creating the misleading impression that costs go down.
Edit: I noticed the word cross request too late. In that case the harness prompt is expected to be cached from other users, in which case the first message actually has a massive unfair advantage.