3 ms·
Nope, you gotta understand. This is the move. It is appended at the end of every single message. It is never saved in the conversation that get sent back for in
by bitexploder 1mo ago
Nope, you gotta understand. This is the move. It is appended at the end of every single message. It is never saved in the conversation that get sent back for inference. So you send it. But when you go back for inference, it’s at the top of the stack so all of the cashing works you’re not pre-filling every time or anything like that. It burns plus N tokens, where N is my prompt stack. It is not really that expensive. I have measured it to within an inch of it its life. Think of it this way every bit of the prefix and the conversation stays exactly the same you’re only adding to the very end of the conversation. So after the first turn, it is basically always cashing within the KV cache for a given context. Sorry I am using voice dictation. My hands are tired this week. Basically you sculpt the conversation history to ensure prefix caching
- bluegatty 1mo agoYes - with arbitrary prefix caching that might work. That's tricky though, not everyone is going to provide that. Did you have to build your own harness for this? Or hack Claude Code or something?
- bitexploder 1mo agoI don’t use their CLIs. OMP/OpenCode. If you are stuck on them, you can do it in a proxy layer.
- thatguymike 1mo agoIt should burn N + len(answer), because you have to re-cache the whole answer without the prompt stack. Perhaps more persnickety, it pushes the LLM out of distribution - if it’s unnatural for it to write in plain language without the prompt stack, your prefix will be an unnatural conversation which can reduce intelligence in hard to measure ways, especially over long conversations. Not saying don’t do it, clarity is perhaps worth the intelligence hit, but it’s not going to be a free lunch.
- bitexploder 1mo agoYes, maybe. Evidence around caveman shows this isn’t a big deal for token consumption (forcing it to respond in a way it was not tuned) and I don’t see a big difference either way. And maybe it increases intelligence in hard to measure ways. Lots of parameters in these models. I feel I fight them less with this setup. They get so lost in their own invented bullshit they stop being useful pretty often without it. So I would take bets on that :)