4 ms·
Can anyone explain why the prefix cache is tied to effort? I frequently run Fable at xhigh effort to run statistical modeling way above my undergraduate unders
by jnwatson 2mo ago
Can anyone explain why the prefix cache is tied to effort?
I frequently run Fable at xhigh effort to run statistical modeling way above my undergraduate understanding. Claude Fable produces Masters-degree level output, and then I spend lots of round trips asking it to explain different parts to me.
The first part absolutely uses the extra effort, but the interrogation exercise is something a much simpler model, or the same model with much less effort, could answer.
- hellohello2 2mo agoOne trick is to simply as for a fast answer when talking to a high effort model, when working interactively. Sounds stupid but I do this all the time and it works. Just tell it you are working interactively now and need ultrafast answers with no thinking. I would be really curious to know as well, why effort is linked to cache as its quite inconveniant. Is it possible the token used to indicate effort is only passed once at the start, not per thinking trace, or quite simply that different efforts have different model weights?
- janalsncm 2mo agoI’m guessing that there’s a system prompt at the top telling the model about its reasoning budget. So when you switch reasoning effort it busts the cache.
- foota 2mo agoHmmm... Why wouldn't this be handled like other end of prompt things like the current mode?
- janalsncm 2mo agoPrompt caching only works based on the prefix. Let’s call your output Y and the low reasoning prompt A and a medium reasoning prompt B. Previously you were at A+Y. Switching to medium reasoning makes it B+Y. There’s no prefix which can be cached, so the entire B+Y needs to be reprocessed.
- foota 1mo agoWhy can't effort just be at the end of the prompt? My understanding is that that is how modes and other "dynamic updates" like the time are handled.
- SoMomentary 2mo agoThat makes sense to me. The output styles work the same way.
- cellularmitosis 2mo agoMaybe switching effort routes you to a different rack of gpu’s which don’t have the cache
- janalsncm 2mo agoThe KV cache is probably offloaded to RAM after a generation is complete. Then it is pulled back into any rack in the data center that has the model you’re using.