3 ms·
How do you use the cache efficiently?
by oblio 2mo ago
How do you use the cache efficiently?
- eulo__ 2mo agoI’m currently running /compact “explain the core next step” whenever over 20% context. Also doing /clear with a md file handover if I think the next input is diverse enough from the previous work. You get a lot more out of it. Not sure if this is best practice though.
- zarzavat 2mo agoYou have to keep your session warm in cache. Keep the AI talking/thinking. If you have not touched a session for few minutes then /clear and start a new session. Providers will generally keep your session in cache for at least 5 minutes, possibly hours. The exact cache policy depends on the provider. If your session expires from cache then the next time you send a message you will have to pay for all the tokens you had used in context up until that point again. e.g. if you have 200k tokens in context then if your session goes cold and you send a message after expiry you will have to pay for those 200k tokens again. With 1M contexts especially you have to be extremely careful that you don't end up resubmitting requests for hundreds of thousands of tokens again and again. Try to get yourself and the model to use disk for medium-term context rather than model context, that way it's much easier to /clear and restart if you need to go to the bathroom or something.