4 ms·
We do both: We compress tool outputs at each step, so the cache isn't broken during the run. Once we hit the 85% context-window limit, we preemptively trigger
by thebeas 7mo ago
We do both:
We compress tool outputs at each step, so the cache isn't broken during the run. Once we hit the 85% context-window limit, we preemptively trigger a summarization step and load that when the context-window fills up.
- esperent 7mo ago> we preemptively trigger a summarization step and load that when the context-window fills up. How does this differ from auto compact? Also, how do you prove that yours is better than using auto compact?
- ivzak 7mo agoFor auto-compact, we do essentially the same Anthropic does, but at 85% filled context window. Then, when the window is 100% filled, we pull this precompaction + append accumulated 15%. This allows to run compaction instantly