3 ms·
What about cache? When you change the context the prefill stage will be much slower?
by urvader 2mo ago
What about cache? When you change the context the prefill stage will be much slower?
- kruxigt 2mo ago[dead]
- esperent 2mo agoThat's always going to be a trade off with anything like this so I guess it's better to think of it as an alternative to compaction. Another use case that comes to mind is that sometimes I'll include some detail early in a conversation and I mean it as incidentals information but the AI fixates on it. If I could selectively edit that out rather than start a whole new conversation it would be worth the cache miss.
- chatchan 2mo agoYes, that is exactly the failure mode I care about and drove me to develop ThoughtDAG! In ThoughtDAG, removing that edge excludes the detail from the next request without deleting the original branch. Thinking of this as user-directed compaction is a useful framing.
- chatchan 2mo agoI have not noticed a measurable slowdown in practice so far, including canvases with around a hundred nodes. A request only includes the wired ancestors of the current node, not the entire canvas, so node count alone is not a good measure of prefill cost. That said, your concern is valid for very long contexts. Editing an early ancestor may reduce prefix-cache reuse, while pruning a branch also makes the resulting prompt shorter. ThoughtDAG does not manage its own KV cache today, so this is something I need to benchmark properly rather than claim is solved. Have you encountered this mainly with local models or hosted APIs?
- olejorgenb 2mo agoI think many providers store the cache such that you can reuse any prefix, so it might not be as bad as you naively expect.