4 ms·
You also want prompt/prefix cache. Otherwise there is a lot of duplicate prefill processing if you fork conversation, have a different chat window or anything l
by nicce 27d ago
You also want prompt/prefix cache. Otherwise there is a lot of duplicate prefill processing if you fork conversation, have a different chat window or anything like that. It therefore also makes subagents much faster.
- latentsea 26d agoI'm getting 99% cache hit in DeepSeek Harness?