2 ms·
Thx, ‘similarity search mixing semantically related data with genuinely valuable data’ and about this ‘adding up during ingestion’ are exactly why we moved from
by Egeozin 6mo ago
Thx, ‘similarity search mixing semantically related data with genuinely valuable data’ and about this ‘adding up during ingestion’ are exactly why we moved from v1 to v2 for companion and conversational use cases. In this domain, scratchpad-like systems work well, and there’s usually no need to over-engineer retrieval.
I think v3 is categorically different. First, the LLM decides what matters, and we believe that scales better than having the engineer impose too much structure upfront and fail to create the right environment for the model, which was part of the limitation in v1. Second, this does not need to be irreversible if you support it with a simple harness, in our case, git and worktrees. V3 is also more applicable to companions or agents that require stronger problem-solving capabilities, such as coding.
We plan to publish our benchmarking results soon, so others can evaluate the approach for themselves.
- gkanellopoulos 6mo agoUnless I misread what is contained in the essay, the worktrees seem to be there to keep sessions separate. The agent doesn't use them to look at old memories while it runs. You even say the agent still can't reason about time in v3. When it updates memory, it only reads the current hot_context file. So if something goes wrong, an engineer can reset a bad merge. The agent can't do that itself (or you are planning to enable this capability?). That's still useful. But it's not the same as the agent looking at its own history. About "the LLM decides what matters": this makes sense for coding agents. They build files, so the file tree IS their work. For conversational memory I'm not so sure. You can get the same thing when you save the memory. Let the LLM pick its own labels, with no fixed list. Then you don't have to do that work every time you read. Looking forward to the v3 numbers.