3 ms·
Benched cold start recently: 19k LongMemEval sessions laid down in the real ~/.claude and ~/.codex layouts, 100 questions whose answer sits in exactly one sessi
by vshulcz 26d ago
Benched cold start recently: 19k LongMemEval sessions laid down in the real ~/.claude and ~/.codex layouts, 100 questions whose answer sits in exactly one session, scored by whether that session comes back (bias: I built deja, one of the six)
deja: 29s to index, 24ms query, 18/100 hit@1, 67 found@50. Plain BM25, no vectors
agentmemory: 95s import, 14 hit@1, 65 found@50, plus a worker and engine on four ports
MemPalace: ~3h mining, 2.6s query, 14 hit@1
CASS: 56m index, every NL query fails with "query fuel exhausted" on the release build (fixed on their main)
claude-mem: no-op out of the box, only records forward from install
funes: the documented 1 min first pass indexed 189 of 19k sessions (0/100); full index still embedding, ~3 sessions/s
Numbers look low because 19k sessions is brutal; on the standard 500-session LongMemEval-S the same BM25 gets ~85% hit@1
The funny thing is BM25 basically ties embeddings here at 1/100th the cost. The real cliff is reranking (found@50 67 vs hit@5 35) and staleness. Vector search has zero concept of "superseded info" only fix I found was letting explicit user corrections outrank the transcript.
Repro scripts and corpus: https://vshulcz.github.io/deja-vu/guide/day-zero.html https://vshulcz.github.io/deja-vu/guide/day-zero.html
@skeledrew: cross-agent across 23 harnesses, but yeah, it's an index over logs, not a source of truth :)
- opwizardx 26d agoI did similar tests on my own corpus when was considering whether to keep semantic search in default path for pond. On 3 months of my own sessions I’ve seen that BM25 search was finding the correct answer in ~61%, where semantic had shown only ~37% of success. After that it was easy for me to make the decision. Got all info on how I did evals in here, if interested: https://github.com/tenequm/pond/tree/main/docs/researches/2608-21-semantic-vs-fts-usage-eval https://github.com/tenequm/pond/tree/main/docs/researches/26...