4 ms·
I don't get it. Why benchmark the latency instead of recall/precision? Optimizing for millisecond-level latency is meaningless in the context of LLM calls. Accu
by langs 28d ago
I don't get it. Why benchmark the latency instead of recall/precision? Optimizing for millisecond-level latency is meaningless in the context of LLM calls. Accuracy is the tool's greatest value, yet there is no testing for it?
- esafak 28d agoHas anyone else benchmarked all these tools for precision/recall? I too want to know if agent memory is something I should add. I only do session memory for now and that is quite useful.
- glub 27d agoYes. See LongMemEval, LoCoMo. Tons of research here. But precision/recall is relatively "solved". What nobody has gotten close to solving is maintenance and provenance - what goes into memory, what qualifies as truth, how stale memory gets invalidated/superseded. We're now in the phase of re-discovering 30+ years of pain of knowledgebases.
- olenzma 26d agoInvalidation is not an algorithm, but a judgment made in context
- nedomolkovivan 26d ago[flagged]
- entity002 26d ago[dead]
- okf_memory 21d ago[flagged]
- vshulcz 27d agoBenched cold start recently: 19k LongMemEval sessions laid down in the real ~/.claude and ~/.codex layouts, 100 questions whose answer sits in exactly one session, scored by whether that session comes back (bias: I built deja, one of the six) deja: 29s to index, 24ms query, 18/100 hit@1, 67 found@50. Plain BM25, no vectors agentmemory: 95s import, 14 hit@1, 65 found@50, plus a worker and engine on four ports MemPalace: ~3h mining, 2.6s query, 14 hit@1 CASS: 56m index, every NL query fails with "query fuel exhausted" on the release build (fixed on their main) claude-mem: no-op out of the box, only records forward from install funes: the documented 1 min first pass indexed 189 of 19k sessions (0/100); full index still embedding, ~3 sessions/s Numbers look low because 19k sessions is brutal; on the standard 500-session LongMemEval-S the same BM25 gets ~85% hit@1 The funny thing is BM25 basically ties embeddings here at 1/100th the cost. The real cliff is reranking (found@50 67 vs hit@5 35) and staleness. Vector search has zero concept of "superseded info" only fix I found was letting explicit user corrections outrank the transcript. Repro scripts and corpus: https://vshulcz.github.io/deja-vu/guide/day-zero.html https://vshulcz.github.io/deja-vu/guide/day-zero.html @skeledrew: cross-agent across 23 harnesses, but yeah, it's an index over logs, not a source of truth :)
- opwizardx 27d agoI did similar tests on my own corpus when was considering whether to keep semantic search in default path for pond. On 3 months of my own sessions I’ve seen that BM25 search was finding the correct answer in ~61%, where semantic had shown only ~37% of success. After that it was easy for me to make the decision. Got all info on how I did evals in here, if interested: https://github.com/tenequm/pond/tree/main/docs/researches/2608-21-semantic-vs-fts-usage-eval https://github.com/tenequm/pond/tree/main/docs/researches/26...