Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
SmithersBot
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
46 ms
·
1.
▲
by
SmithersBot
10d ago
We used Claude Code agents running Opus 5 with each of the 12 memory systems. Each agent got the same work across multiple sessions and 150 questions. Each memory tool runs through weeks of simulated work. Each agent chooses what to store a
2.
▲
A Markdown wiki outscored every AI agent memory product we benchmarked
(verginglabs.com)
2 points
by
SmithersBot
2mo ago
|
0 comments
3.
▲
Perplexity came dead last after testing agentic search tools 3,537 times
(agenticresourceradar.com)
2 points
by
SmithersBot
3mo ago
|
1 comments
4.
▲
by
SmithersBot
3mo ago
I just launched the Agentic Search Index, a public benchmark of the web-search tools AI agents use. Perplexity Sonar came last of the nine tools tested. I ran the same Claude agent on 121 web-search tasks through every tool, three times for
5.
▲
Show HN: My open source agent built and launched its own business in 48 hours
(github.com)
1 points
by
SmithersBot
4mo ago
|
0 comments
6.
▲
by
SmithersBot
4mo ago
I built an agent that pursues your goals over weeks or months until they're achieved: https://github.com/smithersbot/smithersbot
7.
▲
by
SmithersBot
4mo ago
Doing it in the same session does save a ton of tokens but I find it's too biased towards its own implementation even if you tell it to use "fresh eyes" or to "act like a code reviewer in a bad mood." Including thos
8.
▲
by
SmithersBot
4mo ago
Opus work so well for now... until they quantize next week...
9.
▲
by
SmithersBot
4mo ago
as long as OpenAI and Anthropic keep subsidizing dirt cheap Codex or Claude Code usage, I'll just keep using them as evaluators. The trick is to have a fresh instance doing the reviewing, not the one that did the work.