Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
yohji1984
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Why My Open-Source Project Hasn't Done Better
4 points
by
yohji1984
2mo ago
|
1 comments
2.
▲
by
yohji1984
3mo ago
Sorry I will edit the post
3.
▲
Claude Code team should try macro so users can complete 3x as many tasks
1 points
by
yohji1984
3mo ago
|
2 comments
4.
▲
Is GPT-5.6 Sol Max Worth It?
2 points
by
yohji1984
3mo ago
|
4 comments
5.
▲
Why people chasing after useless token saving plugins and ignoring real solution
2 points
by
yohji1984
3mo ago
|
0 comments
6.
▲
Show HN: Agent use 83.1% fewer turns; 16.7 percentage points higher success
(turaai.net)
2 points
by
yohji1984
3mo ago
|
0 comments
7.
▲
That is why the tokensaving plugin and skill is completely useless
(turaai.net)
2 points
by
yohji1984
3mo ago
|
0 comments
8.
▲
Agent runtime reduces LLM turns by 80% with a higher success rate in DeepSWE
(github.com)
2 points
by
yohji1984
3mo ago
|
1 comments
9.
▲
by
yohji1984
3mo ago
Hi HN, I've been working on an agentic runtime framework. The earlier benchmark eval is promising. But I can see the limits of the test design and the fragility of the runtime itself. I would like to ask for your reviews of the framewo
10.
▲
by
yohji1984
4mo ago
Sorry, I missed the Open Items section. You're right about that, designing a good eval harness can be difficult and expensive. Maybe we need some kind of community project for agentic evals, where people can share eval harnesses and ru
11.
▲
by
yohji1984
4mo ago
I'm wondering why all these token-saving solutions focus their benchmarks exclusively on simple Q&A tasks. If their tools truly saved money in real, long-term programming tasks, they would have definitely published those benchmark