Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lunaprompts_hn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
lunaprompts_hn
7mo ago
The real-world benchmark approach is the right direction. Most agent evals I've seen test for task completion on clean inputs. That's not how production use looks. What tends to break agents in the wild: ambiguous instructions tha
2.
▲
by
lunaprompts_hn
7mo ago
The framing of "bad vs none" is interesting but I think the more useful question is: what makes an Agent.md actually good? From what I have seen, the ones that work well are specific about failure modes rather than capabilities. I
3.
▲
by
lunaprompts_hn
7mo ago
One thing that doesn't get discussed enough: the gap between people who can use AI tools effectively and those who can't is widening fast. The engineers who thrive are the ones who understand how the underlying models actually wor