Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aSidorenkoCode
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
aSidorenkoCode
7mo ago
Good benchmark results don't mean identical outputs. The task completion rate is the same: both pass the same exercises. The paths the model takes differ, but the end result is the same -> pass the tests The full benchmarking method
2.
▲
by
aSidorenkoCode
7mo ago
the benchmarks show no degradation in task completion with the shorter descriptions. We're in the age where frontier LLMs don't need instructions on how to read or edit a file. The descriptions aren't dynamically summarized
3.
▲
by
aSidorenkoCode
7mo ago
Every API call sends the full tool schema for all available tools. In a 10-20 step session, you're paying for the same verbose descriptions over and over. Models don't need a paragraph-long explanation of read on the 15th call. Th
4.
▲
Show HN: OpenSlimedit – Cut AI coding token usage by 21-45% with zero config
(github.com)
2 points
by
aSidorenkoCode
7mo ago
|
8 comments
5.
▲
by
aSidorenkoCode
7mo ago
I made a benchmark on the top tier models as this is not the case in the article. Also I did several cases and the result is self speaking. Hashline is not an improvement that speaks for itself. It is the overhead of harness itself in tool