3 ms·
the benchmarks show no degradation in task completion with the shorter descriptions. We're in the age where frontier LLMs don't need instructions on how to rea
by aSidorenkoCode 8mo ago
the benchmarks show no degradation in task completion with the shorter descriptions. We're in the age where frontier LLMs don't need instructions on how to read or edit a file.
The descriptions aren't dynamically summarized either. They're static in the plugin, same every call, every session. Zero overhead, fully deterministic.
This has been validated in over 3000 benchmark runs in OpenCode and I ran the entire Exercism Python practice suite (https://github.com/exercism/python/tree/main/exercises/practice https://github.com/exercism/python/tree/main/exercises/pract...) with and without the plugin with identical results. An initial dataset is shared in the repo.
- verdverm 8mo agoHave you made that benchmarking process open so others could reproduce it? > with identical results If your results are identical, you should be very sus, something is wrong if this is true. Nothing in agentic is reliable of fully deterministic
- aSidorenkoCode 8mo agoGood benchmark results don't mean identical outputs. The task completion rate is the same: both pass the same exercises. The paths the model takes differ, but the end result is the same -> pass the tests The full benchmarking methodology and tooling will be published alongside the paper.
- verdverm 8mo agoyou used the word "identical" to describe it, not me words matter which is why I still think this is a terrible idea, I don't think it holds up in the general case and would, as a peer reviewer, be inclined to believe there is benchmark filtering that makes for good results. You should use the same benchmarks everyone else is when you write your paper
- verdverm 8mo agotimely result, https://arxiv.org/pdf/2602.12670 https://arxiv.org/pdf/2602.12670