3 ms·
I checked out the development benchmarks for agents, and was suprized to find that they're not really representative of many of my own day-to-day development wo
by pistoriusp 2mo ago
I checked out the development benchmarks for agents, and was suprized to find that they're not really representative of many of my own day-to-day development workflows.
It appears that they're mostly testing the ability to make business tasks autonomous, with ~20% associated to development tasks (ssh here, install this, etc.), but not actual programming.
- andai 2mo agoThat's probably because the actual coding benchmarks were saturated several years ago.
- pistoriusp 2mo agoWhich is why you should perform your own benchmarks against your own software stack.
- andai 2mo agoAgree. To this I would add, many things are saturated even for smaller models, which tend to be much cheaper and faster. On many of my tests, there was no difference in the result between the smaller and bigger model, but there was a big difference in speed and price.