4 ms·
Computer Anthology: A continuously evolving benchmark family for AI agents
- Betaantunes11 2mo agoThe methodology was the most interesting part for me. The paper spends as much time explaining how the benchmark was built as the benchmark itself.
- Luvison_Rafael 2mo agonice!
- mrhectograma 2mo agoRefreshing to see something practical instead of another leaderboard battle. Also, props to the team for being so meticulous.
- kmiens 2mo agoThe comparison between harnesses is very nice. Interesting to see that using a different harness can bump the performance of the model as much as a new version (e.g., GPT 5.5+Codex ~= GPT 5.6+Terminus, at lower cost)
- dvaplima 2mo agoThat’s a nice discussion. Some people say that with current model capabilities, the real differentiator is the harness. What are the best harnesses you guys are using?
- dvaplima 2mo ago[flagged]
- alinebindel 2mo agothorough work, good stuff.. it even runs a selection-bias analysis against their own benchmark and reports that some tasks that were disproportionately hard for a model. Rare to see a benchmark paper attack itself like that.
- ltononro 2mo ago[flagged]
- pedroaugusto-me 2mo agoThis approach of not only producing the benchmark tasks, but also focusing on creating a data engine that will improve over time and produce up-to-date tasks that challenge the cutting-edge models is very interesting and valuable.
- aamdias 2mo agoUsing semantic perturbation to test whether difficulty survives rewording is really smart. Great work!
- gramulho 2mo ago[flagged]
- MstTK 2mo ago[flagged]
- abernat 2mo ago[dead]
- caue_gimenez 2mo ago[dead]
- VicElko 2mo agogood work!
- MatheusFelipe 2mo agoInteresting the idea of treating the benchmark as an evolving system rather than a static dataset.
- stefani16bit 2mo ago[dead]
- lorenaaraujo 2mo agoImpressive!
- GIULIAFEDERIGHI 2mo ago[flagged]