3 ms·
Regarding the Pareto frontier and related benchmarks, I have a hard time taking anything seriously that claims that Opus is anywhere near the intelligence of Fa
by a13n 1mo ago
Regarding the Pareto frontier and related benchmarks, I have a hard time taking anything seriously that claims that Opus is anywhere near the intelligence of Fable. Are there any benchmarks that haven't just been benchmaxxed that more accurately represent actual usage?
- kakugawa 1mo agoFrontierCode is prob the closest. [1] It's closed source (so no direct benchmaxxing), and it was calibrated by 20+ open source maintainers. It shows Opus 5 (medium), beating out the other reasoning levels by a large margin. i.e. Opus 5 w/ higher reasoning levels actually reduces performance. [2] However, you'll have to gauge for yourself how closely their tasks resemble your tasks. 1/ https://cognition.com/blog/frontier-code https://cognition.com/blog/frontier-code 2/ https://cognition.com/frontiercode https://cognition.com/frontiercode