3 ms·
I understand that running these benchmarks can get expensive, but it would be really nice to see AA include more benchmarks of models at reasoning settings othe
by AnodicElegy 2mo ago
I understand that running these benchmarks can get expensive, but it would be really nice to see AA include more benchmarks of models at reasoning settings other than the maximum, at least for the biggest releases. They have that nice graph of cost vs. composite benchmark score with the Pareto frontier line, but who knows if those are actually the optimal choices? There are already a few non-max-reasoning models on the Pareto line, among the few that were tested.
- apitman 2mo agoYou can turn on various levels of some of many of the models in the UI
- AnodicElegy 2mo agoYes, they have multiple levels of Claude, GPT, Gemini, and Kimi, but not the other top models (I would put GLM, Qwen, Muse, Grok, and Deepseek in that bucket).