4 ms·
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort
by GodelNumbering 1mo ago
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
- XCSme 1mo agoThat's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.
- GodelNumbering 1mo ago> That's quite common with many models Such as? I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.
- XCSme 1mo agoIn my own tests on aibenchy.com, where questions are quite simple, higher reasoning efforts consistently used to do worse than medium for most models. The reasoning effort should match the complexity of the task against the model's capability. Hard task with low reasoning = bad Easy task with very high reasoning = bad
- minatoaqua1 1mo agogrok 4.6
- JacobAsmuth 1mo agoLOL
- desterothx 1mo agoiirc, some of the original fable bemchmarks showed this. Definitely saw it in other frontier releases though
- m0zzie 1mo agoI find this very amusing, given we humans are also highly susceptible to this.