3 ms·
> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. It's not useful because the b
by re-thc 29d ago
> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses.
It's not useful because the benchmarks often measure the wrong thing. They're here yapping about AGI and yet the benchmarks treat it like a trained dog. Fetch this. 100 points.
Each "problem" in these benchmarks likely has more than 1 solution that can be considered correct and even should be graded in many ways. Yet we see in many benchmarks higher effort (or thinking levels) don't help because the benchmark penalizes for doing "more" than what the answers asks for. So what did you ask for?
In human school you often get marks on the process and not just the end result. Thinking tokens have been cut. All we group on is things like cost, turns and time but not the what else.