3 ms·
Yea this may explain part of it or all of it, it’s likely a case by case kind of thing. Also to respond to the parent comment: benchmarks have a variety of dif
by astro1234 2mo ago
Yea this may explain part of it or all of it, it’s likely a case by case kind of thing.
Also to respond to the parent comment: benchmarks have a variety of difficulty levels. Humanity’s Last Exam, though now hitting the beginning of a saturation phase with Fable, was long unsaturated while other benchmarks saturated awhile ago. So that’s what I meant by Epoch capability index: using IRT models this effect so that you gather robust signals from variety of benchmark difficulties and can track progress over time as model capabilities have evolved (and so benchmarks have had to evolve to keep up).
But yes like I was saying: all benchmarks are problematic, some are useful. Benchmark quality problems abound, so 90% being the true ceiling is not surprising. There may be other factors at play here too, I haven’t studied this problem that deeply to have a good thorough answer to this. But keep in mind there are probably 50,000 benchmarks in the literature and that is not a joke number. A crapload of noise in that signal but it’s not all noise.
- ACCount37 2mo agoSaturation is mostly just selection effects in play. Throw out the "90% easiest" of tasks, and what remains is a jagged ladder of high difficulty outliers. Hard to climb, and hard to measure the climb - because you have less effective data points and the datapoints themselves are less linear, while you're still being subject to the measurement noise. Not having the mislabeled tasks would reduce the saturation, but it wouldn't drive it to zero. Even without the "infinite difficulty tasks", bell curve would do its thing.
- scotty79 2mo agoYou almost can't give a correct answer to a task with expected wrong answer unless you know it expects wrong answer. ... maybe on a yes/no question by chance. But as you become more knowledgeable, the probability that you give wrong but expected answer reduces. If a benchmark saturates to 100% it's very likely that answers leaked into the training data. In college I had a funny exam. It was on C++. One question I had to answer incorrectly because there was a mistake in the question. So I gave two answers for it, one that answered the question as it was and the other that answered the question as I inferred it was intended to be. It was appreciated. I got a honorary mark above the top possible (I gave correct answers to all other questions). I wouldn't be surprised if across so many, so huge benchmarks, there were tasks with wrong questions or answers in the key, that some LLM answered in and expected manner in the same fashion.