3 ms·
I'm already quite put off by the title (it's science -- if you have a better benchmark, publish it!), but the contents aren't great either. It keeps citing numb
by bbor 11mo ago
I'm already quite put off by the title (it's science -- if you have a better benchmark, publish it!), but the contents aren't great either. It keeps citing numbers about "445 LLM benchmarks" without confirming whether any of the ones they deem insufficiently statistical are used by any of the major players. I've seen a lot of benchmarks, but maybe 20 are used regularly by large labs, max.
"For example, if a benchmark reuses questions from a calculator-free exam such as AIME," the study says, "numbers in each problem will have been chosen to facilitate basic arithmetic. Testing only on these problems would not predict performance on larger numbers, where LLMs struggle."
For a math-based critique, this seems to ignore a glaring problem: is it even possible to randomly sample all natural numbers? As another comment pointed out we wouldn't even want to ("LLMs can't accurately multiply 6-digit numbers" isn't something anyone cares about/expected them to do in the first place), but regardless: this seems like a vacuous critique dressed up in a costume of mathematical rigor.
At least some of those who design benchmark tests are aware of these concerns.
In related news, at least some scientists studying climate change are aware that their methods are imperfect. More at 11!
If anyone doubts my concerns and thinks this article is in good faith, just check out this site's "AI+ML" section: https://www.theregister.com/software/ai_ml/ https://www.theregister.com/software/ai_ml/
- daveguy 11mo agoThe article references this review: https://openreview.net/pdf?id=mdA5lVvNcU https://openreview.net/pdf?id=mdA5lVvNcU And the review is pretty damning regarding statistical validity of LLM benchmarks.
- dang 11mo ago(We've since changed both title and URL - see https://news.ycombinator.com/item?id=45860056 https://news.ycombinator.com/item?id=45860056)