4 ms·
Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks" John Sous from Yale
by qt31415926 17d ago
Article: "How Good Are Frontier Models at Physics?
Expert Re-Grading Reveals Broken Evaluations and
Near-Saturation of Leading Benchmarks"
John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.
When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.
- fsh 17d agoI would be very surprised if any of the frontier models wasn't trained on all public physics benchmarks. Training data providers have been hiring people for exactly this task.
- letmevoteplease 17d agoThis study appears to be evidence against that: the model failed the benchmark but arrived at the correct answer.
- fsh 17d agoHalf of the benchmarks don't have published answers. That's why training data providers have been hiring physicists to solve them.
- deleted 17d ago[deleted]
- red75prime 17d agoThis is unconventional benchmaxxing then, when they decrease the benchmark scores to allow models to generalize on correct solutions.
- bobmarleybiceps 17d agoyeah, it would be almost shocking if an open source benchmark was NOT used ~somewhere in training. Perhaps just pre-training, but still. Neural networks can be fairly robust to some mistakes in their training data, so maybe it doesn't even matter if some of them are incorrect. Who knows.
- redwood 17d agoI'd have thought the same but this article from yesterday blew my mind https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit https://www.amazon.science/blog/why-dont-machine-learning-re... As it essential implies that these models compress knowledge well in a way that what remains is what's generalizeable more so than remembering every specific detail... Anyway more understanding necessary but thought provoking