3 ms·
> I'm a "practitioner", well a researcher. Most of my peers that I talk to believe evals/benchmarks are a strong indicator of real world performance as well as
by maeil 2y ago
> I'm a "practitioner", well a researcher. Most of my peers that I talk to believe evals/benchmarks are a strong indicator of real world performance as well as much more abstract notions such as ability to think
If this includes the benchmarks used by the frontier model providers, I'm stunned to hear this. It's abundantly clear that they're trained on and gamed to oblivion, rendered meaningless.
Or are these a different set of benchmarks? Or only when applied to research, non-commercial models?
- godelski 2y agoHumanEval has 60 authors. They thought that they could make up "leetcode style questions" and that because they were "hand written" that training on GitHub wouldn't spoil the test set. Hand crafted questions like returning the decimal part of a float or parsing nested parentheses. Have they seriously never heard of lisp? But you can find many of the canonical solutions verbatim in code prior to the cutoff. I guess no one checked. I'm not sure if people are just being quiet and going along with it or they are that disconnected. All those questions were obviously spoiled. The same is true, as you're noting. And if my CVPR reviews are any indication, they're laughable. I rarely see papers that do proper hypothesis testing. I shoot down poster papers that don't and will champion papers that do, even when I'm outnumbered. But I get why people don't write papers following the scientific style. If you isolate variables and make your work clear, reviewers take the easy way out and say the work is not novel. I've lost a lot of faith in my community. But then again, are we surprised? Look at how many experts supported Rabbit when it debuted. Every single one of them should have been able to sniff out a that scam. They were claiming the ability to do stuff beyond state of the art, on a tiny device, and did a really poor job faking the demo. I think there's an incentive structure for blinders and it's hindering our progress towards AGI or even more useful ML