4 ms·
The obvious solution is to have non-public benchmarks. It is exceedingly difficult to train on a proprietary benchmark administered by someone with half a brai
by throw10920 2mo ago
The obvious solution is to have non-public benchmarks.
It is exceedingly difficult to train on a proprietary benchmark administered by someone with half a brain (i.e. don't sign up for a ChatGPT account with your benchmark@artificialanalysis.ai email) - you have to find a tiny needle in a vast haystack.
In fact, it can be difficult enough that it's simply not economically viable - that is, that it's cheaper to make the model better than it is to try to find the account running the benchmark.
In the limit case, the benchmark is indistinguishable from...normal problems that need to be solved.