4 ms·
How do the organisers keep the private test set private? Does openAI hand them the model for testing? If they use a model API, then surely OpenAI has access to
by miga89 2y ago
How do the organisers keep the private test set private? Does openAI hand them the model for testing?
If they use a model API, then surely OpenAI has access to the private test set questions and can include it in the next round of training?
(I am sure I am missing something.)
- 7734128 2y agoI suppose that's why they are calling it "semi-private".
- owenpalmer 2y agoI wouldn't be surprised if the term "benchmark fraud" will soon been coined.
- PhilippGille 2y agoBenchmark fraud is not a novel concept. Outside of LLMs for example smartphone manufacturers detect benchmarks and disable or reduce CPU throttling: https://www.theregister.com/2019/09/30/samsung_benchmarking_settlement/ https://www.theregister.com/2019/09/30/samsung_benchmarking_...
- hmottestad 2y agoCPU frequency ramp curve is also something that can be adjusted. You want the CPU to ramp up really quickly to make everything feel responsive, but at the same time you want to not have to use so much power from your battery. If you detect that a benchmark is running then you can just ramp up to max frequency immediately. It’ll show how fast your CPU is, but won’t be representative of the actual performance that users will get from their device.
- deneas 2y agoThey have two sets, a fully private one where the models run isolated and the semi-private one where they run models accessed over the internet.
- gritzko 2y agoThat is the top question, actually. Given all the billions at stake.
- PoignardAzur 2y agoIf we really want to imagine a cold-war-style solution, the two teams could meet in an empty warehouse, bring one computer with the model, one with the benchmarks, and connect them with a USB cable. In practice I assume they just gave them the benchmarks and took it on the honor system they wouldn't cheat, yeah. They can always cook up a new test set for next time, it's only 10% of the benchmark content anyway and the results are pretty close.
- andrepd 2y agoThere's no honor system when there's billions of dollars at stake x) I'm highly highly skeptical of these benchmarks because of intentional cheating and accidental contamination.
- bjornsing 2y agoIsn’t that why they call it “ Semi-Private”? There’s a fully private test set too as I understand it, that o3 hasn’t run on yet.
- daveguy 2y agoAnd o3 will not run on the private set unless it is a truly free and open source model (presumably also the case for ARC-AGI-2). This is the distinction between private and semi-private. In private you provide all the knowledge/weights/logic to operate without any external communication. Private benchmark results are the only true evaluation of performance on any benchmark -- reserved for a final evaluation. It is the only way to prevent shenanigans.