4 ms·
I haven’t seen it discussed anywhere that closed models can essentially cheat benchmarks right? What Anthropic or OpenAI brand as a model doesn’t necessarily ha
by cedws 3mo ago
I haven’t seen it discussed anywhere that closed models can essentially cheat benchmarks right? What Anthropic or OpenAI brand as a model doesn’t necessarily have to be just weights, it can be a whole backend system that augments the model itself. With this they can score better benchmarks than an open source model that is weights alone.
- snthpy 3mo agoGood point
- jstanley 3mo agoSure, I think that's fine, that all counts. It counts for open source too, it's not like they're somehow running these benchmarks without any harness. Nobody cares if your AGI is 100% made out of neural networks or if it's like 50% neural networks and 50% perl scripts.
- stkdump 3mo agoI think they mean cheat in a Dieselgate sense. You detect that you are being tested with a specific benchmark question and heuristically give the correct (manually programmed) answer. That wouldn't be AGI.