3 ms·
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, a
by throw10920 2mo ago
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).
Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.
- staticman2 2mo agoDoes OpenAI give Artificial Analysis early access to test models? If so it's definitely not "some random account".
- throw10920 2mo agoArtificial Analysis was an clearly meant to be an example. I was obviously talking about the ideal scenario of proprietary benchmarking, not how it might be being screwed up in practice.
- inigyou 2mo agoSearch for accounts @artificialanalysis.com. read their chat history. optimise for that.
- throw10920 2mo agoPlease imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.
- inigyou 2mo agoWhy would I assume anything other than maximum incompetence from the AI ecosystem?
- throw10920 2mo agoThis is such a ridiculous and shallow cop-out. And it also is completely irrelevant to my challenge to show how proprietary benchmarking can be gamed, because it presumes (absolutely insane and divorced from reality) circumstances that have nothing to do with benchmarking as a concept or process.