3 ms·
Weren't they caught multiple times gaming the benchmark even more so then the rest?
by qpricjalcbeu 3mo ago
Weren't they caught multiple times gaming the benchmark even more so then the rest?
- alansaber 3mo agoLet me assure you, literally everybody does this
- sheepscreek 3mo agoI don't think it even matters. Because noone will continue to use an LLM that doesn't work well for them, whether or not it has a good bench result. So for their own sake, the correct representation can actually win them some loyalty: eg. Model X is weaker than Fable, but competes well with Opus/Sonnet and costs 1/5th as much etc - something similar playing out with Grok 4.5.
- zmmmmm 3mo agoYes and Zuck effectively disbanded the entire team that did that. Not saying we shouldn't cast a critical eye on it, but it probably does warrant a second chance.
- qpricjalcbeu 3mo agoZuck was part of that team.