3 ms·
I really don't get these companies posting disingenuous benchmarks. Every time, they pick and choose who to compare against. Not comparing to the latest 5.3-cod
by deanc 8mo ago
I really don't get these companies posting disingenuous benchmarks. Every time, they pick and choose who to compare against. Not comparing to the latest 5.3-codex is absurd when it's been out a couple of weeks now. Who are they trying to kid?
- AdamConwayIE 8mo agoThere aren't really any of the typical benchmark suites targeting Codex 5.3 because it's still not in the API. SWE bench for example creates a predictions file and evaluates the results in the harness. Without Codex 5.3 being in the API, it can't.
- rvz 8mo ago> Who are they trying to kid? People who do not know how reproducible research works. Any benchmark that is presented by AI labs must be reproduced reliably by someone else independent of that AI lab presenting these results. Otherwise, not only it is biased, these numbers can be just made up for marketing purposes.
- falloon 8mo agoIf you were writing a promotional post for your new model, would you include benchmarks of a competitor that's spanking you across the board? This is marketing.
- tomlis 8mo agogpt-5.3-codex isn't available via the API yet. Pretty sure they were only testing via API access.