4 ms·
The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like t
by mchinen 3mo ago
The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations.
I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post.
If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressive over Fable, I will lose what remaining trust I had in these announcements.
- therobots927 3mo agoThe proof is in the pudding and these benchmark stats will only work for so long before people lose interest.
- thurn 3mo agoThey do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...
- mchinen 3mo agoI totally missed that, because in the charts they showcase for coding, the SWEBench score is not present, they only include it at the end of the post in tables. Hmm. Great catch.
- saberience 3mo agoThe SWEBench benchmarks are really gamed at this point and should not be trusted period. The solutions are effectively in the training sets and have been for a while.
- mnicky 3mo agoSWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.
- saberience 3mo agoSWE Bench Pro is also gamed and shouldn't be trusted.
- simianwords 3mo ago> SWEBench Pro, which is a pretty high-quality benchmark No, doesn't seem like it https://openai.com/index/separating-signal-from-noise-coding-evaluations/ https://openai.com/index/separating-signal-from-noise-coding...
- bel8 3mo agoSWEBench Pro should be ignored until they fix it or disprove the broken task accusations.
- danielsamuels 3mo agoIt's worth noting that OpenAI recently came out saying, "We don't think SWEBench Pro is worth reporting any more" - https://openai.com/index/separating-signal-from-noise-coding-evaluations/ https://openai.com/index/separating-signal-from-noise-coding...
- 37374848 3mo ago[flagged]
- William_BB 3mo agoPeople will downvote you because this comment is "not appropriate" for HN, but there were countless conversations on HN about how important these benchmarks are. I am literally LOLing at HN right now
- Sabinus 3mo agoYou're free to lol but at least put a substantive comment with it. 'Lol' by itself is low quality.
- anthonyrstevens 3mo agoIt's almost like there are multiple different people commenting on HN.
- lwansbrough 3mo agoDidn't they also just post about how SWEBench is broken?
- futureshock 3mo agoThere has been a lot of chatter ever since the Mythos scores had been release that SWEbench pro had major contamination and that Mythos had memorized many questions that lacked the context to be solvable on their own. And now with OpenAI saying a large number of the questions are broken, I think it's worth taking that single outlier benchmark with some salt when the overall trend is that 5.6 is very competitive with Mythos at about half the price.
- enraged_camel 3mo agoThe charts are also extremely difficult to parse. They seem auto-generated. Dataset coloring is atrocious. Regarding your main point, yes, I agree. My impression (as someone who uses both Codex and Claude Code daily) is that OpenAI does a fair amount of benchmaxxing.
- deleted 3mo ago[deleted]