2 ms·
Though this Codex version isnt on the leaderboard, GPT-5.2-Medium already seems to be a bit better than Opus 4.5: https://swe-rebench.com/ https://swe-rebench.c
by Mkengin 10mo ago
Though this Codex version isnt on the leaderboard, GPT-5.2-Medium already seems to be a bit better than Opus 4.5: https://swe-rebench.com/ https://swe-rebench.com/
- gizmodo59 10mo agoIs that your website or something? You keep promoting it
- Mkengin 10mo agoNo, I am not affiliated with the website, I just want to see more discussions based on uncontaminated benchmarks and feel that people rely too much on benchmarks that companies can conduct themselves. If that is the case, I don't feel I can trust them. For general LLM capabilities, for example, I would also tend to rely on dubesor [1] rather than artificial analysis or similar leaderboards. [1] https://dubesor.de/benchtable https://dubesor.de/benchtable