8 ms·
I think Anthropic rushed out the release before 10am this morning to avoid having to put in comparisons to GPT-5.3-codex! The new Opus 4.6 scores 65.4 on Termi
by granzymes 8mo ago
I think Anthropic rushed out the release before 10am this morning to avoid having to put in comparisons to GPT-5.3-codex!
The new Opus 4.6 scores 65.4 on Terminal-Bench 2.0, up from 64.7 from GPT-5.2-codex.
GPT-5.3-codex scores 77.3.
- __jl__ 8mo agoImpressive jump for GPT-5.3-codex and crazy to see two top coding models come out on the same day...
- granzymes 8mo agoInsane! I think this has to be the shortest-lived SOTA for any model so far. Competition is amazing.
- the_duke 8mo agoI do not trust the AI benchmarks much, they often do not line up with my experience. That said ... I do think Codex 5.2 was the best coding model for more complex tasks, albeit quite slow. So very much looking forward to trying out 5.3.
- NitpickLawyer 8mo agoJust some anecdata++ here but I found 5.2 to be really good at code review. So I can have something crunched by cheaper models, reviewed async by codex and then re-prompt with the findings from the review. It finds good things, doesn't flag nits (if prompted not to) and the overall flow is worth it for me. Speed loss doesn't impact this flow that much.
- kilroy123 8mo agoPersonally, I have Claude do the coding. Then 5.2-high do the reviewing.
- seunosewa 8mo agoThen I pass the review back to Claude Opus to implement it.
- VladVladikoff 8mo agoJust curious is this a manual process or you guys have automated these steps?
- ricketycricket 8mo agoI have a `codex-review` skill with a shell script that uses the Codex CLI with a prompt. It tells Claude to use Codex as a review partner and to push back if it disagrees. They will go through 3 or 4 back-and-forth iterations some times before they find consensus. It's not perfect, but it does help because Claude will point out the things Codex found and give it credit.
- bryanlarsen 8mo agoMind sharing the skill/prompt?
- dror 8mo agoNot the OP, but I use the same approach. https://gist.github.com/drorm/7851e6ee84a263c8bad743b037fb7abc https://gist.github.com/drorm/7851e6ee84a263c8bad743b037fb7a... I typically use github issues as the unit of work, so that's part of my instruction.
- _zoltan_ 8mo agozen-mcp (now called pal-mcp I think) and then claude code can actually just pass things to gemini (or any other model)
- kilroy123 8mo agoSometimes, depends on how big of a task. I just find 5.2 so slow.
- StephenHerlihyy 8mo agoI don’t use OpenAI too much, but I follow a similar work flow. Use Opus for design/architecture work. Move it to Sonnet for implementation and build out. Then finally over to Gemini for review, QC and standards check. There is an absolute gain in using different models. Each has their own style and way of solving the problem just like a human team. It’s kind of awesome and crazy and a bit scary all at once.
- readyforbrunch 8mo agoHow do you orchestrate this workflow? Do you define different skills that all use different models, or something else?
- nitroedge 8mo agoYou should check out the PAL MCP and then also use this process, its super solid: https://github.com/glittercowboy/get-shit-done https://github.com/glittercowboy/get-shit-done The way "Phases" are handled is incredible with research then planning, then execution and no context rot because behind the scenes everything is being saved in a State.md file... I'm on Phase 41 of my own project and the reliability and almost absence of any error is amazing. Investigate and see if its a fit for you. The PAL MCP you can setup to have Gemini with its large context review what Claude codes.
- jahsome 8mo agoAnother day, another hn thread of "this model changes everything" followed immediately by a reply stating "actually I have the literal opposite experience and find competitor's model is the best" repeated until it's time to start the next day's thread.
- malshe 8mo agoThis pretty accurately summarizes all the long discussions about AI models on HN.
- wasmainiac 8mo ago[flagged]
- BoredPositron 8mo agoWhen you keep his ramblings on twitter or company blog in mind I bet he is a shit poster here.
- locknitpicker 8mo ago> All anonymous as well. Who are making these claims? script kiddies? sr devs? Altman? You can take off your tinfoil hat. The same models can perform differently depending on the programming language, frameworks and libraries employed, and even project. Also, context does matter, and a model's output greatly varies depending on your prompt history.
- andrepd 8mo agoIt's hardly tinfoil to understand that companies riding a multi-trillion dollar funding wave would spend a few pennies astroturfing their shit on hn. Or overfit to benchmarks that people take as objective measurements.
- nocman 8mo ago> Who are making these claims? script kiddies? sr devs? Altman? AI agents, perhaps? :-D
- fooker 8mo agoYeah, these benchmarks are bogus. Every new model overfits to the latest overhyped benchmark. Someone should take this to a logical extreme and train a tiny model that scores better on a specific benchmark.
- mrandish 8mo ago> Yeah, these benchmarks are bogus. It's not just over-fitting to leading benchmarks, there's also too many degrees of freedom in how a model is tested (harness, etc). Until there's standardized documentation enabling independent replication, it's all just benchmarketing .
- scoring1774 8mo agoThis has been done: https://arxiv.org/abs/2510.04871v1 https://arxiv.org/abs/2510.04871v1
- bunderbunder 8mo agoAll shared machine learning benchmarks are a little bit bogus, for a really “machine learning 101” reason: your test set only yields an unbiased performance metric if you agree to only use it once. But that just isn’t a realistic way to use a shared benchmark. Using them repeatedly is kind of the whole point. But even an imperfect yardstick is better than no yardstick at all. You’ve just got to remember to maintain a healthy level of skepticism is all.
- 8mo ago
- aurareturn 8mo ago5.2 Codex became my default coding model. It “feels” smarter than Opus 4.5. I use 5.2 Codex for the entire task, then ask Opus 4.5 at the end to double check the work. It's nice to have another frontier model's opinion and ask it to spot any potential issues. Looking forward to trying 5.3.
- koakuma-chan 8mo agoOpus 4.5 is more creative and better at making UIs
- hypercube33 8mo agoUnless it's scroll bar theming then my God it's bad. it told me it gives up. Gemini 3 got stuck but the right prompt it did work.
- nerdsniper 8mo agoOpus 4.5 still worked better for most of my work, which is generally "weird stuff". A lot of my programming involves concepts that are a bit brain-melting for LLMs, because multiple "99% of the time, assumption X is correct" are reversed for my project. I think Opus does better at not falling into those traps. Excited to try out 5.3
- nubg 8mo agowhat do you do?
- audience_mem 8mo agoHe works on brain-melting stuff, the understanding of which is far beyond us.
- nerdsniper 8mo agoIt's relatively easy for people to grok, if a bit niche. Just sometimes confuses LLMs. Humans are much better at holding space for rare exceptions to usual rules than LLMs are.
- mmaunder 8mo agoARG-AGI-2 leaderboard has a strong correlation with my Rust/CUDA coding experience with the models.
- int_19h 8mo agoCodex 5.3 seems to be a lot chattier. As in, it comments in the chat about things it has done or is about to do. They don't show up as "thinking" CoT blocks, but as regular outputs, but overall the experience is somewhat more like Claude is in that you can spot the problems in model's reasoning much earlier if you keep an eye on it as it works, and steer it away.
- nurettin 8mo agoOpus was quite useless today. Created lots of globals, statics, forward declarations, hidden implementations in cpp files with no testable interface, erasing types, casting void pointers, I had to fix quite a lot and decouple the entangled mess. Hopefully performance will pick up after the rollout.
- nickstinemates 8mo agoDid you give it any architecture guidance? An architecture skill that it can load to make sure it lays out things according to your taste?
- nurettin 8mo agoYes, it has a very tight CLAUDE.md which it used to follow. Feels like this happens a couple of times a month.
- leumon 8mo agothey tested it at xhigh reasoning though, which is probably double the cost of Anthropic's model. Cost to Run Artificial Analysis Intelligence Index: GPT-5.2 Codex (xhigh): $3244 Claude Opus 4.5-reasoning: $1485 (and probably similar values for the newer models?)
- wilg 8mo agoIn my personal experience the GPT models have always been significantly better than the Claude models for agentic coding, I’m baffled why people think Claude has the edge on programming.
- dudeinhawaii 8mo agoI think for many/most programmers = 'speed + output' and webdev == "great coding". Not throwing shade anyone's way. I actually do prefer Claude for webdev (even if it does cringe things like generate custom CSS on every page) -- because I hate webdev and Claude designs are always better looking. But the meat of my code is backend and "hard" and for that Codex is always better, not even a competition. In that domain, I want accuracy and not speed. Solution, use both as needed!
- whynotminot 8mo ago> Solution, use both as needed! This is the way. People are unfortunately starting to divide themselves into camps on this — it’s human nature we’re tribal - but we should try to avoid turning this into a Yankees Redsox. Both companies are producing incredible models and I’m glad they have strengths because if you use them both where appropriate it means you have more coverage for important work.
- falloutx 8mo ago> I actually do prefer Claude for webdev Ah and let me guess all your frontends look like cookie cutter versions of this: https://openclaw.dog/ https://openclaw.dog/
- Yiin 8mo agoYes and I love it.
- flir 8mo agoThat's the best theory I've heard. Or at least, it's the one that fits with my usage. I'm mostly-backend, and I'm mostly-GPT. (I'm also a "small steps under guidance" user rather than a "fire and forget" user, so maybe that plays into it too).
- jronak 8mo agoDid you look at the ARC AGI 2? Codex might be overfit for terminal bench
- tedsanders 8mo agoARC AGI 2 has a training set that model providers can choose to train on, so really wouldn't recommend using it as a general measure of coding ability.
- janalsncm 8mo agoMore fundamentally, ARC is for abstract reasoning. Moving blocks around on a grid. While in theory there is some overlap with SWE tasks, what I really care about is competence on the specific task I will ask it to do. That requires a lot of domain knowledge. As an analogy, Terence Tao may be one of the smartest people alive now, but IQ alone isn’t enough to do a job with no domain-specific training.
- mrandish 8mo agoA key aspect of ARC AGI is to remain highly resistant to training on test problems which is essential for ARC AGI's purpose of evaluating fluid intelligence and adaptability in solving novel problems. They do release public test sets but hold back private sets. The whole idea is being a test where training on public test sets doesn't materially help. The only valid ARC AGI results are from tests done by the ARC AGI non-profit using an unreleased private set. I believe lab-conducted ARC AGI tests must be on public sets and taken on a 'scout's honor' basis that the lab self-administered the test correctly, didn't cheat or accidentally have public ARC AGI test data slip into their training data. IIRC, some time ago there was an issue when OpenAI published ARC AGI 1 test results on a new model's release which the ARC AGI non-profit was unable to replicate on a private set some weeks later (to be fair, I don't know if these issues were resolved). Edit to Add: Summary of what happened: https://grok.com/share/c2hhcmQtMw_66c34055-740f-43a3-a63c-4b03d336c214 https://grok.com/share/c2hhcmQtMw_66c34055-740f-43a3-a63c-4b... I have no expertise to verify how training-resistant ARC AGI is in practice but I've read a couple of their papers and was impressed by how deeply they're thinking through these challenges. They're clearly trying to be a unique test which evaluates aspects of 'human-like' intelligence other tests don't. It's also not a specific coding test and I don't know how directly ARC AGI scores map to coding ability.