6 ms·
GPT-5 vs. Sonnet: Complex Agentic Coding
- arcticfox 1y ago> Note that Claude 4 Sonnet isn’t the strongest model from Anthropic’s Claude series. Claude Opus is their most capable model for coding, but it seemed inappropriate to compare it with GPT-5 because it costs 10 times as much. Well - I would have been interested in GPT-5 vs. Opus. Claude Code Max is affordable with Opus.
- swader999 1y agoYou're absolutely right!
- intellectronica 1y ago:D
- rubslopes 1y ago"I see it now!"
- qeternity 1y ago> Claude Code Max is affordable with Opus Because Anthropic is presumably massively subsidizing the usage.
- kvirani 1y agoIsn't it all heavily subsidized by VC money at this time?
- adventured 1y agoOpenAI for its part is tracking to $12-$15 billion in annual sales. If they slapped a basic ad model for referring onto what they're already doing, it's an easily profitable enterprise doing $30+ billion in sales next year. Frankly they should have already built and deployed that, it would make their free versions instantly profitable and they could boost usage limits and choke off the competition. It's the very straight-forward path to financially ruining their various weaker competition. Anthropic is Lyft in this scenario (and I say that as a big fan of Claude).
- inquirerGeneral 1y ago[dead]
- kingstnap 1y agoThe APIs are marginally profitable. You can calculate the lifecycle costs of the open models on clusters in batched inference and figure out its less than than what they charge. The training and researches are very expensive. The fixed price subscriptions are 100% a sweetheart deal.
- Filligree 1y agoWhich doesn’t factor into my immediate decisions.
- carterparks 1y agoI'm getting an SSL error in Chrome: ERR_SSL_PROTOCOL_ERROR
- SV_BubbleTime 1y ago> but when I'd point out the missing implementation, it would give its usual "you're absolutely right" and try to fix it. I really trying to not be annoyed by Claude’s “You’re absolutely right” because I know I cannot control it but this is an increasingly difficult task.
- jpalawaga 1y agoI think it's because "you're right!" somehow presupposes it knew the answer and was just testing you. an intern never says that. they say "oh, I see."
- AlecSchueler 1y agoDoes it also seem to be getting worse this way?
- CamperBob2 1y agoYou can control it in the chat page, at least (User name at lower left->Settings). I use this: Answer concisely when appropriate, more extensively when necessary. Avoid rhetorical flourishes, bonhomie, and (above all) cliches. Take a forward-thinking view. OK to be mildly positive and encouraging but NEVER sycophantic or cloying. Above all, NEVER use the phrase "You're absolutely right." Rather than "Let me know if..." style continuations, list a set of prompts to explore further topics. That last bit causes some clutter at the end of each response, not sure if I'm going to keep it. But it does do a good job at following these guidelines in my experience. The same basic instructions also work well in ChatGPT and Gemini. Does Claude Code not support anything like this?
- unshavedyak 1y agoClaude Code has "memory" files which can be layered (global, project, private). I'll probably add this to mine haha, but mine currently mostly consist of things that i can't automate. Eg i have hooks for tests/lints, so i don't need to tell claude to do those things. However general styling preferences and naming conventions can be a bit more difficult to automate, so i put those in the memory files. I find it does decently, but it's far from perfect. Eg before hooks i had "ALWAYS format and lint" style entries in the memory file and it probably had a 70% success rate. Often it would go on little side paths cleaning work up and forget to run lints after or w/e. Formatting was my biggest gripe. Deterministic wrappers have been the biggest gain for me personally. I suspect they'll get a lot better over time too. Eg i want to try and find a way to write more personal style guides in a linter to enforce claude not break various conventions i prefer. But of course, that can be difficult.
- blurbleblurble 1y agoClaude is just so well rounded and considerate. A lot of this probably comes down to prompt and context engineering, though surely there's something magical about Anthropic's principled training methodologies. They invented constitutional AI and I can only imagine that behind the scenes they're doing really cool stuff. Can't wait to see Claude 5!
- stitched2gethr 1y agoThis take rings true for me after admittedly only a couple of hours of use of gpt-5. I had an issue I had been working with Claude on but it was difficult to give it real-time feedback so it floundered. gpt-5 struggled in the same areas but after about $2 of tokens it did fix the issue. It was far from a 1 shot like I might have expected from the hype, but it did get the job in about an hour done where Claude could not in 3. For reference my Claude usage was mostly Sonnet, but with consulting from Opus.
- 0xfaded 1y agoWould you be comfortable sharing a brief description of what the issue was?
- indigodaddy 1y agoWhat does the 1x and .33x mean on the list of models in copilot? (Never used but thinking about trying on the free tier)
- commandar 1y agoThey're multipliers against your quota of requests. GPT-4.1 is "free" with a copilot sub, and then the premium models would burn credits against a multiplier. So higher multipliers count more against your monthly quota. GPT5, Sonnet 4, and Gemini Pro 2.5 are all 1x. Opus is 10x, for comparison. https://docs.github.com/en/copilot/reference/ai-models/supported-models https://docs.github.com/en/copilot/reference/ai-models/suppo... Also worth keeping in mind that Copilot has reduced context windows even for the premium models, which has a very real impact on agentic performance.
- indigodaddy 1y agoThanks for the info. Would you consider GH Copilot the best bang for buck currently, or would you recommend just going with the Claude $20 plan? I'm definitely not looking to spend a lot of money, just want to see what kind of mileage I can get on low-end plans
- commandar 1y agoIt's going to depend heavily on your usage. I use Copilot because work is paying for it and it can be made usable, but requires being really deliberate about managing context to keep things on the rails. It's nice that it gives you access to a pretty decent selection of models, though. At home, I'm mostly using the $100 Claude plan. It's definitely not cheap, but I've found it has a pretty decent balance for my casual experiments with agentic coding. Another option to seriously consider is setting up an account with OpenRouter and just tossing some cash into your bucket on occasion. OpenRouter lets you arbitrarily make API requests to pretty much any model you want. I've been occasionally tossing $10 or so into mine and I'll use it when I've hit my usage limits with Claude or if I want to see how another model will attack a particular task. FWIW, I use Roo code for all of this, so it's pretty easy for me to switch between models/providers as I need to.
- arresin 1y agoGithub copilot is utter garbage. The diffing crawls along at a snail’s pace. I think it’s coming up on two years and this must criticised aspect of it still isn’t fixed—-even with all the reverse engineering of how cursor did it. I wish I could find an alternative to cursor (which has other issues). Honestly, that company just threw away a golden opportunity as the first mover.
- sourcecodeplz 1y agoWhy did they throw it away? Because of the new opaque pricing?
- swader999 1y agoThey let their moat dry right up.
- bredren 1y agoI've done evaluations of Github Copilot, Sourcegraph Cody and Gitlab Duo and Copilot is not garbage, but rather the by far leader among these other options.
- zhivota 1y agoDid you compare to Cursor? We gave up on Copilot a while back after Cursor blew us away. In the context of this article though, Cursor is very obviously tuned better towards Claude than OpenAI in my experience.
- grumple 1y agoCursor's agent is better, but the in-editor suggestions by Copilot when you're actually the one coding are very useful. Claude's agent is better than Cursor, so I'm not sure where Cursor fits in with this ecosystem.
- mkozlows 1y ago"Best option among loser tools" isn't the high praise you think it is, though.
- chromejs10 1y agoThis should have been compared with Opus... I know OP says he didn't because of cost but if you're comparing who is better then you need to compare the best to the best... if Claude Opus 4.1 is significantly better than GPT 5 then that could offset the extra expense. Not saying it will... but forget cost if we want to compare solely the quality
- qeternity 1y ago> but forget cost if we want to compare solely the quality I think this is the whole reason not to compare it to Opus...
- bgirard 1y agoI agree. Opus is cost prohibitive for most longer coding tasks. The increase output doesn't justify the cost.
- fouc 1y agogpt-5 isn't supposed to be the best, it's supposed to be cost effective
- senko 1y agoFrom OpenAI website: > Our smartest, fastest, and most useful model yet I'd say it's definitely supposed to be the best, it just doesn't deliver.
- cheema33 1y ago>> Our smartest, fastest, and most useful model yet > I'd say it's definitely supposed to be the best, it just doesn't deliver. What part of "Our" is difficult to understand in that statement? Or are you claiming that OpenAI owns another model that is clearly better than GPT-5?
- senko 1y ago
- anotheryou 1y agoI think we need to stop testing models raw. Claude is trained for claude code and that's how it's used in the field too.
- nightshift1 1y agounless you use it through copilot
- anotheryou 1y agoClaude? why would you
- fouc 1y agoClaude Code can be used through VS code too. As for why? Some people prefer IDEs over terminals. Personally I think the attempts to combine LLM coding with current IDE UIs, a la Cursor/Windsurf/VS Code is probably the wrong way to go, it feels too awkward and cumbersome. I like a more interactive interface, and Claude Code is more in line with that.
- anotheryou 1y agoany ide is fine. Quick free edits with windsurf if I'm just too lazy to format css or something is nice
- rezistik 1y agoMy work pays for copilot subscriptions, they don't pay for claude code.
- anotheryou 1y agoI see. id still suggest benchmarking the sota setup first
- Nizoss 1y agoI have been using Claude Code with TDD through hooks, which significantly improved my workflow for production code. Watching the ChatGPT 5 demo yesterday, I noticed most of the code seemed oriented towards one-off scripts rather than maintainable codebases which limits its value for me. Does anyone know if ChatGPT 5 or Copilot have similar extensibility to enforce practices like TDD? For context on the approach: https://github.com/nizos/tdd-guard https://github.com/nizos/tdd-guard I use pre/post operation commands to enforce TDD rules.
- deleted 1y ago[deleted]
- MrGreenTea 1y agoI just recently stumbled upon your tdd-guard when looking for inspiration for Claude hooks. I've been so impressed with what it allowed me to improve the workflow and quality. Then I was somewhat disappointed that almost no one seems to talk about this potential and how they're using hooks. Yours was the only interesting project I found in this regard and hope to give it a spin this weekend . You don't happen to have a short video where you go into a bit more detail on how you use it though?
- Nizoss 1y agoThank you for the kind words, it means a lot! I spent my summer holiday on this because I truly believe in the potential of hooks in agentic coding. I'm equally surprised that this space hasn't been explored more. I'm currently working on making the validation faster and more customizable, plus adding reporters to support more languages. I think there is an Amazon backed vscode forked that is also exploring this space. I think they market it as spec driven development. Edit: I found it, its called Kiro: https://kiro.dev/ https://kiro.dev/
- Nizoss 1y agoSorry, I missed the second part of your comment! I don't have a detailed video beyond the short demo on the repo, but I'll look into recording something more comprehensive or cover it in a blog post. Happy to ping you when it's ready! In the meantime: I simply set it up and go about my work. The only thing I really do is just nudge the agent into making architectural simplifications and make sure that it follows the testing strategies that I like: dependency injection, test helpers, test data factories and such. Things that I would do regardless of the hook. I like to give my tests the same attention and care that I give production code. They should be meaningful and resilient. The code base contains plenty of examples but I will look into putting something together.
- chisleu 1y agoHow was he doing "complex agentic coding" when the APIs have such extreme context and throughput limitations?
- DrNosferatu 1y agoThen instruct GPT5 to write more structured and annotated code.
- mewpmewp2 1y agoSo far from my testing I have found Claude Code with Sonnet 4 better than Cursor + GPT-5 still. I started exact same projects at the same time, and it seemed Claude Code was just more reliable. It was just much slower in terms of setting up the project and didn't setup the project up as scalably (despite them highlighting that in the demo), and when I tried to instruct it to set it up DRY, modular, etc it kind of didn't just go where I wanted it to, while Claude Code did. It was a game involving OOP, three.js. I think both are probably great at good design and CRUD things.
- nextworddev 1y agoGPT-5 is much cheaper though
- mewpmewp2 1y agoI'm using Claude Code which is $200/month, and I do multiple agents, subagents, terminals at the same time, much faster than Cursor. I get almost 24/7 dev time from that.
- natiman1000 1y agoI was initially excited about GPT5, and I quickly switched to it but still can't use it for some reason it is clearly smart but not useful.
- mewpmewp2 1y agoI'm getting the same thing I got with Codex when I tried right now. I give it a command, and it keeps reading files and thinking for 5 min+, this never happens with Claude Code.
- dwaltrip 1y agoOh I tried codex for the first time last night with gpt-5. It looked stuck twice when it used the grep tool (after working successfully for minute or so), and both times I canceled after seeing no output for more than a minute. It would have eventually finished?
- nojs 1y agoThis pretty much matches my experience today. GPT5 (in Cursor) feels smarter in isolation, but CC with Opus is faster and better at real tasks involving a large codebase.
- mvATM99 1y agoThe manual approval of commands in GHCP can be circumvented, there's an experimental setting that allows you to accept all commands automatically. I wish you could be a bit more specific though, you can't set which commands you want to auto-accept in detail.
- typpilol 1y agoPretty sure you can set a terminal whitelist and blacklist for it.
- patcon 1y ago> One continuous difference: while GPT-5 would do lots of thinking then do something right the first time, Claude frantically tried different things — writing code, executing commands, making pretty dumb mistakes [...], but then recovering. This meant it eventually got to correct implementation with many more steps. Sounds like Claude muddles. I consider that the stronger tactic. I sure hope GPt-5 is muddling on the backend, else I suspect it will be very brittle. Re: https://contraptions.venkateshrao.com/p/massed-muddler-intelligence https://contraptions.venkateshrao.com/p/massed-muddler-intel... > Lindblom’s paper identifies two patterns of agentic behavior, “root” (or rational-comprehensive) and “branch” (or successive limited comparisons), and argues that in complicated messy circumstances requiring coordinated action at scale, the way actually effective humans operate is the branch method, which looks like “muddling through” but gradually gets there, where the root ["godding through"] method fails entirely.
- sudohalt 1y agoIsn't the issue with that the prohibitive costs, it can easily be 5 to 10 (maybe even more for long running tasks). Currently they are probably subsidizing the compute costs to some extent.
- quijoteuniv 1y agoToday I used GPT-5 for some OpenTelemetry Collector configs that both Claude and OpenAI models struggled with before and it was surprisingly impressive. It got the replies right on the first try. Previously, both had been tripped up by outdated or missing docs (OTel changes so quickly). For home projects, I wish I could have GPT-5 plugged into Claude’s code CLI interface. iteration just works! Looking forward to less baby sitting in the future!
- mattnewton 1y agoCursor CLI is pretty close to Claude code- it’s missing a bunch of features like being able to manually compact or define sub agents, but the basic workflow is there and if you squint it’s pretty close to gpt-5 in Claude code. I haven’t tried codex cli recently yet, I think it just got an update. That would be another to investigate.
- endorphine 1y agoWhat is the way to use this agentic stuff with neovim? Do I have to resort to OpenAI's Codex or a nvim plugin is sufficient? Or Claude Code?
- mkozlows 1y agoClaude Code or cursor-cli or Codex or any of the command-line tools should be good. (Claude Code seems so far to be the option people like best of those, though.)
- Cyphus 1y agoI've been using codecompanion.nvim[0] combined with mcp-hub.nvim[1]. Code Companion works well for interactive chat but falls short for agentic coding. It's limited to some pre-configured and user-defined "workflows" which are basically templated steps with prompts, actions, and loops. I've been meaning to give avante.nvim[2] a try since it aims to provide a "Cursor like" experience, but for now I've been alternating between Code Companion for simple prompts and Claude CLI (in a tmux pane next to Neovim) for agentic stuff. [0] https://codecompanion.olimorris.dev/ https://codecompanion.olimorris.dev/ [1] https://ravitemer.github.io/mcphub.nvim/ https://ravitemer.github.io/mcphub.nvim/ [2] https://github.com/yetone/avante.nvim https://github.com/yetone/avante.nvim
- roguesherlock 1y agoI've found these two to be really good https://github.com/dlants/magenta.nvim https://github.com/dlants/magenta.nvim https://github.com/NickvanDyke/opencode.nvim https://github.com/NickvanDyke/opencode.nvim
- h4ny 1y agoI have been seeing different people reporting different results with different tasks. Watched a live stream that compared GPT-5, Gemini Pro 2.5, Claude 4 Sonnet, and GLM 4.5, and GPT-5 appeared to not follow instructions as well as the other three. At the moment it feels like most people "reviewing" models depends on their believes and agenda, and there are no objective ways to evaluate and compare models (many benchmarks can be gamed). The blurring boundaries between technical overview, news, opinions and marketing is truly concerning.
- x187463 1y agoThis has been ubiquitous for a while. Even here on HN every thread about these models (even this one, I'm sure) features an inordinate amount of disagreement between people vehemently declaring one model more useful than another. There truly seems to be no objective measurement of quality that can discern the difference between frontier models.
- physix 1y agoI think this is actually good, because it means there is no clear winner who can sit back and demand rent. Instead they all work as hard as they can to stay competitive, hopefully thereby accelerating AI software engineering capabilities, with the investors footing the bill.
- NitpickLawyer 1y agoYeah, I agree. And prices are slowly coming down. Gemini 2.5 was cheaper than claude4, and (again depending on task) either on par or slightly below in quality. Now gpt5 is cheaper still (I think their -main is 10$/M?) and they also have -mini and -nano versions. The more choices we have the better it will be. As you said, without a clear winner we're about to get spoiled for choice, and there's no clear way for them to just sit on stuff and increase prices (yet). Plus there's some pressure coming from the open source releases. Not there in quality, but they are runnable "on prem", pretty cheap and keep getting better.
- isaacremuant 1y ago
- olddustytrail 1y agoFrom reactions I've seen it appears that GPT-5 hallucinates less than previous models but the flip side is that it's worse for creative tasks. This makes logical sense: you don't want a model to get creative if you need functioning code, but if you want a story idea it should basically be all hallucination. I think it makes sense to have different models for these tasks.
- lherron 1y agoDid I miss the total cost for each run in the article? Can't seem to find it. If Sonnet is more expensive AND more chatty/requires more attempts for the same result, seems like that would favor GPT5 for daily driver.
- animex 1y agoWonderful, timely article. It sounds like a hybrid approach might produce good results: Using ChatGPT-5 for planning/analysis and using Claude for execution.
- blurbleblurble 1y agoAnother option would be to add something to your AGENTS.md or whatever giving examples of the kind of code organization you want. You could ask Claude to explain its approach in terms explicitly that GPT-5 can understand. GPT-5 seems much more sensitive in its responsiveness to instructions. My sense is that in the long run this will be really nice, but that the current prompts in these mainstream LLM coding tools are designed for models with a different style of responsiveness to instructions.
- Surac 1y agoTypescript to rust. I mostly test models on c code. C is much less boilerplate and more code per word. Models need to be ready smart to see all the pointer magic and misuse of lib functions. Claude really makes a very competent c coder in my test
- doctoboggan 1y agoI really like Claude code's context engineering and prompt engineering, is it possible to plug in GPT-5 into Claude code? I think that would be a more apples to apples test as it's just testing the models and not the agentic framework around them.
- koakuma-chan 1y agoI imagine Claude Code is optimized for Claude specifically, and GPT-5 would not be great there. You should probably use Codex if you want to use GPT-5.
- jjani 1y agoIs it really this easy now to get your article high on HN with 100 comments? The findings are completely meaningless. "Agenticness" depends so much on the specific tooling (harness) and system prompts. It mentions Copilot - did it use this for both? Given it's created by Microsoft there's good reason to believe it'd be built yo do especially well with GPT (they'll have had 5 available in preview for months by now). Or it could be the opposite and be tuned towards Sonnet. At the very minimum you'd need to try a few different harnesses, preferably ones not closely related to either OpenAI/MS or Anthropic. This article even mentions things like "Sonnet is much faster" which is very dependent on the specific load at the time of usage. Today everyone is testing GPT-5 so it's slow and Sonnet is much faster.
- debarshri 1y agoI guess agents are voting it up
- intellectronica 1y agoOP here. I actually agree with you that the "findings" here are meaningless. This is pure vibe. Also regarding "Sonnet is faster" I did explicitly mention that I believe this is because GPT-5 is in preview and hours from the release. The speed I experienced doesn't say anything about the model performance you can expect.
- jjani 1y agoIf they're meaningless why'd you post it besides getting views on your blog? > Also regarding "Sonnet is faster" I did explicitly mention that I believe this is because GPT-5 is in preview and hours from the release. I genuinely don't see this mentioned, where is it?
- ramesh31 1y ago>Is it really this easy now to get your article high on HN with 100 comments? Everyone wants to know the answer to GPT5 vs Claude without wasting the tokens personally because we can all more or less guess what the result will be.
- ramoz 1y agoNot using Claude code is a crime.
- siamtttt 1y ago[dead]
- siamtttt 1y ago[dead]
- ctbellmar 1y agoI know it's been mentioned a few times, but worth repeating: these LLMs tend to do noticeably better in their own native environments. Claude (Opus or Sonnet) in Copilot != Claude in Claude Code. Same applies to Cursor, Windsurf, Augment, etc. This likely has a lot to do with context manipulation (and compression), which affects the resulting output. I imagine that GPT-5 likewise will do better in Codex vs 3rd party plugin/VS Code fork.
- fragmede 1y agoThe system prompts aren't shared either, and probably accounts for quite a bit of difference as well.