8 ms·
This should have been compared with Opus... I know OP says he didn't because of cost but if you're comparing who is better then you need to compare the best to
by chromejs10 1y ago
This should have been compared with Opus... I know OP says he didn't because of cost but if you're comparing who is better then you need to compare the best to the best... if Claude Opus 4.1 is significantly better than GPT 5 then that could offset the extra expense. Not saying it will... but forget cost if we want to compare solely the quality
- qeternity 1y ago> but forget cost if we want to compare solely the quality I think this is the whole reason not to compare it to Opus...
- bgirard 1y agoI agree. Opus is cost prohibitive for most longer coding tasks. The increase output doesn't justify the cost.
- fouc 1y agogpt-5 isn't supposed to be the best, it's supposed to be cost effective
- senko 1y agoFrom OpenAI website: > Our smartest, fastest, and most useful model yet I'd say it's definitely supposed to be the best, it just doesn't deliver.
- cheema33 1y ago>> Our smartest, fastest, and most useful model yet > I'd say it's definitely supposed to be the best, it just doesn't deliver. What part of "Our" is difficult to understand in that statement? Or are you claiming that OpenAI owns another model that is clearly better than GPT-5?
- senko 1y agoAh so your reading of that statement is "our best, which we acknowledge is not THE best, but it's not supposed to, it's supposed to be cost effective"? I would suggest reading the entire comment thread before attacking people.
- mat0 1y agoNot the person you were responding to but, if a company provides a service, they want you to use it instead of their competitors. No company is going to say “use ours unless you want to use the best, then use our competitor’s”. so even though I agree with you that they are not explicitly saying “this is the best model in the world”, they are definitely saying “hey this is the best we got, use it”.
- fouc 1y agoI was going by what sama was saying on twitter. he was mostly hyping the cost-effectiveness of it. which can be considered a factor what is "best" too.
- sergiotapia 1y agoYou compare what can be used by most engineers. Most engineers are not going to spend that insane price of Opus. It's extremely high compared to all other models, so even if it is slightly better, it's a non-starter for engineering workloads.
- andsoitis 1y ago> t insane price of Opus I believe Opus starts at $20 a month, similar to GPT5 if you want more than just cursory usage. Or am I missing something?
- sergiotapia 1y agoYes you are missing something: Claude Opus 4.1 Most intelligent model for complex tasks Input $15 / MTok Output $75 / MTok Prompt caching Write $18.75 / MTok Read $1.50 / MTok
- andsoitis 1y agoI see. Do you know of a resource that does an across-the board apples-to-apples comparison between the different services (knowing they all price slightly differently)? It would be useful to be able to easily compare what it costs across the big providers: Gemini, Grok, Claude, ChatGPT.
- inquirerGeneral 1y ago[dead]
- michaelt 1y agoFor $20/month you get Opus in-browser chat access, and Sonnet claude code access. If you want to use Opus in claude code, you've got to get the $100/month plan - or pay API prices. And agentic coding uses a lot of tokens.
- markbao 1y agoMost engineers spending their own money maybe, but the cost of Opus is not that much compared to the output when the company is paying for it.
- deleted 1y ago[deleted]
- runako 1y agore: the comments that Opus is not cost effective...The whole sales pitch behind these tools, and quite specifically the pitch OpenAI made yesterday, is that they will replace people, specifically programmers. Opus is cheaper than a US-based engineer. It's totally reasonable to use it as the benchmark if it's best. Also keep in mind that many employees are not paying out of pocket for LLM use at work. A $1,000 monthly bill for LLM usage is high for an individual but not so much for a company that employees engineers.
- michaelt 1y agoMy experience with coding agents is they need a lot of hand-holding. They're impressive despite that. But if Sonnet is $20/month and I have to intervene every 3 minutes, while Opus is $100/month and I have to intervene every 5 minutes? ¯\_(ツ)_/¯
- runako 1y agoReally depends on who's paying the bill, and how much gets done between interventions, right? Inverting the problem, one might ask how best to spend (say) $5,000 monthly on coding agents. I don't know the answer to that.
- epolanski 1y ago> My experience with coding agents is they need a lot of hand-holding. So do engineers. The difference is that IRL engineers know a lot about the context of the business, features, product, ux, stakeholders, expectations, etc, etc which means that the hand-holding is a long running process. LLMs need all of these things to be clearly written down and specified in one shot.
- nearbuy 1y agoFor what it's worth, I've been trying Opus 4.1 in VS Code through GitHub Copilot and it's been really bad. Maybe worse than Sonnet and GPT 4.1. I'm not sure why it was doing so poorly. In one instance, I asked it to optimize a roughly 80 line C# method that matches some object positions by object ID and delta encodes their positions from the previous frame. It seemed to be confused about how all this should work and output completely wrong code. It has all the context it needs in the file and the method is fairly self-contained. Other models did much better. GPT-5 understood what to do immediately. I tried a few other tasks/questions that also had underwhelming results. Now I've switched to using GPT-5. If you have a quick prompt you'd like me to try, I can share the results.
- cpursley 1y agoUse Claude Code, the rest aren't worth the bother.
- addandsubtract 1y agoWhat does Claude Code do differently to Copilot Agent? Shouldn't they produce the same(ish) result if they're using the same model?
- DannyBee 1y agoIf they prompt the same and ..., They should. But they definitely don't taking into account whatever prompts the tools are really using (or ms is using a neutered version to reduce cost). So I would agree with the suggestion. Using sonnet through copilot seems very very different than cursor or cline or Claude code. Using the same exact model, Copilot consistently often fails to finish tasks or makes a mess. It is consistent at this across ides (ie using the jetbrains plugin generates nearly identical bad results as vscode copilot). I then discard all it did and try the exact same (user) prompt in cursor or Claude code or cline with the same model and it does the same task perfectly.
- akmarinov 1y agoCopilot sucks more at applying what the model is instructing it to do
- intellectronica 1y agoOpus costs 10X more. Maybe it's better, but I can't afford to use it, so who cares.