8 ms·
Between Opus aand GPT-5, it's not clear there's a substantial difference in software development expertise. The metric that I can't seem to get past in my attem
by aliljet 1y ago
Between Opus aand GPT-5, it's not clear there's a substantial difference in software development expertise. The metric that I can't seem to get past in my attempts to use the systems is context awareness over long-running tasks. Producing a very complex, context-exceeding objective is a daily (maybe hourly) ocurrence for me. All I care about is how these systems manage context and stay on track over extended periods of time.
What eval is tracking that? It seems like it's potentially the most imporatnt metric for real-world software engineering and not one-shot vibe prayers.
- realusername 1y agoPersonally I think I'll wait for another 10x improvement for coding because with the current way it's going, they clearly need that.
- fsloth 1y agoFrom my experience when used through IDE such as Cursor the current gen Claude model enables impressive speedruns over commodity tasks. My context is a CAD application I’ve been writing as a hobby. I used to work in that field for a decade so have a pretty good touch on how long I would expect tasks to take. I’m using mostly a similar software stack as that at previous job and am definetly getting stuff done much faster on holiday at home than at that previous work. Of course the codebase is also a lot smaller, intrinsic motivation, etc, but still.
- BoredPositron 1y agoHow often do you have to build the simple scaffolding though?
- fsloth 1y agoAt a real job? Not that often! And it's miserable in large scale architecture. However, at leas for me there is lots of "small enough context" boilerplate that the context can deal with. Clearly this is not a tool in the sense it's predictable.
- realusername 1y agoI've done pretty much the same as you (Cursor/Claude) for our large Rails/React codebase at work and the experience has been horrific so far, I reverted back to vscode.
- fsloth 1y agoYeah! It's quite possible my scenario is in the "happy accident" valley. I'm using it mostly for C#, WPF and OpenTK. The type system seems to help a lot. The UI logic it recommends is mostly god awful. But at least for me when it's given a pattern it can apply, it does so pretty well.
- bdangubic 1y agocontext awareness over long-running tasks don’t have long-running tasks, llms or not. break the problem down into small manageable chunks and then assemble it. neither humans nor llms are good at long-running tasks.
- beoberha 1y agoA series of small manageable chunks becomes a long running task :) If LLMs are going to act as agents, they need to maintain context across these chunks.
- bastawhiz 1y ago> neither humans nor llms are good at long-running tasks. That's a wild comparison to make. I can easily work for an hour. Cursor can hardly work for a continuous pomodoro. "Long-running" is not a fixed size.
- echelon 1y agoHumans can error correct. LLMs multiply errors over time.
- bdangubic 1y agoI just finished my workday, 8hrs with Claude Code. No single task took more than 20 minutes total. Cleared context after each task and asked it to summarize for itself the previous task before I cleared context. If I ran this as a continuous 8hr task it would have died after 35-ish minutes. Just know the limitations (like with any other tool) and you’ll be good :)
- 0x457 1y agoI always find it wild that none of these tools use VCS - completed logical unit of work, make a commit, drop entire context related to that commit, while referencing said commit, continue onto the next stage, rinse and repeat. Claud always misunderstands how API exported by my service works and every compaction it forgets all over and commits "oh api has changed since last time I've used, let me use different query parameters", my brother Christ nothing has changed, and you are the one who made this API.
- swader999 1y agoIf GPT 5 truly has 400k context, that might be all it needs to meaningfully surpass Opus.
- AS04 1y ago400k context with 100% on the fiction livebench would make GPT-5 the undisputably best model IMHO. Don't think it will achieve that though, sadly.
- simonw 1y agoIt's 272,000 input tokens and 128,000 output tokens.
- zurfer 1y agoWoah that's really kind of hidden. But I think you can specify max output tokens. Need to test that!
- 6thbit 1y agoOh, I had not grasped that the “context window” size advertised had to include both input and output. But is it really 272k even if the output was say 10k? Cause it does say “max output” in the docs, so I wonder
- simonw 1y agoThis is the only model where the input limit and the context limit are different values. OpenAI docs team are working on updating that page.
- dudeinhawaii 1y agoThe website clearly lays them out as 400k input and 128k output [1]. I just updated my AI apps to support the new models. I routinely fill the entire context on large code calls. Input is not a "shared" context. I found 100k was barely enough for a single project without spillover, so 4x allows for linking more adjacent codebases for large scale analysis. [1] https://platform.openai.com/docs/models/gpt-5 https://platform.openai.com/docs/models/gpt-5
- logicchains 1y ago>Between Opus aand GPT-5, it's not clear there's a substantial difference in software development expertise. If there's no substantial difference in software development expertise then GPT-5 absolutely blows Opus out of the water due to being almost 10x cheaper.
- spiderice 1y agoDoes OpenAI provide a $200/month option that lets me use as much GPT-5 I want inside of Codex? Because if not, I'd still go with Opus + Claude Code. I'd rather be able to tell my employer, "this will cost you $200/month" than "this might cost you less than $200/month, but we really don't know because it's based on usage"
- mh- 1y agoTo be clear, Claude doesn't provide that either. You can get "usage limited" off of Opus on the $200/mo plan.
- konarkm 1y agoThe ChatGPT paid subscriptions now come with Codex CLI usage included
- t1amat 1y agoIs this actually true? Last I checked (a week ago?) Codex the agents were free at some tiers in a preview capacity (with future rate limits based on tier), but codex cli was not. With codex cli you can log in but the purpose of that is to link it to an API key where you pay per use. The sub tiers give one time credits you would burn through quickly.
- Deradon 1y agoFound this in the GPT-5 Announcement: > Availability and access > GPT‑5 is starting to roll out today to all Plus, Pro, Team, and Free users, with access for Enterprise and Edu coming in one week. Pro, Plus, and Team users can also start coding with GPT‑5 in the Codex CLI (opens in a new window) by signing in with ChatGPT.
- nadis 1y agoIt's pretty vague, but the OP had this callout: >"GPT‑5 is the strongest coding model we’ve ever released. It outperforms o3 across coding benchmarks and real-world use cases, and has been fine-tuned to shine in agentic coding products like Cursor, Windsurf, GitHub Copilot, and Codex CLI. GPT‑5 impressed our alpha testers, setting records on many of their private internal evals."
- RobinL 1y agoTotally agree. At the moment I find that frontier LLMs are able to solve most of the problems I throw at them given enough context. Most of my time is spent working out what context they're missing when they fail. So the thing that would help me most is much a much more focussed ability to gather context. For my use cases, this is mostly needing to be really home in on relevant code files, issues, discussions, PRs. I'm hopeful that GPT5 will be a step forward in this regard that isn't fully captured in the benchmark results. It's certainly promising that it can achieve similar results more cheaply than e.g. Opus.
- deleted 1y ago[deleted]
- abossy 1y agoAt my company (Charlie Labs), we've had a tremendous amount of success with context awareness over long-running tasks with GPT-5 since getting access a few weeks ago. We ran an eval to solve 10 real Github issues so that we could measure this against Claude Code and the differences were surprisingly large. You can see our write-up here: https://charlielabs.ai/research/gpt-5 https://charlielabs.ai/research/gpt-5 Often, our tasks take 30-45 minutes and can handle massive context threads in Linear or Github without getting tripped up by things like changes in direction part of the way through the thread. While 10 issues isn't crazy comprehensive, we found it to be directionally very impressive and we'll likely build upon it to better understand performance going forward.
- bartman 1y agoI am not (usually) photosensitive, but the animated static noise on your websites causes noticable flickering on various screens I use and made it impossible for me to read your article. For better accessibility and a safer experience[1] I would recommend not animating the background, or at least making it easily togglable. [1] https://developer.mozilla.org/en-US/docs/Web/Accessibility/Guides/Seizure_disorders https://developer.mozilla.org/en-US/docs/Web/Accessibility/G...
- MPSFounder 1y agoI concur. Awful UI
- neom 1y agoRemoved- sorry, and thank you for the feedback.
- joshmlewis 1y agoI've been testing it against Opus 4.1 the last few hours and it has done better and solved problems Claude kept failing at. I would say it's definitely better, at least so far.
- cyanydeez 1y agoreal context is a graph of objectives and results. The power of these models has peaked and simply arn't going to manage the type of awareness being promised.
- 1659447091 1y ago> Producing a very complex, context-exceeding objective is a daily (maybe hourly) ocurrence for me. All I care about is how these systems manage context and stay on track over extended periods of time. For whatever reason Github's Copilot is treated like the redheaded stepchild of coding assistants. Even through there are Anthropic, OpenAI, and Google models to choose from. And there is a "spaces"[0] website feature that may be close to what you are looking for. I got better results for testing some larger task using that than I did through the IDE version. But have not used it much. Maybe others have more experience with it. Trying to gather all the context and then review the results was taking longer than doing it myself; having the context gathered already or building it up over time is probably where its value is. [0] https://docs.github.com/en/copilot/concepts/spaces https://docs.github.com/en/copilot/concepts/spaces
- ilaksh 1y agoThe pricing is dramatically better than Opus for gpt-5 since that is now comparable to Gemini 2.5 Pro.
- user3939382 1y agoSorry if this is repetitive but you have to break the problem down just like any complex computing task. The difference is how. You have to break the problems into context windows that you anticipate being able to sow together later. It’s not the same way you would break down a source code authoring task in its absence but the theory is the same.
- altitudinous 1y agoIndeed context awareness is the big difference here, GPT5 is a vast improvement. It doesn't lose track (as easily)
- greymalik 1y ago> it's not clear there's a substantial difference in software development expertise But GPT-5 is substantially cheaper[0]. [0] https://simonwillison.net/2025/Aug/7/gpt-5/#pricing-is-aggressively-competitive https://simonwillison.net/2025/Aug/7/gpt-5/#pricing-is-aggre...
- andhuman 1y agoIs this because it’s now a Moe? They now match price with Gemini 2.5 Pro, which is also a moe.