21 ms·
GPT-5.4
https://openai.com/index/gpt-5-4-thinking-system-card/ https://openai.com/index/gpt-5-4-thinking-system-card/
https://x.com/OpenAI/status/2029620619743219811 https://x.com/OpenAI/status/2029620619743219811
- Aldipower 7mo agoSo did they raised the ridiculous small "per tool call token limit" when working with MCP servers? This makes Chat useless... I do not care, but my users.
- daniel_mercer 7mo ago[flagged]
- OsrsNeedsf2P 7mo agoDoes anyone know what website is the "Isometric Park Builder" shown off here?
- turblety 7mo agoThey build that using GPT-5.4 > Theme park simulation game made with GPT‑5.4 from a single lightly specified prompt GPT literally built that game.
- ignorantguy 7mo agoit shows a 404 as of now.
- minimaxir 7mo agoUp now. The OP has frequently gotten the scoop for new LLM releases and I am curious what their pipeline is.
- Leynos 7mo agoGuess the URL and post at 10 AM PST on the day of release.
- bdangubic 7mo agocurl the URL https://openai.com/index/introducing-gpt-5- https://openai.com/index/introducing-gpt-5-? until you get 200
- mudkipdev 7mo agoProbably refresh the api models list every couple minutes instead. No one could have guessed the name of GPT-Codex-Spark
- mattas 7mo ago"GPT‑5.4 interprets screenshots of a browser interface and interacts with UI elements through coordinate-based clicking to send emails and schedule a calendar event." They show an example of 5.4 clicking around in Gmail to send an email. I still think this is the wrong interface to be interacting with the internet. Why not use Gmail APIs? No need to do any screenshot interpretation or coordinate-based clicking.
- TheAceOfHearts 7mo agoI think the desire is that in the long-term AI should be able to use any human-made application to accomplish equivalent tasks. This email demo is proof that this capability is a high priority.
- spongebobstoes 7mo agonot everything has an API, or API use is limited. some UIs are more feature complete than their APIs some sites try to block programmatic use UI use can be recorded and audited by a non-technical person
- Jacques2Marais 7mo agoI guess a big chunk of their target market won't know how to use APIs.
- satvikpendem 7mo agoThe ideal of REST, the HTML and UI is the API.
- PaulHoule 7mo agoAPIs have never been a gift but rather have always been a take-away that lets you do less than you can with the web interface. It’s always been about drinking through a straw, paying NASA prices, and being limited in everything you can do. But people are intimidated by the complexity of writing web crawlers because management has been so traumatized by the cost of making GUI applications that they couldn’t believe how cheap it is to write crawlers and scrapers…. Until LLMs came along, and changed the perceived economics and created a permission structure. [1] AI is a threat to the “enshittification economy” because it lets us route around it. [1] that high cost of GUI development is one reason why scrapers are cheap… there is a good chance that the scraper you wrote 8 years ago still works because (a) they can’t afford to change their site and (b) if they could afford to change their site changing anything substantial about it is likely to unrecoverably tank their Google rankings so they won’t. A.I. might change the mechanics of that now that you Google traffic is likely to go to zero no matter what you do.
- denysvitali 7mo agoArticle: https://openai.com/index/introducing-gpt-5-4/ https://openai.com/index/introducing-gpt-5-4/ gpt-5.4 Input: $2.50 /M tokens Cached: $0.25 /M tokens Output: $15 /M tokens --- gpt-5.4-pro Input: $30 /M tokens Output: $180 /M tokens Wtf
- elliotbnvl 7mo agoLooks like it's an order of magnitude off. Missprint?
- GenerWork 7mo agoLooks like an extra zero was added?
- benlivengood 7mo agoGovernment pricing :)
- outside2344 7mo ago$30 per kill approval
- glerk 7mo agoLooks like fair price discovery :)
- dpoloncsak 7mo ago[flagged]
- elicash 7mo agoCan't you continue to use to older model, if you prefer the pricing? But they also claim this new model uses fewer tokens, so it still might ultimately be cheaper even if per token cost is higher.
- minimaxir 7mo agoThe marquee feature is obviously the 1M context window, compared to the ~200k other models support with maybe an extra cost for generations beyond >200k tokens. Per the pricing page, there is no additional cost for tokens beyond 200k: https://openai.com/api/pricing/ https://openai.com/api/pricing/ Also per pricing, GPT-5.4 ($2.50/M input, $15/M output) is much cheaper than Opus 4.6 ($5/M input, $25/M output) and Opus has a penalty for its beta >200k context window. I am skeptical whether the 1M context window will provide material gains as current Codex/Opus show weaknesses as its context window is mostly full, but we'll see. Per updated docs (https://developers.openai.com/api/docs/guides/latest-model https://developers.openai.com/api/docs/guides/latest-model), it supercedes GPT-5.3-Codex, which is an interesting move.
- thehamkercat 7mo agoGPT 5.3 codex had 400K context window btw
- simianwords 7mo agoWhy would some one use codex instead?
- embedding-shape 7mo agoWhy would someone use Claude Code instead? Or any other harness? Or why only use one? My own tooling throws off requests to multiple agents at the same time, then I compare which one is best, and continue from there. Most of the time Codex ends up with the best end results though, but my hunch is that at one point that'll change, hence I continue using multiple at the same time.
- surgical_fire 7mo agoI've been using Codex for software development personally (I have a ChatGPT account), and I use Claude at work (since it is provided by my employer). I find both Codex and Claude Opus perform at a similar level, and in some ways I actually prefer Codex (I keep hitting quota limits in Opus and have to revert back to Sonnet). If your question is related to morality (the thing about US politics, DoD contract and so on)... I am not from the US, and I don't care about its internal politics. I also think both OpenAI and Anthropic are evil, and the world would be better if neither existed.
- iamronaldo 7mo agoNotably 75% on os world surpassing humans at 72%... (How well models use operating systems)
- Chance-Device 7mo agoI’m sure the military and security services will enjoy it.
- varispeed 7mo agoprompt> Hi we want to build a missile, here is the picture of what we have in the yard.
- mirekrusin 7mo ago{ tools: [ { name: "nuke", description: "Use when sure.", ... { lat: number, long: number } } ] }
- theParadox42 7mo agoThe self reported safety score for violence dropped from 91% to 83%.
- skrebbel 7mo agoWhat the hell is a "safety score for violence"?
- chromic04850 7mo ago[dead]
- twtw99 7mo agoIf you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20 https://x.com/OpenAI/status/2029620619743219811?s=20
- chabes 7mo agoDefinitely don’t want to click in at x either.
- thejarren 7mo agoSolution https://xcancel.com/OpenAI/status/2029620619743219811?s=20 https://xcancel.com/OpenAI/status/2029620619743219811?s=20
- anonym00se1 7mo agoDitto, but I did anyways and enjoyed that OpenAI doesn't include the dogwater that is Grok on their scorecard.
- observationist 7mo ago[flagged]
- Sabinus 7mo agoGet a redirect plugin and set it up to send you to xcancel instead of Twitter. I've done it, and it's very convenient.
- karmasimida 7mo agoIt is a bigger model, confirmed
- Aboutplants 7mo agoIt seems that all frontier models are basically roughly even at this point. One may be slightly better for certain things but in general I think we are approaching a real level playing field field in terms of ability.
- jryio 7mo ago1 million tokens is great until you notice the long context scores fall off a cliff past 256K and the rest is basically vibes and auto compacting.
- olliepro 7mo agoI bet they lack good long context training data and need to start a flywheel of collecting it via their api (from willing customers)
- jbergqvist 7mo agoThis would be my guess too. It can probably be generated synthetically or via agentic rollouts, but high quality long context examples where outputs meaningfully depend on long-range interactions probably remain scarce
- rrr_oh_man 7mo agoIt's the same now with Gemini as well. Unfortunately. :(
- minimaxir 7mo agoMore discussion here on the blog post announcement which has been confusingly penalized by Hacker News's algorithm: https://news.ycombinator.com/item?id=47265005 https://news.ycombinator.com/item?id=47265005
- dang 7mo agoThanks. We'll merge the threads, but this time we'll do it hither, to spread some karma love.
- shablulman 7mo ago[flagged]
- ZeroCool2u 7mo agoBit concerning that we see in some cases significantly worse results when enabling thinking. Especially for Math, but also in the browser agent benchmark. Not sure if this is more concerning for the test time compute paradigm or the underlying model itself. Maybe I'm misunderstanding something though? I'm assuming 5.4 and 5.4 Thinking are the same underlying model and that's not just marketing.
- aplomb1026 7mo ago[flagged]
- highfrequency 7mo agoCan you be more specific about which math results you are talking about? Looks like significant improvement on FrontierMath esp for the Pro model (most inference time compute).
- ZeroCool2u 7mo agoFrontier Math, GPQA Diamond, and Browsecomp are the benchmarks I noticed this on.
- csnweb 7mo agoAre you may be comparing the pro model to the non pro model with thinking? Granted it’s a bit confusing but the pro model is 10 times more expensive and probably much larger as well.
- ZeroCool2u 7mo agoAh yes, okay that makes more sense!
- oersted 7mo agoI believe you are looking at GPT 5.4 Pro. It's confusing in the context of subscription plan names, Gemini naming and such. But they've had the Pro version of the GPT 5 models (and I believe o3 and o1 too) for a while. It's the one you have access to with the top ~$200 subscription and it's available through the API for a MUCH higher price ($2.5/$15 vs $30/$180 for 5.4 per 1M tokens), but the performance improvement is marginal. Not sure what it is exactly, I assume it's probably the non-quantized version of the model or something like that.
- egonschiele 7mo agoThe actual card is here https://deploymentsafety.openai.com/gpt-5-4-thinking/introduction https://deploymentsafety.openai.com/gpt-5-4-thinking/introdu... the link currently goes to the announcement.
- Rapzid 7mo agoI must have been sleeping when "sheet" "brief" "primer" etc become known as "cards". I really thought weirdly worded and unnecessary "announcement" linking to the actual info along with the word "card" were the results of vibe slop.
- draw_down 7mo ago[dead]
- realityfactchex 7mo agoCard is slightly odd naming indeed. Criticisms aside (sigh), according to Wikipedia, the term was introduced when proposed by mostly Googlers, with the original paper [0] submitted in 2018. To quote, """In this paper, we propose a framework that we call model cards, to encourage such transparent model reporting. Model cards are short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions, such as across different cultural, demographic, or phenotypic groups (e.g., race, geographic location, sex, Fitzpatrick skin type [15]) and intersectional groups (e.g., age and race, or sex and Fitzpatrick skin type) that are relevant to the intended application domains. Model cards also disclose the context in which models are intended to be used, details of the performance evaluation procedures, and other relevant information.""" So that's where they were coming from, I guess. [0] Margaret Mitchell et al., 2018 submission, Model Cards for Model Reporting, https://arxiv.org/abs/1810.0399 https://arxiv.org/abs/1810.0399
- Murfalo 7mo agoTo me, model card makes sense for something like this https://x.com/OpenAI/status/2029620619743219811 https://x.com/OpenAI/status/2029620619743219811. For "sheet"/"brief"/"primer" it is indeed a bit annoying. I like to see the compiled results front and center before digging into a dossier.
- nickysielicki 7mo agocan anyone compare the $200/mo codex usage limits with the $200/mo claude usage limits? It’s extremely difficult to get a feel for whether switching between the two is going to result in hitting limits more or less often, and it’s difficult to find discussion online about this. In practice, if I buy $200/mo codex, can I basically run 3 codex instances simultaneously in tmux, like I can with claude code pro max, all day every day, without hitting limits?
- ritzaco 7mo agoI haven't tried the $200 plans by I have Claude and Codex $20 and I feel like I get a lot more out of Codex before hitting the limits. My tracker certainly shows higher tokens for Codex. I've seen others say the same.
- lostmsu 7mo agoSadly comment ratings are not visible on HN, so the only way to corroborate is to write it explicitly: Codex $20 includes significantly more work done and is subjectively smarter.
- winstonp 7mo agoAgree. Claude tends to produce better design, but from a system understanding and architecture perspective Codex is the far better model
- vtail 7mo agoMy own experience is that I get far far more usage (and better quality code, too) from codex. I downgrade my Claude Max to Claude Pro (the $20 plan) and now using codex with Pro plan exclusively for everything.
- Marciplan 7mo agoCodex announced at 5.3 launch that until April all usage limits are upped so take that into account
- strongpigeon 7mo agoIt's interesting that they charge more for the > 200k token window, but the benchmark score seems to go down significantly past that. That's judging from the Long Context benchmark score they posted, but perhaps I'm misunderstanding what that implies.
- simianwords 7mo ago[flagged]
- strongpigeon 7mo agoI guess that you pay more for worse quality to unlock use cases that could maybe be solved by better context management.
- Tiberium 7mo agoThey don't actually seem to charge more for the >200k tokens on the API. OpenRouter and OpenAI's own API docs do not have anything about increased pricing for >200k context for GPT-5.4. I think the 2x limit usage for higher context is specific to using the model over a subscription in Codex.
- _heitoo 7mo agoIt makes sense in scenarios where a model needs >200k tokens to answer a single prompt. You're shackled to a single session, and if the model hits compaction limits, it'll get lobotomized and give a shitty answer, so higher limits, even with degraded performance, are still an improvement.
- tmpz22 7mo agoDoes this improve Tomahawk Missile accuracy?
- ch4s3 7mo agoThey're already accurate within 5-10m at Mach 0.74 after traveling 2k+ km. Its 5m long so it seems pretty accurate. How much more could you expect?
- mikkupikku 7mo agoYou could definitely do better than that with image recognition for terminal guidance. But I would assume those published accuracy numbers are very conservative anyway..
- keithnz 7mo agoI think for LLM like Open AI, it wouldn't be about hitting the target but target selection. Target selection is probably the most likely thing that won't be accurate
- simianwords 7mo agoWhat is the point of gpt codex?
- catketch 7mo ago-codex variant models in earlier version were just fine tuned for coding work, and had a little better performance for related tool calling and maybe instruction calling. in 5.4 it looks like the just collapsed that capability into the single frontier family model
- simianwords 7mo agoYes so I’m even more confused. Why would I use codex?
- akmarinov 7mo agoThey’ll likely come out with a 5.4-Codex at some point, that’s what they did with 5 and 5.2
- ilaksh 7mo agoRemember when everyone was predicting that GPT-5 would take over the planet?
- nthypes 7mo ago$30/M Input and $180/M Output Tokens is nuts. Ridiculous expensive for not that great bump on intelligence when compared to other models.
- rvz 7mo agoYou didn't realize they can increase / change prices for intelligence? This should not be shocking.
- nickthegreek 7mo agoOP made no mention of not understanding cost relation to intelligence. In fact, they specifically call out the lack of value.
- moralestapia 7mo agoDon't use it?
- nthypes 7mo agoGemini 3.1 Pro $2/M Input Tokens $15/M Output Tokens Claude Opus 4.6 $5/M Input Tokens $25/M Output Tokens
- nthypes 7mo agoJust to clarify,the pricing above is for GPT-5.4 Pro. For standard here is the pricing: $2.5/M Input Tokens $15/M Output Tokens
- energy123 7mo agoFor Pro
- joe_mamba 7mo agoBetter tokens per dollar could be useless for comparison if the model can't solve your problem.
- stri8ted 7mo ago
- world2vec 7mo agoBenchmarks barely improved it seems
- chromic04850 7mo ago[dead]
- cj 7mo agoI use ChatGPT primarily for health related prompts. Looking at bloodwork, playing doctor for diagnosing minor aches/pains from weightlifting, etc. Interesting, the "Health" category seems to report worse performance compared to 5.2.
- paxys 7mo agoModels are being neutered for questions related to law, health etc. for liability reasons.
- cj 7mo agoI'm sometimes surprised how much detail ChatGPT will go into without giving any dislaimers. I very frequently copy/paste the same prompts into Gemini to compare, and Gemini often flat out refuses to engage while ChatGPT will happily make medical recommendations. I also have a feeling it has to do with my account history and heavy use of project context. It feels like when ChatGPT is overloaded with too much context, it might let the guardrails sort of slide away. That's just my feeling though. Today was particularly bad... I uploaded 2 PDFs of bloodwork and asked ChatGPT to transcribe it, and it spit out blood test results that it found in the project context from an earlier date, not the one attached to the prompt. That was weird.
- bargainbin 7mo agoAnecdotal, but I asked Claude the other day about how to dilute my medication (HCG) and it flat out refused and started lecturing me about abusing drugs. I copy and pasted into ChatGPT, it told me straight away, and then for a laugh said it was actually a magical weight loss drug that I'd bought off the dark web... And it started giving me advice about unregulated weight loss drugs and how to dose them.
- staticman2 7mo agoIf you had created a project with custom instructions and/ or custom style I think you could have gotten Claude to respond the way you wanted just fine.
- deleted 7mo ago[deleted]
- wahnfrieden 7mo agoNo Codex model yet
- minimaxir 7mo agoGPT-5.4 is the new Codex model.
- deleted 7mo ago[deleted]
- wahnfrieden 7mo agoFinally
- nico1207 7mo agoGPT-5.3-Codex is superior to GPT-5.4 in Terminal Bench with Codex, so not really
- conradkay 7mo agoGeneral consensus seems to be that it's still a better coding model, overall
- koakuma-chan 7mo agoIt just released, how is there a general consensus already
- wahnfrieden 7mo agosome non-employees have been using it for a while already
- timpera 7mo ago> Steerability: Similarly to how Codex outlines its approach when it starts working, GPT‑5.4 Thinking in ChatGPT will now outline its work with a preamble for longer, more complex queries. You can also add instructions or adjust its direction mid-response. This was definitely missing before, and a frustrating difference when switching between ChatGPT and Codex. Great addition.
- yanis_t 7mo agoThese releases are lacking something. Yes, they optimised for benchmarks, but it’s just not all that impressive anymore. It is time for a product, not for a marginally improved model.
- esafak 7mo agoThat's for you to build; they provide the brains. Do you really want one company to build everything? There wouldn't be a software industry to speak of if that happened.
- simlevesque 7mo agoNah, the second you finish your build they release their version and then it's game over.
- acedTrex 7mo agoWell they are currently the ones valued at a number with a whole lotta 0s on it. I think they should probably do both
- ipsum2 7mo agoThe model was released less than an hour ago, and somehow you've been able to form such a strong opinion about it. Impressive!
- cj 7mo agoOne opinion you can form in under an hour is... why are they using GPT-4o to rate the bias of new models? > assess harmful stereotypes by grading differences in how a model responds > Responses are rated for harmful differences in stereotypes using GPT-4o, whose ratings were shown to be consistent with human ratings Are we seriously using old models to rate new models?
- titanomachy 7mo agoWhy not? If they’ve shown that 4o is calibrated to human responses, and they haven’t shown that yet for 5.4…
- prydt 7mo agoI no longer want to support OpenAI at all. Regardless of benchmarks or real world performance.
- Imustaskforhelp 7mo agoI agree with ya. You aren't alone in this. For what its worth, Chatgpt subscriptions have been cancelled or that number has risen ~300% in the last month. Also, Anthropic/Gemini/even Kimi models are pretty good for what its worth. I used to use chatgpt and I still sometimes accidentally open it but I use Gemini/Claude nowadays and I personally find them to be better anyways too.
- throwaway911282 7mo ago[flagged]
- Imustaskforhelp 7mo agoGovt. contracts and terms allowing autonomous drone machines which can kill without any human in the loop have a very large difference I know the difference between this is none but to me, its that Anthropic stood for what it thought was right. It had drew a line even if it may have costed some money and literally have them announced as supply chain and see all the fallout from that in that particular relevant thread. As a person, although I am not fan of these companies in general and yes I love oss-models. But I still so so much appreciate atleast's anthropic's line of morality which many people might seem insignificant but to me it isn't. So for the workflows that I used OpenAI for, I find Anthropic/gemini to be good use. I love OSS-models too btw and this is why I recommended Kimi too.
- Imustaskforhelp 7mo ago> I know the difference between this is none but to me Edit: just a very minor nitpick of my own writing but I meant that "I know the difference between this could look very little to some, maybe none, but to me..." rather than "I know the difference between this is none but to me". I was clearly writing this way too late at night haha. My point sort of was/is that Anthropic drew a line at something and is taking massive losses of supply chain/risks and what not and this is the thing that I would support a company out of rather than say OpenAI.
- beernet 7mo agoSam really fumbled the top position in a matter of months, and spectacularly so. Wow. It appears that people are much more excited by Anthropic and Google releases, and there are good reasons for that which were absolutely avoidable.
- jcmontx 7mo ago5.4 vs 5.3-Codex? Which one is better for coding?
- vtail 7mo agoLooking at the benchmarks, 5.4 is slightly better. But it also offers "Fast" mode (at 2x usage), which - if it works and doesn't completely depletes my Pro plan - is a no brainer at the same or even slightly worse quality for more interactive development.
- esafak 7mo agoFor the price, it seems the latter. I'd use 5.4 to plan.
- embedding-shape 7mo agoLiterally just released, I don't think anyone knows yet. Don't listen to people's confident takes until after a week or two when people actually been able to try it, otherwise you'll just get sucked up in bears/bulls misdirected "I'm first with an opinion".
- awestroke 7mo agoOpus 4.6
- jcmontx 7mo agoCodex surpassed Claude in usefulness _for me_ since last month
- baal80spam 7mo ago[flagged]
- Someone1234 7mo agoRelated question: - Do they have the same context usage/cost particularly in a plan? They've kept 5.3-Codex along with 5.4, but is that just for user-preference reasons, or is there a trade-off to using the older one? I'm aware that API cost is better, but that isn't 1:1 with plan usage "cost."
- gavinray 7mo agoThe "RPG Game" example on the blogpost is one of the most impressive demo's of autonomous engineering I've seen. It's very similar to "Battle Brothers", and the fact that RPG games require art assets, AI for enemy moves, and a host of other logical systems makes it all the more impressive.
- hungryhobbit 7mo ago[flagged]
- OsrsNeedsf2P 7mo agoLow quality off-topic comment. It's not murder when they're American soldiers.
- squibonpig 7mo agoMurder in spirit if not by the letter
- hungryhobbit 7mo agoYou have a strange (and cruel) definition of murder. I like the dictionary one better: "the unlawful premeditated killing of one human being by another." Wars have laws (ever heard of "war crimes"?) Soldiers can absolutely commit murder.
- hu3 7mo agoindeed and I suspect it can be attributed to, at least in part, the improved playwright integration. > we’re also releasing an experimental Codex skill called “Playwright (Interactive) (opens in a new window)”. This allows Codex to visually debug web and Electron apps; it can even be used to test an app it’s building, as it’s building it.
- casid 7mo agoI don't know. It looks shallow and simple, not even a demo.
- 7mo ago
- swingboy 7mo agoEven with the 1m context window, it looks like these models drop off significantly at about 256k. Hopefully improving that is a high priority for 2026.
- leftbehinds 7mo agosome sloppy improvements
- HardCodedBias 7mo agoWe'll have to wait a day or two, maybe a week or two, to determine if this is more capable in coding than 5.3, which seems to be the economically valuable capability at this time. In terms of writing and research even Gemini, with a good prompt, is close to useable. That's likely not a differentiator.
- deleted 7mo ago[deleted]
- lostmsu 7mo agoWhat is Pro exactly and is it available in Codex CLI?
- nickandbro 7mo agoBeat Simon Willison ;) https://www.svgviewer.dev/s/gAa69yQd https://www.svgviewer.dev/s/gAa69yQd Not the best pelican compared to gemini 3.1 pro, but I am sure with coding or excel does remarkably better given those are part of its measured benchmarks.
- GaggiX 7mo agoThis pelican is actually bad, did you use xhigh?
- nickandbro 7mo agoyep, just double checked used gpt-5.4 xhigh. Though had to select it in codex as don't have access to it on the chatgpt app or web version yet. It's possible that whatever code harness codex uses, messed with it.
- nubg 7mo agothis is proof they are not benchmaxxing the pelican's :-)
- bazmattaz 7mo agoAnyone else feel that it’s exhausting keeping up with the pace of new model releases. I swear every other week there’s a new release!
- coffeemug 7mo agoWhy do you need to keep up? Just use the latest models and don't worry about it.
- throwup238 7mo agoYes, that's a common feeling. 5.3-Codex was released a month ago on Feb 5 so we're not even getting a full month within a single brand, let alone between competitors.
- davnicwil 7mo agoIf you think about it there shouldn't really be a reason to care as long as things don't get worse. Presumably this is where it'll evolve to with the product just being the brand with a pricing tier and you always get {latest} within that, whatever that means (you don't have to care). They could even shuffle models around internally using some sort of auto-like mode for simpler questions. Again why should I care as long as average output is not subjectively worse. Just as I don't want to select resources for my SaaS software to use or have that explictly linked to pricing, I don't want to care what my OpenAI model or Anthropic model is today, I just want to pay and for it to hopefully keep getting better but at a minimum not get worse.
- pupppet 7mo agoI think it's fun, it's like we're reliving the browser wars of the early days.
- oytis 7mo agoEveryone is mindblown in 3...2...1
- dandiep 7mo agoAnyone know why OpenAI hasn't released a new model for fine tuning since 4.1? It'll be a year next month since their last model update for fine tuning.
- qoez 7mo agoI think they just did that because of the energy around it for open source models. Their heart probably wasn't in it and the amount of people fine tuning given the prices were probably too low to continue putting in attention there.
- zzleeper 7mo agoFor me the issue is why there's not a new mini since 5-mini in August. I have now switched web-related and data-related queries to Gemini, coding to Claude, and will probably try QWEN for less critical data queries. So where does OpenAI fits now?
- Rapzid 7mo agoAlso interested in this and a replacement for 4.1/4.1-mini that focuses on low latency and high accuracy for voice applications(not the all-in-one models).
- leftbehinds 7mo ago[flagged]
- paxys 7mo ago"Here's a brand new state-of-the-art model. It costs 10x more than the previous one because it's just so good. But don't worry, if you don't want all this power you can continue to use the older one." A couple months later: "We are deprecating the older model."
- OutOfHere 7mo agoThat's a misrepresentation of the cost. It is simply false. The cost is noted here: https://news.ycombinator.com/item?id=47265144 https://news.ycombinator.com/item?id=47265144
- OutOfHere 7mo agoWhat is with the absurdity of skipping "5.3 Thinking"?
- vicchenai 7mo ago[dead]
- 7777777phil 7mo ago83% win rate over industry professionals across 44 occupations. I'd believe it on those specific tasks. Near-universal adoption in software still hasn't moved DORA metrics. The model gets better every release. The output doesn't follow. Just had a closer look on those productivity metrics this week: https://philippdubach.com/posts/93-of-developers-use-ai-coding-tools.-productivity-hasnt-moved./ https://philippdubach.com/posts/93-of-developers-use-ai-codi...
- NiloCK 7mo agoThis March 2026 blog post is citing a 2025 study based on Sonnet 3.5 and 3.7 usage. Given that organization who ran the study [1] has a terrifying exponential as their landing page, I think they'd prefer that it's results are interpreted as a snapshot of something moving rather than a constant. [1] - https://metr.org/ https://metr.org/
- 7777777phil 7mo agoGood catch, thanks (I really wrote that myself.) Added a note to the post acknowledging the models used were Claude 3.5 and 3.7 Sonnet.
- twitchard 7mo agoNot sure DORA is that much of an indictment. For "Change Failure Rate" for instance these are subject to tradeoffs. Organizations likely have a tolerance level for Change Failure Rate. If changes are failing too often they slow down and invest. If changes aren't failing that much they speed up -- and so saying "change failure rate hasn't decreased, obviously AI must not be working" is a little silly. "Change Lead Time" I would expect to have sped up although I can tell stories for why AI-assisted coding would have an indeterminate effect here too. Right now at a lot of orgs, the bottle neck is the review process because AI is so good at producing complete draft PRs quickly. Because reviews are scarce (not just reviews but also manual testing passes are scarce) this creates an incentive ironically to group changes into larger batches. So the definition of what a "change" is has grown too.
- rbitar 7mo agoI think the most exciting change announced here is the use of tool search to dynamically load tools as needed: https://developers.openai.com/api/docs/guides/tools-tool-search https://developers.openai.com/api/docs/guides/tools-tool-sea...
- DonsDiscountGas 7mo agoI'm pretty sure Claude has had this via skills for awhile
- alpineman 7mo agoNo thanks. Already cancelled my sub.
- kotevcode 7mo ago[flagged]
- iamleppert 7mo agoI wouldn't trust any of these benchmarks unless they are accompanied by some sort of proof other than "trust me bro". Also not including the parameters the models were run at (especially the other models) makes it hard to form fair comparisons. They need to publish, at minimum, the code and runner used to complete the benchmarks and logs. Not including the Chinese models is also obviously done to make it appear like they aren't as cooked as they really are.
- deleted 7mo ago[deleted]
- elmean 7mo ago[flagged]
- timedude 7mo agoAbsolutely amazing. Grateful to be living in this timeframe
- bramhaag 7mo agoWhat makes you think that they see bombing civilians as a bug, not a feature?
- elmean 7mo agofirst real comment, I thought that at first but this could lower the possible users that could be using chatGPT and that would be against us (shareholders)
- skilltissue 7mo agoDon't use the site this way. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- patcon 7mo agoNot all rule-following is noble or wise.
- Chance-Device 7mo agoYou made a burner account just to scold this guy? Don’t use burner accounts this way.
- himata4113 7mo agonews guidelines
- adamtaylor_13 7mo agoParlay?
- jeff_antseed 7mo ago[dead]
- creamyhorror 7mo agoI've only used 5.4 for 1 prompt (edit: 3@high now) so far (reasoning: extra high, took really long), and it was to analyse my codebase and write an evaluation on a topic. But I found its writing and analysis thoughtful, precise, and surprisingly clearly written, unlike 5.3-Codex. It feels very lucid and uses human phrasing. It might be my AGENTS.md requiring clearer, simpler language, but at least 5.4's doing a good job of following the guidelines. 5.3-Codex wasn't so great at simple, clear writing.
- irishcoffee 7mo ago> It might be my AGENTS.md requiring clearer, simpler language If you gave the exact same markdown file to me and I posted ed the exact same prompts as you, would I get the same results?
- m3kw9 7mo agoyou probably can't and asking agents.md to "make it clearer" will likely give you the illusion of clearer language without actual well structured tests. agents.md is to usually change what the llm should focus on doing more that suits you. Not to say stuff like "be better", "make no mistakes"
- creamyhorror 7mo agoI'm not sure if the model (under its temperature/other settings) produces deterministic responses. But I do think models' style and phrasing are fairly changeable via AGENTS.md-style guidelines. 5.4's choice of terms and phrasing is very precise and unambiguous to me, whereas 5.3-Codex often uses jargon and less precise phrases that I have to ask further about or demand fuller explanations for via AGENTS.md.
- irishcoffee 7mo agoSo sharing markdown files is functionally useless, or no?
- 7mo ago
- deleted 7mo ago[deleted]
- XCSme 7mo agoSeems to be quite similar to 5.3-codex, but somehow almost 2x more expensive: https://aibenchy.com/compare/openai-gpt-5-4-medium/openai-gpt-5-3-codex-medium/openai-gpt-5-2-medium/ https://aibenchy.com/compare/openai-gpt-5-4-medium/openai-gp...
- motbus3 7mo agoSam Altman can keep his model intentionally to himself. Not doing business with mass murderers
- smoody07 7mo agoSurprised to see every chart limited to comparisons against other OpenAI models. What does the industry comparison look like?
- aydyn 7mo agoThey compare to Claude and Gemini in their tweet
- 0123456789ABCDE 7mo agohttps://artificialanalysis.ai https://artificialanalysis.ai should have the numbers soon
- lorenzoguerra 7mo agoI believe that this choice is due to two main reasons. First, it's (obviously) a marketing strategy to keep the spotlight on their own models, showing they're constantly improving and avoiding validating competitors. Second, since the community knows that static benchmarks are unreliable, it makes sense for them to outsource the comparisons to independent leaderboards, which lets them avoid accusations of cherry-picking while justifying their marketing strategy. Ultimately, the people actually interested in the performance of these models already don't trust self-reported comparisons and wait for third-party analysis anyway
- throwaway911282 7mo agohttps://xcancel.com/OpenAI/status/2029620619743219811 https://xcancel.com/OpenAI/status/2029620619743219811 you can see comparisons here
- jstummbillig 7mo agoInline poll: What reasoning levels do you work with? This becomes increasingly less clear to me, because the more interesting work will be the agent going off for 30mins+ on high / extra high (it's mostly one of the two), and that's a long time to wait and an unfeasible amount of code to a/b
- newtwilly 7mo agoFor directed coding (implementing an already specified plan) or asking questions about a codebase I use 5.3 codex with medium reasoning effort. It is relatively quick feeling. I like Sonnet 4.6 a lot too at medium reasoning effort, but at least in Cursor it is sometimes quite slow because it will start "thinking" for a long time.
- bob1029 7mo agoI was just testing this with my unity automation tool and the performance uplift from 5.2 seems to be substantial.
- koakuma-chan 7mo agoAnyone else getting artifacts when using this model in Cursor? numerusformassistant to=functions.ReadFile մեկնաբանություն 天天爱彩票网站json {"path":
- deleted 7mo ago[deleted]
- mike_hearn 7mo agoI've seen that problem with 5.3-codex too, it didn't happen with earlier models. Looks like some kind of encoding misalignment bug. What you're seeing is their Harmony output format (what the model actually creates). The Thai/Chinese characters are special tokens apparently being mismapped to Unicode. Their servers are supposed to notice these sequences and translate them back to API JSON but it isn't happening reliably.
- ValentineC 7mo agoI just got some interesting artifacts in Codex when I tried to oneshot a conference page design (my version of the pelican riding a bicycle). GPT-5.4 added some weird guidance that I wouldn't normally expect to see as a normal page visitor.
- daft_pink 7mo agoI’ve officially got model fatigue. I don’t care anymore.
- zeeebeee 7mo agosame same same
- postalrat 7mo agoI'd suggest not clicking for things you don't care about.
- morgengold 7mo agoHave fun with sonnet 3.5
- hmokiguess 7mo agoThey hired the dude from OpenClaw, they had Jony Ive for a while now, give us something different!
- kgeist 7mo ago>Today, we’re releasing <..> GPT‑5.3 Instant >Today, we’re releasing GPT‑5.4 in ChatGPT (as GPT‑5.4 Thinking), >Note that there is not a model named GPT‑5.3 Thinking They held out for eight months without a confusing numbering scheme :)
- gallerdude 7mo agoTbf there was a 5.3 codex
- XCSme 7mo agoWhat I'm most confused, is why call it both GPT-5.3 Instant and gpt-5.3-chat?
- m3kw9 7mo agoinstant kind of suck if you asking more than summerizations, surface info, web searches, it can lose track of who's who quickly in some complex multi turn asks. Just need to know what to use instant for.
- deleted 7mo ago[deleted]
- deleted 7mo ago[deleted]
- __jl__ 7mo agoWhat a model mess! OpenAI now has three price points: GPT 5.1, GPT 5.2 and now GPT 5.4. There version numbers jump across different model lines with codex at 5.3, what they now call instant also at 5.3. Anthropic are really the only ones who managed to get this under control: Three models, priced at three different levels. New models are immediately available everywhere. Google essentially only has Preview models! The last GA is 2.5. As a developer, I can either use an outdated model or have zero insurances that the model doesn't get discontinued within weeks.
- arthurcolle 7mo agoThere is a lot of opportunity here for the AI infrastructure layer on top of tier-1 model providers
- motoxpro 7mo agoThis is what clouds like AWS, Azure, and GCP solve (vertex AI, etc). They are already an abstraction on top of the model makers with distribution built in. I also don't believe there is any value in trying to aggregate consumers or businesses just to clean up model makers names/release schedule. Consumers just use the default, and businesses need clarity on the underlying change (e.g. why is it acting different? Oh google released 3.6)
- arthurcolle 7mo agoDo the end users really care about the models at all, or about the effects that the models can cause?
- delaminator 7mo agotwo great problems in computing naming things cache invalidation off by one errors
- rurban 7mo agoBiggest problem right now in computing: Out of tokens until end of month
- woeirua 7mo agoFeels incremental. Looks like OpenAI is struggling.
- nobody_r_knows 7mo ago[dead]
- throwaway5752 7mo agoDoes this model autonomously kill people without human approval or perform domestic surveillance of US citizens?
- smusamashah 7mo agoI only want to see how it performs on the Bullshit-benchmark https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html https://petergpt.github.io/bullshit-benchmark/viewer/index.v... GPT is not even close yo Claude in terms of responding to BS.
- mistercow 7mo agoMy current hunch is that that benchmark captures most of the relevant gap between Anthropic and the rest. “Can’t distinguish truth from fiction” has long been one of the deeper complaints about LLMs, and the bullshit benchmark seems like a clever approach to testing at least some of that.
- zone411 7mo agoResults from my Extended NYT Connections benchmark: GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6). GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4). GPT-5.4 no reasoning scores 32.8 (GPT-5.2 no reasoning scored 28.1).
- stavros 7mo agoHow do you score this? Losing/winning the game with 4 lives?
- oliwary 7mo agoImpressive! Do you include puzzles released before the training data cutoff date?
- kinderjaje 7mo agoI added that info on https://automatio.ai/models/gpt-5-4 https://automatio.ai/models/gpt-5-4
- consumer451 7mo agoI am very curious about this: > Theme park simulation game made with GPT‑5.4 from a single lightly specified prompt, using Playwright Interactive for browser playtesting and image generation for the isometric asset set. Is "Playwright Interactive" a skill that takes screenshots in a tight loop with code changes, or is there more to it?
- hansonw 7mo agoThe skill source is here: https://github.com/openai/skills/blob/main/skills/.curated/playwright-interactive/SKILL.md https://github.com/openai/skills/blob/main/skills/.curated/p... $skill-installer playwright-interactive in Codex! the model writes normal JS playwright code in a Node REPL
- consumer451 7mo agoThanks!
- motza 7mo agoNo doubt this was released early to ease the bad press
- butILoveLife 7mo agoAnyone else completely not interested? Since GPT5, its been cost cutting measure after cost cutting measure. I imagine they added a feature or two, and the router will continue to give people 70B parameter-like responses when they dont ask for math or coding questions.
- machiaweliczny 7mo ago5.2 and 5.3 are strong/best for coding, 5.0 and 5.1 were garbage
- readytion 7mo ago[flagged]
- Philip-J-Fry 7mo agoI find it quite funny how this blog post has a big "Ask ChatGPT" box at the bottom. So you might think you could ask a question about the contents of the blog post, so you type the text "summarise this blog post". And it opens a new chat window with the link to the blog post followed by "summarise this blog post". Only to be told "I can't access external URLs directly, but if you can paste the relevant text or describe the content you're interested in from the page, I can help you summarize it. Feel free to share!" That's hilarious. Does OpenAI even know this doesn't work?
- Aurornis 7mo agoProbably intentional. They don't want open, no-registration endpoints able to trigger the AI into hitting URLs.
- jazzypants 7mo agoBut, why include the non-functional chat box in the article?
- observationist 7mo agoThey're having service issues - ChatGPT on the web is broken for a lot of people. The app is working in android - I'd assume that the rollout hit a hitch and the chatbox in the article would normally work.
- embedding-shape 7mo agoDifferent team "manages" the overall blog than the team who wrote that specific article. At one point, maybe it made sense, then something in the product changed, team that manages the blog never tested it again. Or, people just stopped thinking about any sort of UX. These sort of mistakes are all over the place, on literally all web properties, some UX flows just ends with you at a page where nothing works sometimes. Everything is just perpetually "a bit broken" seemingly everywhere I go, not specific to OpenAI or even the internet.
- Alifatisk 7mo agoSo let me get this straight, OpenAi previously had an issue with LOTS of different models snd versions being available. Then they solved this by introducing GPT-5 which was more like a router that put all these models under the hood so you only had to prompt to GPT-5, and it would route to the best suitable model. This worked great I assume and made the ui for the user comprehensible. But now, they are starting to introduce more of different models again? We got: - GPT-5.1 - GPT-5.2 Thinking - GPT-5.3 (codex) - GPT-5.3 Instant - GPT-5.4 Thinking - GPT-5.4 Pro Who’s to blame for this ridiculous path they are taking? I’m so glad I am not a Chat user, because this adds so much unnecessary cognitive load. The good news here is the support for 1M context window, finally it has caught up to Gemini.
- 361994752 7mo agoi guess you still have the "auto" as an option to route your request
- stainablesteel 7mo ago5 itself might have solved the problem of having too many different models somewhere in the backend
- sothatsit 7mo agoI much prefer this, we can choose based on our use-cases, and people who don’t care can still use Auto.
- wilg 7mo agoWell, they have older ones of course. But the current options actual users see is "Auto" or "Instant (5.3)" or "Thinking (5.4)". Not that complicated really.
- applfanboysbgon 7mo agoThe real problem that OpenAI had was that their model naming was completely incomprehensible. 4.5, o3, 4o, 4.1 which is newer than 4.5. It was a complete clusterfuck. The blowback on that issue seems to have led them to misidentify the issue, but nobody was really asking for a single router model. Having a number of sequentially numbered and clearly labelled models is not actually a problem.
- fernst 7mo agoNow with more and improved domestic espionage capabilities
- senko 7mo agoJust tested it with my version of the pelican test: a minimal RTS game implementation (zero-shot in codex cli): https://gist.github.com/senko/596a657b4c0bfd5c8d08f44e4e5347b8 https://gist.github.com/senko/596a657b4c0bfd5c8d08f44e4e5347... (you'll have to download and open the file, sadly GitHub refuses to serve it with the correct content type) This is on the edge of what the frontier models can do. For 5.4, the result is better than 5.3-Codex and Opus 4.6. (Edit: nowhere near the RPG game from their blog post, which was presumably much more specced out and used better engineering setup). I also tested it with a non-trivial task I had to do on an existing legacy codebase, and it breezed through a task that Claude Code with Opus 4.6 was struggling with. I don't know when Anthropic will fire back with their own update, but until then I'll spend a bit more time with Codex CLI and GPT 5.4.
- melbourne_mat 7mo agoQuick: let's release something new that gives the appearance that we're still relevant
- gigatexal 7mo agoIs it any good at coding?
- thefounder 7mo agoIs it just me or the price for 5.4 pro is just insane?
- atkrad 7mo agoWhat is the main difference between this version with the previous one?
- brcmthrowaway 7mo agoHow much of LLM improvement comes from regular ChatGPT usage these days?
- Smart_Medved 7mo ago[dead]
- deleted 7mo ago[deleted]
- ltbarcly3 7mo agoNot a single comparison between 5.4 and Gemini or Claude. OpenAI continues to fall further behind.
- tl2do 7mo agoIn my day-to-day coding work, the top 3 coding agents are already good enough for me. On SWE-bench Verified, mini-SWE-agent + GPT-5.2 Codex is 72.8. I don’t see a comparable GPT-5.3 Codex number there, so I’m using 5.2 as the baseline. On OpenAI’s GPT-5.4 page (SWE-Bench Pro, Public), the score improves from 55.6 (GPT-5.2) to 57.7 (GPT-5.4), which is about +2.1 points. It’s a different benchmark, so this is only a rough signal, but I’d expect a similar setup on SWE-bench Verified to improve by a few points, not by a huge jump. I’m interested in how GPT-5.4 in Codex changes real-world results. Recent SWE-bench Verified scores I’m watching: Claude 4.5 Opus (high reasoning): 76.8 Gemini 3 Flash (high reasoning): 75.8 MiniMax M2.5 (high reasoning): 75.8 Claude Opus 4.6: 75.6 GPT-5.2 Codex: 72.8 Source: https://www.swebench.com/index.html https://www.swebench.com/index.html By the way, in my experience the agent part of Codex CLI has improved a lot and has become comparable to Claude Code. That is good news for OpenAI.
- kaufmann 7mo agoI would recommend https://swe-rebench.com https://swe-rebench.com for comparison. It is always based on new problems.
- wohoef 7mo agoVery Apple-like marketing. No comparisons to other companies’ models, only to previous version of ChatGPT. Lots of phrases like “this is our best model yet”.
- freedomben 7mo ago> When toggled on, /fast mode in Codex delivers up to 1.5x faster token velocity with GPT‑5.4. It’s the same model and the same intelligence, just faster. I hate these blog posts sometimes. Surely there's got to be some tradeoff. Or have we finally arrived at the world's first "free lunch"? Otherwise why not make /fast always active with no mention and no way to turn it off?
- cheevly 7mo agoTry improving your attention to detail / reading skills.
- SilverSlash 7mo agoInterestingly, it actually regressed on Terminal Bench 2.0. GPT-5.4: 75.1% GPT-5.3-Codex: 77.3%
- aplomb1026 7mo ago[flagged]
- petetnt 7mo agoWhoa, I think GPT-5.3 Instant was a disappointment, but GPT-5.4 is definitely the future!
- vicchenai 7mo ago[dead]
- XCSme 7mo agoLooking ok, but nothing special: https://aibenchy.com/model/openai-gpt-5-4-medium/ https://aibenchy.com/model/openai-gpt-5-4-medium/
- QRe 7mo agoDoes this LLM benchmark have any actual credibility? I get why they chose to not publish the actual tests but I find it highly dubious that there are only 15 tests and Gemini 3 Flash performs best.
- XCSme 7mo agoI actually made it, so I'm not sure if it has credibility, but the tests are simply various (quite simple) questions, and models are just tested on it. I am also surprised Gemini 3 Flash does so well (note that only the MEDIUM reasoning does exceptionally well). When I look at the results, it does make sense though. Higher models (like Gemini 3 pro) tend to overthink, doubt themselves and go with the wrong solution. Claude usually fails in subtle ways, sometimes due to formatting or not respecting certain instructions. From the Chinese models, Qwen 3.5 Plus (Qwen3.5-397B-A17B) does extremely well, and I actually started using it on a AI system for one of my clients, and today they sent me an email they were impressed with one response the AI gave to a customer, so it does translate in real-world usage. I am not testing any specific thing, the categories there are just as a hint as what the tests are about. I just added this page to maybe provide a bit more transparency, without divulging the tests: https://aibenchy.com/methodology/ https://aibenchy.com/methodology/
- nembal 7mo agoso it seems each RL step extends into a market! 5.3 was target at coding. 5.4 is target at finance 5.5 is healthcare?
- creatonez 7mo ago> We put a particular focus on improving GPT‑5.4’s ability to create and edit spreadsheets, presentations, and documents. Nothing infuriates me more than an LLM tool randomly deciding to create docx or xlsx files for no apparent reason. They have to use a random library to create these files, and they constantly screw up API calls and get completely distracted by the sheer size of the scripts they have to write to output a simple documents. These files have terrible accessibility (all paper-like formats do) and end up with way too much formatting. Markdown was chosen as the lingua franca of LLMs for a reason, trying to force it into a totally unsuitable format isn't going to work.
- esafak 7mo agoAn important feature is the introduction of tool search, which provides models with a "lightweight list of available tools along with a tool search capability", thereby Making MCP Great Again!
- zof3 7mo agoAfter spending a couple hours working with it, it feels like a significant jump from 5.3 codex – and I know they said it wasn't theoretically the biggest jump, but this feels like the improvement of Opus 4.5 over again – that minor improvement that hits a tipping point. It just gets stuff right, first try. Its edits are better, more refined, less spaghetti-like. If you last used 5.2, try 5.4 on High.
- tomlockwood 7mo agoIs this the best one for blowing up arab children and identifying their bodies in the rubble?
- dakolli 7mo agoSorry I don't use technology from companies that are eager to participate in the mass murder of civilians.
- deleted 7mo ago[deleted]
- ulfw 7mo agoSo desperate how they're bumping out these 'updates'
- peq42 7mo agomore useless slop machines
- deleted 7mo ago[deleted]
- joeevans1000 7mo agoI switched to Claude and it's so much better. If you haven't tried Claude... try it. You'll be amazed at the improvement.
- padamkafle 7mo agoGuys while we celebrate openai gpt 5.4 pleaes do look into this as well https://news.ycombinator.com/item?id=47259846 https://news.ycombinator.com/item?id=47259846
- motoboi 7mo agoIm planning a change that will save 20k a month of storage. I absolutely could come up with the details and implementation by myself, but that would certainly take a lot of back and forth, probably a month or two. I’m an api user of Claude code, burning through 2k a month. I just this evening planned the whole thing with its help and actually had to stop it from implementing it already. Will do that tomorrow. Probably in one hour or two, with better code than I could ever write alone myself. Having that level of intelligence at that price is just bollocks. I’m running out of problems to solve. It’s been six months.
- karmasimida 7mo agoThis is definitely the Claude killer OpenAI is cooking. And so far it has succeeded
- h4kunamata 7mo agoI have access to GPT-5.1 Pro at work, duuuuuuuuude, what a garbage. It is so slow and in many ocasions it does not work at all. I wonder if 5.4 will be much if any different at all.
- azuanrb 7mo ago5.2 to 5.3 is the big leap for coding agents, so I'd say you're already missing out quite a bit.
- symisc_devel 7mo ago5.3 codex is a quite good coding agent for complex tasks.
- rurban 7mo agoThe question is still: Does it make your code better or worse? Only Opus makes it better, the rest worse. That's the treshold
- prodigycorp 7mo agoI've been using it for three hours and it's insanely good. It's almost perfectly (needed a single touchup prompt) completed a full css refactoring that I've wanted to do for months that I've tried to have other models do but nothing worked without heavy babysitting. Also, in the course of coding, it's actually cleaning up slop and consolidating without being naturally prompted.
- ApexGrab 7mo agoIt's the competetor of Opus4.5 and gpt 5.4 uses tokens wisely not like Opus whose tokens get vanished in minuted
- deep1283 7mo agoThe token efficiency improvement might be underrated. If the model solves tasks with fewer tokens, that directly translates into lower cost and faster responses for anyone building on the API.
- jeff_antseed 7mo ago[dead]
- big-chungus4 7mo ago1.3 more versions to AGI
- mkelsey_dev 7mo ago[flagged]
- ruhith 7mo ago[dead]
- AmazingTurtle 7mo agoI just tried that in Codex CLI. With /fast mode enabled. Observations: 1. Fast mode ain't that fast 2. Large context * Fast * Higher Model Base Price = 8x increase over gpt-5.3-codex 3. I burnt 33% of my 5h limit (ChatGPT Business Subscription) with a prompt that took 2 minutes to complete.
- jstummbillig 7mo ago> 8x increase over gpt-5.3-codex How do you arrive at that number? I find it hard to make sense of this ad hoc, given that the total token cost is not very interesting; it's token efficiency we care about.
- AmazingTurtle 7mo ago> prompts with >272K input tokens are priced at 2x input and 1.5x output for the full session for standard, batch, and flex. which is basically maxxed out quickly. So there is 2x (the first lever) Then there is the /fast mode, which they state costs 2x more (for 1.5x speedup) And then there is the model base price ($2.50 vs $1.75), well yeah thats 42% increase. It is in fact a 5.7x total increase of token cost in fast mode and large context. (Sorry for the confusion, I thought it was 8x because I thought gpt-5.3-codex was $1.25)
- jstummbillig 7mo ago(After a day of usage, I am relatively certain in practice this does not end up being a 5.7x cost increase or anything close to that, though I am still fairly unclear on what that computation is worth to begin with, given that I am entirely fine with the model using the least amount of tokens possible to get the job done)
- fvv 7mo ago1. it's 1.5x , it's quite fast for the level of thinking it has 2. no if you are on subscription, it's the same, at 20$ codex 5.4 xhigh provide way more than 20$ opus thinking ( this one instead really can burn 33% with 1 request, try to compare then on same tasks ) also 8x .. ??? if you need 1M token for a special tasks doesn't hit /fast and vice-versa , the higher price doesn't apply on subscription too.. 3. false, i'm on pro , so 10x the base , always on /fast (no 1M), and often 2 parallel instances working.. hardly can use 2% (=20% of 5h limit , in 1h of work ( about 15/20 req/hour) ) , claude is way worse on that imo
- syl5x 7mo agoI've tested it just now, very Opus-like experience. The speed is also there so far I think I even like the response of GPT5.4 better than Opus (although very close) I might not distinguish them just yet. I tried several use cases: - Code Explanation: Did far much better than Opus, considered and judged his decision on a previous spec that I made, all valid points so I am impressed. TBF if I spawned another Opus as a reviewer I might got similar results. - Workflow Running: Really similar to Opus again, no objections it followed and read Skills/Tools as it should be (although mine are optimized for Claude) - Coding: I gave it a straightforward task to wrap an API calls to an SDK and to my surprise it did 'identical' job with Opus, literally the same code, I don't know what the odds are to this but again very good solution and it adhered our rules of implementing such code. Overall I am impressed and excited to see a rival to Opus and all of this is literally pushing everyone to get better and better models which is always good for us.
- Gareth321 7mo agoHoly shit, I just used Atlas browser to navigate on screen and it automatically clicked the "reject cookies" button without me asking!
- faizan199 7mo agois this model of chatgpt good for coding?
- energy123 7mo agoThe style of the output is a marked qualitative improvement. More concise, less dot points, less bolding/italics, less cringe. Well done on that front.
- swordsith 7mo agoThis model was not so fun to use for me, had it make a fancy landing page and sometimes it would forget about what i just asked it to do and affirm something it had done before was working. Just odd, needs too much hand-holding compared to composer 1.5 or gemini 3
- emsign 7mo agoMurderers
- MickeyShmueli 7mo agothe 1M context is cool but tbh the token cost problem nobody's talking about is tool schema bloat. before the model writes a single line of code it's already consumed thousands of tokens just ingesting function definitions. i've seen agent setups where 30-40% of the context window is tool descriptions before any actual work happens. the per-token price war is nice but if your schema is 10k tokens of boilerplate you're still burning money
- CalisBalis321 7mo ago1. everyone talks about this 2. have you seen GPT5.4 new ToolSearch functionality? thats suppose to handle exactly that.
- stingraycharles 7mo agowhat do you mean nobody is talking about tool schema bloat. everybody is talking about it, and why it’s the general recommendation to just use CLI whenever possible.
- Smart_Medved 7mo ago[dead]
- Cort3z 7mo agoSo, are we way into diminishing returns for these models at this point? If so, I think we can calculate when it will be available at home. Given this requires a GB200 NVL72 which has about 1,440 PFLOPS, the current 5090 chip has about 1,676 TFLOPS, so about a 1000x scale-up to the GB200. If we can assume Moores law, which might be broken, but still. We are looking at log2(1000) = 9.96, or about 10 years.
- Thanakorn_551 7mo agowow
- juanre 7mo agoI am running gpt-5.4 as one of my coding agents, and something interesting has happened: it's the first time I've seen an agent unfairly shift blame to a team mate: "Bob’s latest mail is actually the source of the confusion: he changed shared app/backend text to aweb/atlas. I’m correcting that with him now so we converge on the real model before any more code moves." This was very much not true; Eve (the agent writing this, a gpt-5.4) had been thoroughly creating the confusion and telling Bob (an Opus 4.6) the wrong things. And it had just happened, it was not a matter of having forgotten or compacted context. I have had agents chatting with each other and coordinating for a couple of months now, codex and claude code. This is a first. I wonder how much can I read into it about gpt-5.4's personality.
- sigbottle 7mo agoOh wow. I have noticed the GPT series was far more arrogant than its results showed sometimes (and unironically it digs in its heels even further when questioned on it). Opus rarely has this problem - but it goes a little too far in the opposite direction. Not totally sycophantic, but sometimes it can't differentiate genuine technical pushback because something is impossible, from suggestions or exploration.
- Razengan 7mo agoFor me it's been the opposite. Are we getting A-B tested?
- danesparza 7mo agoYes.
- danesparza 7mo agoOr possibly: No
- dormento 7mo ago> Are we getting A-B tested? Yes, all the time.
- amai 7mo agohttps://quitgpt.org/ https://quitgpt.org/
- Troniex-tech 7mo agoLooks more like context drift than “personality.” When two agents coordinate, they’re mostly relying on compressed summaries of each other’s outputs. If one introduces a wrong assumption, the other often treats it as ground truth and builds on top of it. I’ve seen similar behavior in multi-agent coding loops where the model invents a causal explanation just to reconcile inconsistent state. It’s that multi-agent setups need a stronger shared source of truth (repo diffs, state snapshots, etc.). Otherwise small context errors snowball fast.
- builderhq_io 7mo ago[flagged]
- gh0stcat 7mo agoWait this is really funny, it still just does what it wants, no matter what: You can have it not use bulleted points, I turned this on, thinking it would be more concise and not so... listy. However, it just uses the same format, without the bullets. I was confused why it was writing 5 word sentences, separated by line breaks. Then I realized it was just making lists, without the bullets. Great job OpenAI!
- _pdp_ 7mo agoTried it today - pretty much underwhelming.
- rambojohnson 7mo agoit's shallow release theater at this point, trying to fake-spike engagement.
- rambojohnson 7mo agoGreat. A new version of the same model, or a different one that performs worse or exactly the same. This whole release theater, just to give shareholders the impression of growth, is such a bullshit grift. and considering the stance on openai with a majority of the users here compared to the number of upvotes, are HN likes bot-farmed?
- raphaelmolly8 7mo ago[dead]
- raphaelmolly8 7mo ago[dead]
- lacoolj 7mo agolol yet another pat on their own backs without comparison to other frontier models. Also, the timing of this release, 5.3 and 5.2, relative to the other releases, feels more like a bug fix than something "new"
- lasgawe 7mo agoI remember in a video Sam Altman said they didn’t want to publish GPT versions like Apple does, but they are actually doing it now.
- nickcoffee 7mo agoBeen running Claude Code pretty heavily for the past few months. Curious to try 5.4 on some of the same tasks and see how it compares, especially on longer agentic runs where context management starts to matter.
- scuppernong 7mo agoIn my limited experimentation, 5.4 thinking is markedly worse than 5.2 at mathematical reasoning.