8 ms·
GPT‑5.4 Mini and Nano
- miltonlost 7mo ago[flagged]
- machinecontrol 7mo agoWhat's the practical advantage of using a mini or nano model versus the standard GPT model?
- aavci 7mo agoCheaper. Every month or so I visit the models used and check whether they can be replaced by the cheapest and smallest model possible for the same task. Some people do fine tuning to achieve this too.
- powera 7mo agoI've been waiting for this update. For many "simple" LLM tasks, GPT-5-mini was sufficient 99% of the time. Hopefully these models will do even more and closer to 100% accuracy. The prices are up 2-4x compared to GPT-5-mini and nano. Were those models just loss leaders, or are these substantially larger/better?
- HugoDias 7mo agoFor us, it was also pretty good, but the performance decreased recently, that forced us to migrate to haiku-4.5. More expensive but much more reliable (when anthropic up, of course).
- throwaway911282 7mo agothey dont change the model weights (no frontier lab does). if you have evals and all prompts, tool calls the same, I'm curious how you are saying performance decreased..
- powera 7mo agoSo far on my (simple) benchmarks, GPT-5.4-mini is looking very good. GPT-5.4-mini is about 30% faster than GPT-5-mini. GPT-5.4-mini gets 80% on the "how many Rs in Strawberry" test, and nearly perfect scores on everything else I threw at it. GPT-5.4-nano is less impressive. I would stick to gpt-5.4-mini where precise data is a requirement. But it is fast, and probably cheaper and better quality than an 8-20B parameter local model would be. ( https://encyclopedia.foundation/benchmarks/dashboard/ https://encyclopedia.foundation/benchmarks/dashboard/ for details - the data is moderately blurry - some outlier (15s) calls are included, a few benchmark questions are ambiguous, and some prices shown are very rough estimates ).
- HugoDias 7mo agoAccording to their benchmarks, GPT 5.4 Nano > GPT-5-mini in most areas, but I'm noticing models are getting more expensive and not actually getting cheaper? GPT 5 mini: Input $0.25 / Output $2.00 GPT 5 nano: Input: $0.05 / Output $0.40 GPT 5.4 mini: Input $0.75 / Output $4.50 GPT 5.4 nano: Input $0.20 / Output $1.25
- karmasimida 7mo agoThose are bigger models. The serving isn’t going to be cheaper. Why expect cheaper then? The performance is also better
- trvz 7mo agoYou seem to have insight into the size of OpenAI’s models. Care to share the parameter counts for them?
- simianwords 7mo agomodels are getting costlier but by performance getting cheaper. perhaps they don't see a point supporting really low performance models?
- HugoDias 7mo agoI would be curious to know if from the enterprise / API consumption perspective, these low-performance models aren't the most used ones. At least it matches our current scenario when it comes to tokens in / tokens out. I'd totally buy the price increase if these are becoming more efficient though, consuming less tokens.
- ryao 7mo agoI will be impressed when they release the weights for these and older models as open source. Until then, this is not that interesting.
- simianwords 7mo agowhy isn't nano available in codex? could be used for ingesting huge amount of logs and other such things
- patates 7mo agoIMHO the best way is to let a SOTA model have a look at bunch of random samples and write you tools to analyze those. I think, no model, SOTA or not, has neither the context nor the attention to be able to do anything meaningful with huge amount of logs.
- BoumTAC 7mo agoTo me, mini releases matter much more and better reflect the real progress than SOTA models. The frontier models have become so good that it's getting almost impossible to notice meaningful differences between them. Meanwhile, when a smaller / less powerful model releases a new version, the jump in quality is often massive, to the point where we can now use them 100% of the time in many cases. And since they're also getting dramatically cheaper, it's becoming increasingly compelling to actually run these models in real-life applications.
- pzo 7mo agothey do are cheaper than SOTA but not getting dramatically cheaper but actually the opposite - GPT 5.4 mini is around ~3x more expensive than GPT 5.0 mini. Similarly gemini 3.1 flash lite got more expensive than gemini 2.5 flash lite.
- BoumTAC 7mo agoBut they are getting dramatically better. What's the point of a crazy cheap model if it's shit ? I code most of the time with haiku 4.5 because it's so good. It's cheaper for me than buying a 23€ subscription from Anthropic.
- philipkglass 7mo agoThe crazy cheap models may be adequate for a task, and low cost matters with volume. I need to label millions of images to determine if they're sexually suggestive (this includes but is not limited to nudity). The Gemini 2.0 Flash Lite model is inexpensive and performs well. Gemini 2.5 Flash Lite is also good, but not noticeably better, and it costs more. When 2.0 gets retired this June my costs are going up.
- dev_hugepages 7mo agoTime to gather a dataset and train your own model!
- 7mo ago
- cbg0 7mo agoBased on the SWE-Bench it seems like 5.4 mini high is ~= GPT 5.4 low in terms of accuracy and price but the latency for mini is considerably higher at 254 seconds vs 171 seconds for GPT5.4. Probably a good option to run at lower effort levels to keep costs down for simpler tasks. Long context performance is also not great.
- beklein 7mo agoAs a big Codex user, with many smaller requests, this one is the highlight: "In Codex, GPT‑5.4 mini is available across the Codex app, CLI, IDE extension and web. It uses only 30% of the GPT‑5.4 quota, letting developers quickly handle simpler coding tasks in Codex for about one-third the cost." + Subagents support will be huge.
- hyperbovine 7mo agoHaving to invoke `/model` according to my perceived complexity of the request is a bit of a deal breaker though.
- serf 7mo agoyou use profiles for that [0], or in the case of a more capable tool (like opencode) they're more confusing referred to as 'agents'[1] , which may or may not coordinate subagents.. So, in opencode you'd make a "PR Meister" and "King of Git Commits" agent that was forced to use 5.4mini or whatever, and whenever it fell down to using that agent it'd do so through the preferred model. For example, I use the spark models to orchestrate abunch of sub-agents that may or may not use larger models, thus I get sub-agents and concurrency spun up very fast in places where domain depth matter less. [0]: https://developers.openai.com/codex/config-advanced#profiles https://developers.openai.com/codex/config-advanced#profiles [1]: https://opencode.ai/docs/agents/ https://opencode.ai/docs/agents/
- system2 7mo agoI am feeling the version fatigue. I cannot deal with their incremental bs versions.
- yomismoaqui 7mo agoNot comparing with equivalent models from Anthropic or Google, interesting...
- Tiberium 7mo agoThey did actually compare them in the tweet, see https://x.com/OpenAI/status/2033953592424731072 https://x.com/OpenAI/status/2033953592424731072 Direct image: https://pbs.twimg.com/media/HDoN4PhasAAinj_?format=png&name=4096x4096 https://pbs.twimg.com/media/HDoN4PhasAAinj_?format=png&name=...
- casey2 7mo agoI googled all the testimonial names and they are all linked-in mouthpieces.
- Tiberium 7mo agoI checked the current speed over the API, and so far I'm very impressed. Of course models are usually not as loaded on the release day, but right now: - Older GPT-5 Mini is about 55-60 tokens/s on API normally, 115-120 t/s when used with service_tier="priority" (2x cost). - GPT-5.4 Mini averages about 180-190 t/s on API. Priority does nothing for it currently. - GPT-5.4 Nano is at about 200 t/s. To put this into perspective, Gemini 3 Flash is about 130 t/s on Gemini API and about 120 t/s on Vertex. This is raw tokens/s for all models, it doesn't exclude reasoning tokens, but I ran models with none/minimal effort where supported. And quick price comparisons: - Claude: Opus 4.6 is $5/$25, Sonnet 4.6 is $3/$15, Haiku 4.5 is $1/$5 - GPT: 5.4 is $2.5/$15 ($5/$22.5 for >200K context), 5.4 Mini is $0.75/$4.5, 5.4 Nano is $0.2/$1.25 - Gemini: 3.1 Pro is $2/$12 ($3/$18 for >200K context), 3 Flash is $0.5/$3, 3.1 Flash Lite is $0.25/$1.5
- coder543 7mo agoI wish someone would also thoroughly measure prompt processing speeds across the major providers too. Output speeds are useful too, but more commonly measured.
- JLO64 7mo agoIn my use case for small models I typically only generate a max of 100 tokens per API call, with the prompt processing taking up the majority of the wait time from the user perspective. I found OAI's models to be quite poor at this and made the switch to Anthropic's API just for this. I've found Haiku to be a pretty fast at PP, but would be willing to investigate using another provider if they offer faster speeds.
- asselinpaul 7mo agoOpenRouter has this information
- coder543 7mo agoI do not see prompt processing, only some kind of nebulous “throughput” that could be output or input+output, but definitely not input only.
- 6thbit 7mo agoLooking at the long context benchmark results for these, sounds like they are best fit for also mini-sized context windows. Is there any harness with an easy way to pick a model for a subagent based on the required context size the subagent may need?
- bananamogul 7mo agoThey could call them something like “sonnet” and “haiki” maybe.
- reconnecting 7mo agoAll three ChatGPT models (Instant, Thinking, and Pro) have a new knowledge cutoff of August 2025. Seriously?
- zild3d 7mo agowhats surprising about that? most of the minor version updates from all the labs are post training updates / not changing knowledge cutoff
- reconnecting 7mo agoThanks for letting me know, I will be waiting for the major update.
- F7F7F7 7mo agoIt's been like this since GPT 3.5. This is not a limitation and is generally considered a natural outcome of the process. So there's no major update in the sense that you might be thinking. Most of the time there's not even an announcement when/if training cut offs are updated. It's just another byline. A 6 month lag seems to be the standard across the frontier models.
- reconnecting 7mo agoI've actually started worrying that the amount of false data produced with LLMs on the public internet might provoke a situation where the knowledge cutoff becomes permanently (and silently) frozen. Like we can't trust data after 2025 because it will poison training data at scale, and models will only cover major events without capturing the finer details.
- gwern 7mo agoI agree. That's why you should write as much as you can now, if you want to get it into the LLMs (https://gwern.net/blog/2024/writing-online https://gwern.net/blog/2024/writing-online). You never know when the window will slam shut and LLM training goes 'hermetic' as they focus on 'civilization in a datacenter' where only extremely vetted whitelisted data gets included in the 'seed' and everything is reconstructed from scratch for the training value & safety.
- varispeed 7mo agoI stopped paying attention to GPT-5.x releases, they seem to have been severely dumbed down.
- pscanf 7mo agoI quite like the GPT models when chatting with them (in fact, they're probably my favorites), but for agentic work I only had bad experiences with them. They're incredibly slow (via official API or openrouter), but most of all they seem not to understand the instructions that I give them. I'm sure I'm _holding them wrong_, in the sense that I'm not tailoring my prompt for them, but most other models don't have problem with the exact same prompt. Does anybody else have a similar experience?
- nikanj 7mo agoSame, and I can't put my finger on the "why" either. Plus I keep hitting guard rails for the strangest reasons, like telling codex "Add code signing to this build pipeline, use the pipeline at ~/myotherproject as reference" and codex tells me "You should not copy other people's code signing keys, I can't help you with this"
- tom1337 7mo agoYea absolutely. I am using GPT 5.2 / 5.2 Codex with OpenCode and it just doesn't get what I am doing or looses context. Claude on the other side (via GitHub Copilot) has no problem and also discovers the repository on it's own in new sessions while I need to basically spoonfeed GPT. I also agree on the speed. Earlier today I tasked GPT 5.2 Codex with a small refactor of a task in our codebase with reasoning to high and it took 20 minutes to move around 20 files.
- furyofantares 7mo agoI don't know any reason to use 5.2, when 5.3 is quite a bit faster.
- spiderfarmer 7mo agoIf using OpenAI models, use the Codex desktop app, it runs circles around OpenCode.
- qaz_plm 7mo agoCan you educate me as to what makes Codex app superior using the same GPT model in both? Thx in advance!
- dack 7mo agoi want 5.4 nano to decide whether my prompt needs 5.4 xhigh and route to it automatically
- exitb 7mo agoLike any work estimation, it will likely disappoint.
- mrtesthah 7mo agoAs per OpenAI themselves, xhigh is only necessary if the agent gets stuck on a long running task. Otherwise it’s thinking trades use so many tokens of context that it’s less effective than high for a great majority of tasks. This has also been my experience.
- dack 7mo agoyes but didn't greg brockman say he just runs on xhigh at all times?
- kseniamorph 7mo agowow, not bad result on the computer use benchmark for the mini model. for example, Claude Sonnet 4.6 shows 72.5%, almost on par with GPT-5.4 mini (72.1%). but sonnet costs 4x more on input and 3x more on output
- PunchTornado 7mo agowhat's the point of this benchmark if sonnet is working great at my tasks and mini can't solve my tasks?
- fastpdfai 7mo agoOne thing I really want to find out, is which model and how to process TONS of pdfs very very fast, and very accurate. For prediction of invoice date, accrual accounting and other accounting related purposes. So a decent smart model that is really good at pdf and image reading. While still being very very fast.
- JLO64 7mo agoI have a use case somewhat similar to this where I need to convert the content of PDFs in a non standard format to a specific YAML format. I currently use Haiku for this and am pleased with the accuracy/speed (I haven't tried scanned PDFs yet tho) however I've been thinking about fine tuning a small Qwen model for just this task. I can't yet justify the effort to investigate it but I imagine it could work out.
- mikkelam 7mo agoWhy are we treating LLM evaluation like a vibe check rather than an engineering problem? Most "Model X > Model Y" takes on HN these days (and everywhere) seem based on an hour of unscientific manual prompting. Are we actually running rigorous, version-controlled evals, or just making architectural decisions based on whether a model nailed a regex on the first try this morning?
- tanaros 7mo agoWhenever somebody makes a benchmark, people complain that the benchmark results are meaningless because they’re gamed. I don’t know why those same people don’t understand that grading on vibes is strictly worse.
- tintor 7mo agoDepends on benchmark. If questions are fixed they are trivial to game.
- pizza 7mo agoThere’s a Dark Forest problem for evals. As soon as they’re made public they start running out of time to be useful. It’s also not clear how to predict how the model will perform on a task based on an eval. Or even whether, given two skills that the model can individually do well on in the evals, it still does well on their composition. It might at this point be better to be scientific in unscientific approaches, than to attribute more power to relatively weakly predictive evals than they actually have
- xandrius 7mo agoIs "Dark Forest problem" an actual name? I just heard of the hypothesis and it has nothing to do with how you used it in this context.
- sebastiennight 7mo agoI believe the correct term is "Goodhart's Law": https://en.wikipedia.org/wiki/Goodhart%27s_law https://en.wikipedia.org/wiki/Goodhart%27s_law
- beernet 7mo agoCrazy how OAI is way behind now and the only one to blame is Sam, his ego and lust for influence. Their downwards trajectory of paying accounts since "the move" (DoW deal) is an open secret. If you had placed a new CEO at OAI six months ago and told him to destroy the company, it would have been hard for that CEO to do a better job at that than Sam did. Should have left when he was let go but decided to go full Greg and MAGA instead. Here we are. Go Dario
- beernet 7mo agoJust to elaborate, as I am getting downvoted by tech bros: OpenAI restructures after Anthropic captures 70% of new enterprise deals. Claude Code hits $2.5B while Codex lags at $1B ahead of dual IPOs. Src: https://www.implicator.ai/openai-cuts-its-side-quests-the-enterprise-already-left/ https://www.implicator.ai/openai-cuts-its-side-quests-the-en...
- tintor 7mo agoSeveral customer testimonials for GPT-5.4 Mini have em dashes in them. Did GPT write them?
- kennywinker 7mo agoUsers of AI used AI? Shocking
- derefr 7mo agoOpenAI don't talk about the "size" or "weights" of these models any more. Anyone have any insight into how resource-intensive these Mini/Nano-variant models actually are at this point? I assume that OpenAI continue to use words like "mini" and "nano" in the names of these model variants, to imply that they reserve the smallest possible resource-units of their inference clusters... but, given OpenAI's scale, that may well be "one B200" at this point, rather than anything consumers (or even most companies) could afford. I ask because I'm curious whether the economics of these models' use-cases and call frequency work out (both from the customer perspective, and from OpenAI's perspective) in favor of OpenAI actually hosting inference on these models themselves, vs. it being better if customers (esp. enterprise customers) could instead license these models to run on-prem as black-box software appliances. But of course, that question is only interesting / only has a non-trivial answer, if these models are small enough that it's actually possible to run them on hardware that costs less to acquire than a year's querying quota for the hosted version.
- technocrat8080 7mo agoHave they ever talked about their size or weights?
- derefr 7mo agoThey never put the parameter counts in their model names like other AI companies did, but back in the GPT3 era (i.e. before they had PR people sitting intermediating all their comms channels), OpenAI engineers would disclose this kind of data in their whitepapers / system cards. IIRC, GPT-3 itself was admitted to be a 175B model, and its reduced variants were disclosed to have parameter-counts like 1.3B, 6.7B, 13B, etc.
- technocrat8080 7mo agoWow, would love to see a source for this.
- derefr 7mo ago
- technocrat8080 7mo ago5.4 Mini's OSWorld score is a pleasant surprise. When SOTA scores were still ~30-40 models were too slow and inaccurate for realtime computer use agents (rip Operator/Agent). Curious if anyone's been using these in production.
- Someone1234 7mo agoPeople seem to dismiss OSWorld as "OpenClaw," but I think they're missing how powerful and flexible that type of full-interaction for safe workflows. We have a legacy Win32 application, and we want to side-by-side compare interactions + responses between it and the web-converted version of the same. Once you've taught the model that "X = Y" between the desktop Vs. web, you've got yourself an automated test suite. It is possible to do this another way? Sure, but it isn't cost-effective as you scale the workload out to 30+ Win32 applications.
- ibrahim_h 7mo agoThe OSWorld numbers are kinda getting lost in the pricing discussion but imo that's the most interesting part. Mini at 72.1% vs 72.4% human baseline is basically noise, so why not just use mini by default unless you're hitting specific failure modes. Also context bleed into nano subagents in multi-model pipelines — I've seen orchestrators that just forward the entire message history by default (or something like messages[-N:] without any real budgeting), so your "cheap" extraction step suddenly runs with 30-50K tokens of irrelevant context. And then what's even the point, you've eaten the latency/cost win and added truncation risk on top. Has anyone actually measured where that cutoff is in practice? At what context size nano stops being meaningfully cheaper/faster in real pipelines, not benchmarks.
- jbellis 7mo agoBenchmarking these now. Preregistering my predictions: Mini: better than Haiku but not as good as Flash 3, especially at reasoning=none. Nano: worse than Flash 3 Lite. Probably better than Qwen 3.5 27b.
- attentive 7mo agoPlease post it here. I'd also like to know if 5.4 mini is better than Flash 3. Include reasoning and timing, if possible.
- Rapzid 7mo agoOh.. I thought maybe these would be upgrades to gpt-4.1 and gpt-4.1-mini and etc.. But the latency is way too high compared to the 400-600. Yeah, different models and etc but the naming is confusing.
- nicpottier 7mo agoI've been struggling on finding a reasonably priced model to use with my toy openclaw instance. Opus 4.6 felt kinda magical but that's just too expensive and I'm not risking my max subscription for it. GPT 5.4 mini is the first alternative that is both affordable and decent. Pretty impressed. On a $20 codex plan I think I'm pretty set and the value is there for me.
- GaggiX 7mo agoOpen source models like MiniMax M2.5, GLM 5, Kimi K2.5 were not decent enough? (via openrouter)
- nicpottier 7mo agoI will confess that I have not had time to play with those. Will give them a try, thanks for the recommendation.
- selfhoster11 7mo agoK2.5 and GLM-4.7/-5 were good in my experience, another vote for those.
- simonw 7mo agoHere's a grid of pelicans for the different models and reasoning levels: https://static.simonwillison.net/static/2026/gpt-5.4-pelican-family.svg https://static.simonwillison.net/static/2026/gpt-5.4-pelican...
- nharada 7mo agoSurely this task must now be in the training data
- Kye 7mo agoIf it does and works well then it seems like mission accomplished and time for a new benchmark.
- elif 7mo agoNano medium must have been run when the servers were on fire
- 6thbit 7mo agoThanks for the grid. The nano xhigh is my favorite pelican
- castral 7mo agoSome of these are nightmare fuel. I love them.
- morpheos137 7mo agoi switched to claude when i found chatgpt would argue with just about anything I said even when it was wrong. they have over optimised antisychophancy. i want a model that simulates critical thinking not one that repeats half baked often incomplete dogmas. the chatgpt 5x range is extraordinarily powerful but also extra ordinarily frustrating to try to use for anything creative or productive that is original in my opinion. claude basically is able to think critically while being neither sycophantic or argumentative most of the time in my option with appropriate user prompting. recent chat gpts seem to fight me every step of the way when not doing boiler plate. i don't want to waste my time fighting a tool.
- XCSme 7mo agoIt's odd, that on many benchmarks, including mine[0], Nano does better than Mini. 5.4 mini seems to struggle with consistency, and even with temperature 0 sometimes gives the correct response, sometimes a wrong one... [0]: https://aibenchy.com/compare/openai-gpt-5-4-medium/openai-gpt-5-4-mini-medium/openai-gpt-5-4-nano-medium/openai-gpt-5-mini-medium/ https://aibenchy.com/compare/openai-gpt-5-4-medium/openai-gp...
- michaelgdwn 7mo agoThe Nano tier is the one I'm watching. For agent workflows where you're making dozens of LLM calls per task, the cost per call matters more than peak capability. Would be interesting to see benchmarks on function calling latency specifically — that's what matters for agents.
- dmix 7mo agoLast time I used GPT-5 mini it seems much slower than the primary GPT model API when we used it for an AI chat agent. Particularly around streaming the responses. But everything I've read implies it's supposed to be faster.
- jerrygoyal 7mo agoIs GPT-5.4Mini drastically or marginally better for writing tasks as compared to GPT-5Mini?
- xyproto 7mo agoOpenAI has "open" in the name without being anything similar to "open source". Additionally, they have not rejected using their technology for automatically killing people and for mass surveillance. I deleted my OpenAI account, and it felt good. Recommended.
- AbstractH24 7mo agoIs anyone else getting numb to new model announcements?
- pugchat 7mo ago[flagged]