35 ms·
OpenAI o3 and o4-mini
- timonofathens 1y ago[dead]
- thm 1y agoI'm starting to be reminded of the razor blade business.
- Jordan-117 1y agoFuck Everything, We're Doing o5
- testfrequency 1y agoAs a consumer, it is so exhausting keeping up with what model I should or can be using for the task I want to accomplish.
- fkyoureadthedoc 1y ago[flagged]
- testfrequency 1y agoI’m assuming when you say “read once”, that implies reading once every single release? It’s confusing. If I’m confused, it’s confusing. This is UX 101.
- mrits 1y agoSome people don't blindly trust the marketing department of the publisher
- fkyoureadthedoc 1y agoThen it doesn't even matter what they name the model since it's just marketing that they wouldn't trust anyway.
- czk 1y ago"good at advanced reasoning", "fast at advanced reasoning", "slower at advanced reasoning but more advanced than the good one but not as fast but cant search the internet", "great at code and logic", "good for everyday tasks but awful at everything else", "faster for most questions but answers them incorrectly", "can draw but cant search", "can search but cant draw", "good for writing and doing creative things"
- fkyoureadthedoc 1y agoPutting the actual list would have made it too clear that I'm right I see
- sebzim4500 1y agoAside from anything else, having one model called o4 and one model called 4o is confusing. And I know they haven't released o4 yet but still.
- taberiand 1y agoWe'll know they have cracked AGI when they solve the hardest problem of all - naming things
- darioush 1y agoIt's becoming a bit like iphone 3, 4... 13, 25... Ok they are all phones that run apps and have a camera. I'm not an "AI power user", but I do talk to ChatGPT + Grok for daily tasks and use copilot. The big step function happened when they could search the web but not much else has changed in my limited experience.
- refulgentis 1y agoThis is a very apt analogy. It confers to the speaker confirmation they're absolutely right - names are arbitrary. While also politely, implicitly, pointing out the core issue is it doesn't matter to you --- which is fine! --- but it may just be contributing to dull conversation to be the 10th person to say as much.
- n2d4 1y agoThis one seems to make it easier — if the promises here hold true, the multi-modal support probably makes o4-mini-high OpenAI's best model for most tasks unless you have time and money, in which case it's o3-pro.
- 1123581321 1y agoI think it can be confusing if you're just reading the news. If you use ChatGPT, the model selector has good brief explanations and teaching you about newly available options if you don't visit the dropdown. Anthropic does similarly.
- CamperBob2 1y agoI asked OpenAI how to choose the right USB cable for my device. Now the objects around me are shimmering and winking out of existence, one by one. Help
- ithkuil 1y agoLol. But that's nothing. Wait until you shimmer and wink in and out of existence, like llms do during each completion
- tempaccount420 1y agoAs another consumer, I think you're overreacting, it's not that bad.
- energy123 1y agoGemini 2.5 Pro for every single task was the meta until this release. Will have to reassess now.
- hollerith 1y agoHuh. I use Gemini 2.0 Flash for many things because it's several times faster than 2.5 Pro.
- mring33621 1y agoAgreed. I pretty much stopped shopping around once Gemini 2.0 Flash came out. For general, cloud-centric software development help, it does the job just fine. I'm honestly quite fond of this Gemini model. I feel silly saying that, but it's true.
- jug 1y agoYes, this one is addictive for its speed and I like how Google was clever and also offered it in a powerful reasoning edition. This helps offset deficiencies from being smaller while still being cheap. I also find it quite sufficient for my kind of coding. I only pull out 2.5 Pro on larger and complex code bases that I think might need deeper domain specific knowledge beyond the coding itself.
- blueprint 1y agohow do you deal with the fact that they use all of your data for training their own systems and review all conversations
- sharkjacobs 1y agogemini-2.5-pro-preview-03-25 is the paid version which doesn't use your data https://ai.google.dev/gemini-api/terms#data-use-paid https://ai.google.dev/gemini-api/terms#data-use-paid
- hirvi74 1y ago
- yoyohello13 1y agoThe answer is to just use the latest Claude model and not worry beyond that.
- boznz 1y agoIt feels like all the AI companies are pulling the versions out of their arse at the moment, I think they should work backwards and work to AGI 1.0 So my guess currently is that most are lingering at about 0.3
- brap 1y agoWhere's the comparison with Gemini 2.5 Pro?
- kridsdale1 1y agoExactly.
- gallerdude 1y agoFor coding, I like the Aider polyglot benchmark, since it covers multiple programming languages. Gemini 2.5 Pro got 72.9% o3 high gets 81.3%, o4-mini high gets 68.9%
- asadm 1y agothanks
- vessenes 1y agowhere do you find those o3 high numbers? https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/ currently has gemini 2.5 pro as the leader at, as you say, 72.9%.
- croemer 1y agoIsn't it easy to train on the specific Exercism exercises that this benchmark uses?
- jumpCastle 1y ago
- falleng0d 1y agoMaybe they should ask the new models to generate a better name for themselves. It's getting quite confusing.
- jdross 1y agoThe pace of notable releases across the industry right now is unlike any time I remember since I started doing this in the early 2000's. And it feels like it's accelerating
- emp17344 1y agoNot really. We’re definitely in the incremental improvement stage at this point. Certainly no indication that progress is “accelerating”.
- nwienert 1y agoChatGPT 3 : iPhone 1 A bunch of models later, we're about on the iPhone 4-5 now. Feels about right.
- int_19h 1y agoIt's more like GPT-3 is the Manchester Baby, and we're somewhere around IBM 700 series right now. Still a long way to go to iPhone, as much as the industry likes to pretend otherwise.
- nwienert 1y agoBoth were big consumer commercial breakouts and far better than predecessors. And several years later both see only iterative improvements. Neither apply to your analogy.
- Workaccount2 1y agoIntegration is accelerating rapidly. Even if model development froze today, we would still probably have ~5 years of adoption and integration before it started to level off.
- littlestymaar 1y agoYou are both correct. It feels like the tech itself is kinda plateauing but it's still massively under-used. It will take a decade or more before the deployment starts slowing down.
- typs 1y agoI’m not sure I fully understand the rationale of having newer mini versions (eg o3-mini, o4-mini) when previous thinking models (eg o1) and smart non-thinking models (eg gpt-4.1) exist. Does anyone here use these for anything?
- sho_hn 1y agoI use o3-mini-high in Aider, where I want a model to employ reasoning but not put up with the latency of the non-mini o1.
- drvladb 1y agoo1 is a much larger, more expensive to operate on OpenAI's end. Having a smaller "newer" (roughly equating newer to more capable) model means that you can match the performance of larger older models while reducing inference and API costs.
- morkalork 1y agoIf the ai is smart, why not have it choose the model for the user
- zvitiate 1y agoThat’s what GPT-5 was supposed to be (instead of a new base or reasoning model) last Sam updated his plans I thought. Did those change again?
- firejake308 1y agoNot sure what the goal is with Codex CLI. It's not running a local LLM right, just a CLI to make API calls from the terminal?
- maheshrijal 1y agoThis might be their answer to claude code more than anything else.
- sho_hn 1y agoLooks more like a direct competitor to Aider.
- whitten 1y agoWhere do I find out more about Aider ?
- tailspin2019 1y agohttps://aider.chat https://aider.chat
- stavros 1y agoJust wait a few seconds and there will be a post here with Aider benchmarks for the new model, or https://aider.chat https://aider.chat
- mpaepper 1y agoYes, that's exactly what I thought as well. An attempt to get more share in the developer tooling space for the long term.
- originalvichy 1y agoIs there a non-obvious reason using something like Python to solve queries requiring calculations was not used from day one with LLMs?
- planb 1y agoBecause it‘s not a feature of the LLM but the product that is built around it (like ChatGPT).
- rahimnathwani 1y agoIt's true that product provides the tools, but the model still needs to be trained to use tools, or it won't use them well or at the right times.
- ipsum2 1y agoLLMs could not use tools on day one.
- ksylvest 1y agoAre these available via the API? I'm getting back 'model_not_found' when testing.
- planb 1y agoWhat is wrong with OpenAI? The naming of their models seems like it is intentionally confusing - maybe to distract from lack of progress? Honestly, I have no idea which model to use for simply everyday tasks anymore.
- sho_hn 1y agoSeems to me like they're somewhat trying to simplify now. GPT-N.m -> Non-reasoning oN -> Reasoning oN+1-mini -> Reasoning but speedy; cut-down version of an upcoming oN model (unclear if true or marketing) It would be nice if they actually stick to this pattern.
- jagger27 1y agoAre the oN models built on top of GPT-N.m models? It would be nice to know the lineage there.
- bogtog 1y agoI suspect that "ChatGPT-4o" is the most confusing part. Absolutely baffling to go with that and then later "oN", but surely they will avoid any "No" models moving forward
- krackers 1y agoBut we have both 4o and 4.1 for non-reasoning. And it's still not clear to me which is better (the comparison on their page was from an older version of 4o).
- dabeeeenster 1y agoIt really is bizarre. If you had asked me 2 days ago I would have said unequivically that these models already existed. Surely given the rate of change a date-based numbering system would be more helpful?
- i_love_retros 1y agoI tend to look at the lmarena leaderboard to see what to use (or the aider polyglot leaderboard for coding)
- zapnuk 1y agoSurprisingly, they didn't provide a comparison to Sonnet 3.7 or Gemini Pro 2.5—probably because, while both are impressive, they're only slightly better by comparison. Lets see what the pricing looks like.
- oofbaroomf 1y agoThey didn't provide a comparison either in the GPT-4.1 release and quite a few past releases, which is telling of their attitude as an org.
- Workaccount2 1y agoLooks like they are taking a page from Apple's book, which is to never even acknowledge other products exist outside your ecosystem.
- stogot 1y agoApple has commercials for a decade making fun of “PCs”
- BeetleB 1y agoPricing is already available: https://platform.openai.com/docs/pricing https://platform.openai.com/docs/pricing
- burke 1y agoIt's pretty frustrating to see a press release with "Try on ChatGPT" and then not see the models available even though I'm paying them $200/mo.
- _bin_ 1y agoI see o4-mini on the $20 tier but no o3.
- deleted 1y ago[deleted]
- TuxSH 1y agoThey're supposed to be released today for everyone, and o3-pro for Pro users in a few weeks: "ChatGPT Plus, Pro, and Team users will see o3, o4-mini, and o4-mini-high in the model selector starting today, replacing o1, o3‑mini, and o3‑mini‑high." with rate limits unchanged
- wilg 1y agoThey are all now available on the Pro plan. Y'all really ought to have a little bit more grace to wait 30 minutes after the announcement for the rollout.
- drcongo 1y agoOr maybe OpenAI could wait until they'd released it before telling people to use it now.
- deleted 1y ago[deleted]
- squeaky-clean 1y agoThey'd probably want their announcement to be the one the press picks up instead of a tweet or reddit post saying "Did anyone else notice the new ChatGPT model?"
- wilg 1y ago
- andrethegiant 1y agoBuried in the article, a new CLI for coding: > Codex CLI is fully open-source at https://github.com/openai/codex https://github.com/openai/codex today.
- ipsum2 1y agoLooks like a Claude Code clone.
- jumpCastle 1y agoBut open source like aider
- dang 1y agoRelated ongoing thread: OpenAI Codex CLI: Lightweight coding agent that runs in your terminal - https://news.ycombinator.com/item?id=43708025 https://news.ycombinator.com/item?id=43708025
- behnamoh 1y agoOpenAI be like: o1, o1-mini, o1-pro, o3, o4-mini, gpt-4, gpt-4o, gpt-4-turbo, gpt-4.5, gpt-4.1, gpt-4o-mini, gpt-4.1-mini, gpt-4.1-nano, gpt-3.5-turbo
- deleted 1y ago[deleted]
- xqcgrek2 1y agoUnderwhelming. Cancelled my subscription in favor of Gemini Pro 2.5
- georgewsinger 1y agoVery impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively best in class now (or likely beating Claude with similar augmentation?).
- oofbaroomf 1y agoClaude got 63.2% according to the swebench.com leaderboard (listed as "Tools + Claude 3.7 Sonnet (2025-02-24)).[0] OpenAI said they got 69.1% in their blog post. [0] swebench.com/#verified
- awestroke 1y agoOpenAI have not shown themselves to be trustworthy, I'd take their claims with a few solar masses of salt
- georgewsinger 1y agoYes, however Claude advertised 70.3%[1] on SWE bench verified when using the following scaffolding: > For Claude 3.7 Sonnet and Claude 3.5 Sonnet (new), we use a much simpler approach with minimal scaffolding, where the model decides which commands to run and files to edit in a single session. Our main “no extended thinking” pass@1 result simply equips the model with the two tools described here—a bash tool, and a file editing tool that operates via string replacements—as well as the “planning tool” mentioned above in our TAU-bench results. Arguably this shouldn't be counted though? [1] https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F08bba4487fb5ac1ba52540ee656d7e4da10ca1be-1920x1145.png&w=1920&q=75 https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-...
- tedsanders 1y agoI think you may have misread the footnote. That simpler setup results in the 62.3%/63.7% score. The 70.3% score results from a high-compute parallel setup with rejection sampling and ranking: > For our “high compute” number we adopt additional complexity and parallel test-time compute as follows: > We sample multiple parallel attempts with the scaffold above > We discard patches that break the visible regression tests in the repository, similar to the rejection sampling approach adopted by Agentless; note no hidden test information is used. > We then rank the remaining attempts with a scoring model similar to our results on GPQA and AIME described in our research post and choose the best one for the submission. > This results in a score of 70.3% on the subset of n=489 verified tasks which work on our infrastructure. Without this scaffold, Claude 3.7 Sonnet achieves 63.7% on SWE-bench Verified using this same subset.
- oofbaroomf 1y agoFinally, a new SOTA model on SWE-bench. Love to see this progress, and nice to see OpenAI finally catching up in the coding domain.
- rahimnathwani 1y agoChatGPT Plus, Pro, and Team users will see o3, o4-mini, and o4-mini-high in the model selector starting today, replacing o1, o3‑mini, and o3‑mini‑high. I subscribe to pro but don't yet see the new models (either in the Android app or on the web version).
- oofbaroomf 1y agoSame...
- oofbaroomf 1y agoIt's there now in the web app for me.
- rahimnathwani 1y agoI see them in the Android app now.
- ApolloFortyNine 1y agoMaybe OpenAI needs an easy mode for all these people saying 5 choices of models (and that's only if you pay) is simply too confusing for them. They even provide a description in the UI of each before you select it, and it defaults to a model for you. If you just want an answer of what you should use and can't be bothered to research them, just use o3(4)-mini and call it a day.
- brokencode 1y agoI personally like being able to choose because I understand the tradeoffs and want to choose the best one for what I’m asking. So I hope this doesn’t go away. But I agree that they probably need some kind of basic mode to make things easier for the average person. The basic mode should decide automatically what model to use and hide this from the user.
- CaptainFever 1y agoWould that be considered a Mixture of Experts system?
- simonw 1y agoNo, Mixture of Experts is a really confusing term. It sounds like it means "have a bunch of models, one that's an expert in physics, one that's an expert in health etc and then pick the one that's a best fit for the user's query". It's not that. The "experts" are each another giant opaque blob of weights. The model is trained to select one of those blobs, but they don't have any form of human-understandable "expertise". It's an optimization that lets you avoid using ALL of the weights for every run through the model, which helps with performance. https://huggingface.co/blog/moe#what-is-a-mixture-of-experts-moe https://huggingface.co/blog/moe#what-is-a-mixture-of-experts... is a decent explanation.
- FergusArgyll 1y agoI thought sama said that that's the plan for gpt-5: a router which'll choose the right model and thinking level for you
- davidkunz 1y agoI wish companies would adhere to a consistent naming scheme, like <name>-<params>-<cut-off-month>.
- meetpateltech 1y agoo3 is cheaper than o1. (per 1M tokens) • o3 Pricing: - Input: $10.00 - Cached Input: $2.50 - Output: $40.00 • o1 Pricing: - Input: $15.00 - Cached Input: $7.50 - Output: $60.00 o4-mini pricing remains the same as o3-mini.
- ben_w 1y ago4o and o4 at the same time. Excellent work on the product naming, whoever did that.
- janderson215 1y agoIt took me reading your comment to realize that they were different and this wasn’t deja vu. Maybe that says more about me than OpenAI, but my gut agrees with you.
- throwuxiytayq 1y agoJust wait until they announce oA and A0. They jokingly admitted that they’re bad at naming in the 4.1 reveal video, so they’re certainly aware of the problem. They’re probably hoping to make the model lineup clearer after some of the older models get retired, but the current mess was certainly entirely foreseeable.
- ben_w 1y agoEnergy Intensive Exceptional Intelligence (Omni-domain), AKA E-I-E-I-O.
- stavros 1y agoOh, that was Altman Sam.
- ai-christianson 1y agoAm Saltman
- stavros 1y agoEnter.
- BriggyDwiggs42 1y agoHi, I’m an OpenAI recruiter. Are you interested in a position with us?
- evaneykelen 1y agoA suggestion for OpenAI to create more meaningful model names: {Size}-{Quarter/Year}-{Speed/Accuracy}-{Specialty} Where: * Size is XS/S/M/L/XL/XXL to indicate overall capability level * Quarter/Year like Q2-25 * Speed/Accuracy indicated as Fast/Balanced/Precise * Optional specialty tag like Code/Vision/Science/etc Example model names: * L-Q2-25-Fast-Code (Large model from Q2 2025, optimized for speed, specializes in coding) * M-Q4-24-Balanced (Medium model from Q4 2024, balanced speed/accuracy)
- jsnell 1y agoI think they should name them after fictional characters. Bonus points if they're trademarked characters. "You gotta try Mickey, it beats the crap out of Gandalf in coding."
- oofbaroomf 1y agoThis is even more incomprehensible to users who don't understand what this naming scheme is supposed to mean. Right now, most power users are keeping track of all the models and know what they are like, so this naming wouldn't help them. Normal consumers don't really know the difference between the models, but this wouldn't help them either - all those letters and numbers aren't super inviting and friendly. They could try just having a linear slider for amount of intelligence and another one for speed.
- pembrook 1y agoThank god we don’t usually let engineers name stuff in the west. While this is entirely logical in theory this is how you get LG style naming like “THE ALL NEW LG-CFT563-X2” I mean, it makes total sense, it tells you exactly the model, region, series and edition! Right??
- LanceJones 1y agoWhat about using Marvel superhero names (with permission, of course)? The studio keeps giving us stronger and stronger examples...
- deleted 1y ago[deleted]
- oofbaroomf 1y agoWhen are they going to release o3-high? I don't think it's in the API, and I certainly don't see it in the web app (Pro).
- wilg 1y ago> We expect to release OpenAI o3‑pro in a few weeks with full tool support. For now, Pro users can still access o1‑pro. https://openai.com/index/introducing-o3-and-o4-mini/ https://openai.com/index/introducing-o3-and-o4-mini/
- _fat_santa 1y agoSo at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much. All these models feel like they took the exact same base model, tweaked a few things and released it as an entirely new model rather than updating the existing ones. In fact based on some of the other comments here it sounds like these are just updates to their existing model, but they release them as new models to create more media buzz.
- amarcheschi 1y agoThe old Chinese strategy of having 7343 different phone models with almost the same specs to confuse the customer better
- MoonGhost 1y agonot only that. filling search lists on eBay with your products is old sellers' tactics. Try to search for used Dell workstation or server and you will see pages and pages from the same seller.
- kylehotchkiss 1y agoThis sounds like recent Dell and Lenovo strategies
- whalesalad 1y agorecent? they've been doing this for decades. person a: "I just got an new macbook pro!" person b: "Nice! I just got a Lenovo YogaPilates Flipfold XR 3299 T92 Thinkbookpad model number SRE44939293X3321" ... person a: "does that have oled?" person b: "Lol no silly that is model SRE44939293XB3321". Notice the B in the middle?!?! That is for OLED.
- 1y ago
- mentalgear 1y agoI have doubts whether the live stream was really live. During the live-stream the subtitles are shown line by line. When subtitles are auto-generated, they pop up word by word, which I assume would need to happen during a real live stream. Line-by-line subtitles are shown if the uploader provides captions by themselves for an existing video, the only way OpenAI could provide captions ahead of time, is if the "live-stream" isn't actually live.
- EcommerceFlow 1y agoA very subtle mention of o3-pro, which I'd imagine is now the most capable programming model. Excited to see when I get access to that. Good thing I stopped working a few hours ago EDIT: Altman tweeted o3-pro is coming out in a few weeks, looks like that guy misspoke :(
- Workaccount2 1y agoo4-mini, not to be confused with 4o-mini.
- spencersolberg 1y agoThe Codex CLI looks nice, but it's a shame I have to bring my own API key when I already subscribe to ChatGPT Plus
- carlita_express 1y ago> we’ve observed that large-scale reinforcement learning exhibits the same “more compute = better performance” trend observed in GPT‑series pretraining. Didn’t the pivot to RL from pretraining happen because the scaling “law” didn’t deliver the expected gains? (Or at least because O(log) increases in model performance became unreasonably costly?) I see they’ve finally resigned themselves to calling these trends, not laws, but trends are often fleeting. Why should we expect this one to hold for much longer?
- anothermathbozo 1y agoThis isn't exactly the case. The trend is a log scale. So a 10x in pretraining should yield a 10% increase in performance. That's not proving to be false per say but rather they are encountering practical limitations around 10x'ing data volume and 10x'ing available compute.
- carlita_express 1y agoI am aware of that, like I said: > (Or at least because O(log) increases in model performance became unreasonably costly?) But, yes, I left implicit in my comment that the trend might be “fleeting” because of its impracticality. RL is only a trend so long as it is fashionable, and only fashionable (i.e., practical) so long as OpenAI is fed an exponential amount of VC money to ensure linear improvements under O(log) conditions. OpenAI is selling to VCs the idea that some hitherto unspecified amount of linear model improvement will kick off productivity gains greater than their exponentially increasing investment. These productivity gains would be no less than a sizeable percentage of American GDP, which Altman has publicly set as his target. But as the capital required increases exponentially, the gap between linearly increasing model capability (i.e., its productivity) and the breakeven ROI target widens. The bigger model would need to deliver a non-linear increase in productivity to justify the exponential price tag.
- mode80 1y agoThis happens once it starts improving itself.
- jcynix 1y agoTo plan a visit to a dark sky place, I used duck.ai (Duckduckgo's experimental AI chat feature) to ask five different AIs on what date the new moon will happen in August 2025. GPT-4o mini: The new moon in August 2025 will occur on August 12. Llama 3.3 70B: The new moon in August 2025 is expected to occur on August 16, 2025. Claude 3 Haiku: The new moon in August 2025 will occur on August 23, 2025. o3-mini: Based on astronomical calculations, the new moon in August 2025 is expected to occur on August 7, 2025 (UTC). [...] Mistral Small 3: To determine the date of the new moon in August 2025, we can use astronomical data or a reliable astronomical calendar. As of my last update in October 2023, I don't have real-time data access, but I can guide you on how to find this information. [...] I got different answers, mostly wrong. My calendars (both paper and app versions) show me 23. august as the date. And btw, when I asked those AIs which entries in a robots.text file would block most Chinese search engines, one of them (Claude) told me that it can't tell because that might be discriminatory: "I apologize, but I do not feel comfortable providing recommendations about how to block specific search engines in a robots.txt file. That could be seen as attempting to circumvent or manipulate search engine policies, which goes against my principles."
- xnx 1y agoGemini gets the new moon right. Better to use one good model than 5 worse ones.
- kenjackson 1y agoI think all the full power LLMs will get it right because they do web search. ChatGPT 4 does as well.
- what_ever 1y agoGemini 2.0 Flash gets it correct too.
- WhatIsDukkha 1y agoI would never ask any of these questions of an LLM (and I use and rely on LLMs multiple times a day), this is a job for a computer. I would also never ask a coworker for this precise number either.
- oofbaroomf 1y agoStill a knowledge cutoff of August 2023. That is a significant bottleneck to devs using it for AI stuff.
- cryptoz 1y agoI've taken to pasting in the latest OpenAI API docs for their python library to each prompt (via API, I'm not pasting each time manually in ChatGPT) so that the AI can write code that uses itself! Like, I get it, the training data thing is hard, but - OpenAI changed their python library with breaking changes and their models largely still do not know about it! I haven't tried 4.1- series yet with their newer cutoff, but, the rest of the models like o3-mini (and I presume these new ones today) still write openai python library code in the old, broken style. Argh.
- kumarm 1y agoAnyone got codex working? After installing and setting up API Key I get this error : system OpenAI rejected the request (request ID: req_06727eaf1c5d1e3f900760d10ca565a7). Please verify your settings and try again. ╭──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
- bratao 1y agoOh god. I´m Brazilian and can´t get the "Verification". Using my passport or id. This is very frighting future.
- pton_xd 1y agoThis reminds me of keeping up with all the latest JavaScript framework trivia circa the ~2010s
- bufferoverflow 1y agoJS framework thing is still ongoing https://krausest.github.io/js-framework-benchmark/2025/table_chrome_135.0.7049.42.html https://krausest.github.io/js-framework-benchmark/2025/table...
- WhitneyLand 1y agoSo it looks like no increase in context window size since it’s not mentioned anywhere. I assume this announcement is all 256k, while the base model 4.1 just shot up this week to a million.
- basisword 1y agoThe user experience needs to be massively improved when it comes to model choice. How are average users supposed to know which model to pick? Why shouldn't I just always pick the newest or most powerful one? Why should I have to choose at all? I say this from the perspective of a ChatGPT user - I understand the different pricing on the API side helps people make decisions.
- iamronaldo 1y agoTyler cowen seems convinced https://marginalrevolution.com/marginalrevolution/2025/04/o3-and-agi-is-april-16th-agi-day.html https://marginalrevolution.com/marginalrevolution/2025/04/o3...
- jonahx 1y agoIt can't solve this puzzle: https://i.imgur.com/AJqbqHJ.png https://i.imgur.com/AJqbqHJ.png Thought for 3m 51s Short answer → you can’t. The breathtaking thing is not the model itself, but that someone as smart as Cowen (and he's not the only one) is uttering "AGI" in the same sentence as any of these models. Now, I'm not a hater, and for many tasks they are amazing, but they are, as of now, not even close to AGI, by any reasonable definition.
- AIPedant 1y agoI think it is AGI, seriously. Try asking it lots of questions, and then ask yourself: just how much smarter was I expecting AGI to be? That's his whole argument!!!! This is so frustrating coming from a public intellectual. "You don't need rigorous reasoning to answer these questions, baybeee, just go with your vibes." Complete and total disregard for scientific thinking, in favor of confirmation bias and ideology.
- neonbjb 1y agoI work for openai. o4-mini gets much closer (but I'm pretty sure it fumbles at the last moment): https://chatgpt.com/share/680031fb-2bd0-8013-87ac-941fa91cea9d https://chatgpt.com/share/680031fb-2bd0-8013-87ac-941fa91cea... We're pretty bad at model naming and communicating capabilities (in our defense, it's hard!), but o4-mini is actually a _considerably_ better vision model than o3, despite the benchmarks. Similar to how o3-mini-high was a much better coding model than o1. I would recommend using o4-mini-high over o3 for any task involving vision.
- jonahx 1y agoThanks for the reply. I am not sure the vision is the failing point here, but logic. I routinely try to get these models to solve difficult puzzles or coding challenges (the kind that a good undergrad math major could probably solve, but that most would struggle with). They fail almost always. Even with help. For example, JaneStreet monthly puzzles. Surprisingly, the new o3 was able to solve this months (previous models were not), which was an easier one. Believe me, I am not trying to minimize the overall achievement -- what it can do incredible -- but I don't believe the phrase AGI should even be mentioned until we are seeing solutions to problems that most professional mathematicians would struggle with, including solutions to unsolved problems. That might not be enough even, but that should be the minimum bar for even having the conversation.
- jawiggins 1y agoIn the examples they demonstrate tool use in the reasoning loop. The models pretty impressively recognize they need some external data, and either complete a web search, or write and execute python to solve intermediate steps. To the extent that reasoning is noisy and models can go astray during it, this helps inject truth back into the reasoning loop. Is there some well known equivalent to Moores Law for token use? We're headed in a direction where LLM control loops can run 24/7 generating tokens to reason about live sensor data, and calling tools to act on it.
- eric-p7 1y agoBabe wake up a new LLM just dropped.
- simianwords 1y agoI feel like the only reason O3 is better than O1 just due to the tool usage. With tool use O1 could be similar to O3.
- Topfi 1y agoI have barely found time to gauge 4.1s capabilities, so at this stage, I’d rather focus on the ever worsening names these companies bestow upon their models. To say that I the USB-IF have found their match would be an understatement.
- deleted 1y ago[deleted]
- neya 1y agoThe most annoying part of all this is they replaced o1 with o3 without any notices or warnings. This is why I hate proprietary models.
- sebzim4500 1y agoMeanwhile we have people elsewhere in the thread complaining about too many models. Assuming OpenAI are correct that o3 is strictly an improvement over o1 then I don't see why they'd keep o1 around. When they upgrade gpt-o4 they don't let you use the old version, after all.
- kgeist 1y ago>Assuming OpenAI are correct that o3 is strictly an improvement over o1 then I don't see why they'd keep o1 around. Imagine if every time your favorite SaaS had an update, they renamed the product. Yesterday you were using Slack S7, and today you're suddenly using Slack 9S-o. That was fine in the desktop era, when new releases happened once a year - not every few weeks. You just can't keep up with all the versions. I think they should just stick with one brand and announce new releases as just incremental updates to that same brand/product (even if the underlying models are different): "the DeepSearch Update" or "The April 2025 Reasoning Update" etc. The model picker should be replaced entirely with a router that automatically detects which underlying model to use. Power users could have optional checkboxes like "Think harder" or "Code mode" as settings, if they want to guide the router toward more specialized models.
- I_am_tiberius 1y agoWhat is again the advantage of pro over plus subscriptions?
- postmaster 1y ago> We expect to release OpenAI o3‑pro in a few weeks with full tool support. For now, Pro users can still access o1‑pro.
- I_am_tiberius 1y agoOk, so currently they pay for nothing (or is o1-pro superior to o3?).
- djohnston 1y agoAny quick impressions of o3 vs o1? We've got one inference in our product that only o1 has seemed to handle well, wondering if o3 can replace it.
- sebzim4500 1y agoThey are replacing o1 with o3 in the UI, at least for me, so they must be pretty confident it is a strict improvement.
- osigurdson 1y agoI have a very basic / stupid "Turing test" which is just to write a base 62 converter in C#. I would think this exact thing would be in github somewhere (thus in the weights) but has always failed for me in the past (non-scientific / didn't try every single model). Using o4-mini-high, it actually did produce a working implementation after a bit of prompting. So yeah, today, this test passed which is cool.
- sebzim4500 1y agoUnless I'm misunderstanding what you are asking the model to do, Gemini 2.5 pro just passed this easily. https://g.co/gemini/share/e2876d310914 https://g.co/gemini/share/e2876d310914
- osigurdson 1y agoAs I mentioned, this is not a scientific test but rather just something that I have tried from time to time and has always (shockingly in my opinion) failed but today worked. It takes a minute of two of prompting, is boring to verify and I don't remember exactly which models I have used. It is purely a personal anecdote, nothing more. However, looking at the code that Gemini wrote in the link, it does the same thing that other LLMs often do, which is to assume that we are encoding individual long values. I assume there must be a github repo or stackoverflow question in the weights somewhere that is pushing it in this direction but it is a little odd. Naturally, this isn't the kind encoder that someone would normally want. Typically it should encode a byte array and return a string (or maybe encode / decode UTF8 strings directly). Having the interface use a long is very weird and not very useful. In any case, I suspect with a bit more prompting you might be able to get gemini to do the right thing.
- jiggawatts 1y agoSimilarly, many of my informal tests have started passing with Gemini 2.5 that never worked before, which makes the 2025 era of AI models feel like a step change to me.
- int_19h 1y ago
- pcdoodle 1y agoIt seems to be getting better. I used to use my custom "Turbo Chad" GPT based on 4o and now the default models are similar. Is it learning from my previous annoyances? It has been getting better IMO.
- fpgaminer 1y agoOn the vision side of things: I ran my torture test through it, and while it performed "well", about the same level as 4o and o1, it still fails to handle spatial relationships well, and did hallucinate some details. OCR is a little better it seems, but a more thorough OCR focused test would be needed to know for sure. My torture tests are more focused on accurately describing the content of images. Both seem to be better at prompt following and have more up to date knowledge. But honestly, if o3 was only at the same level as o1, it'd still be an upgrade since it's cheaper. o1 is difficult to justify in the API due to cost.
- deleted 1y ago[deleted]
- taytus 1y agoThis is a mess. I do follow AI news, and do no know if this is "better/faster/cheaper" than 4.1 Why are they doing this?
- rsanheim 1y ago`ETOOMANYMODELS` Is there a reputable, non-blogspam site that offers a 'cheat sheet' of sorts for what models to use, in particular for development? Not just openAI, but across the main cloud offerings and feasible local models? I know there are the benchmarks, and directories like huggingface, and you can get a 'feel' for things by scanning threads here or other forums. I'm thinking more of something that provides use-case tailored "top 3" choices by collecting and summarizing different data points. For example: * agent & tool based dev (cloud) - [top 3 models] * agent & tool based dev (local) - m1, m2, m,3 * code review / high level analysis - ... * general tech questions - ... * technical writing (ADRs, needs assessments, etc) - ... Part of the problem is how quickly the landscape changes everyday, and also just relying on benchmarks isn't enough: it ignores cost, and more importantly ignores actual user experience (which I realize is incredibly hard to aggregate & quantify).
- ac29 1y ago> Is there a reputable, non-blogspam site that offers a 'cheat sheet' of sorts for what models to use, in particular for development? Below is a spreadsheet I bookmarked from a previous HN discussion. Its information dense but you can just look at the composite scores to get a quick idea how things compare. https://docs.google.com/spreadsheets/u/1/d/1foc98Jtbi0-GUsNySddvL0b2a7EuVQw8MoaQlWaDT-w/htmlview https://docs.google.com/spreadsheets/u/1/d/1foc98Jtbi0-GUsNy...
- departed 1y agoLMArena might have some of the information you are looking for. It offers rankings of LLM models across main cloud offerings, and I feel that its evaluation method, human prompting and voting, is closer to real-world use case and less prone to data contamination than benchmarks. https://lmarena.ai/ https://lmarena.ai/ In the "Leaderboard">"Language" tab, it lists the top models in various categories such as overall, coding, math, and creative writing. In the "Leaderboard">"Price Analysis" tab, it shows a chart comparing models by cost per million tokens. In the "Prompt-to-Leaderboard" tab, there is even an LLM to help you find LLMs -- you enter a prompt, and it will find the top models for your particular prompt.
- Carbonhell 1y ago
- Sol- 1y agoInteresting that using tools to zoom around the image is useful for the model. I was kind of assuming that these models were beyond such things and could attend to all aspects image simultaneously anyway, but perhaps their input is still limited in the resolution? Very cool, in any case, spooky progress as always.
- littlestymaar 1y agoThere's just a certain amount of things the image encoder can process at once. It's pretty apparent when you give the models a big table in an image.
- steinvakt2 1y agoBut isn't this basically what the conv layer does...?
- sbochins 1y agoSo far with my random / coding design question that I asked with o1 last week, it did substantially better with o3. It’s more like a mid level engineer and less like a intern.
- tymscar 1y agoGave Codex a go with o4-mini and it's disappointing... Here you can see my tries. It fully fails on something a mid engineer can do after getting used to the tools: https://xcancel.com/Tymscar/status/1912578655378628847 https://xcancel.com/Tymscar/status/1912578655378628847
- dr_kiszonka 1y agoI want to be excited about this but after chatting with 4.1 about a simple app screenshot and it continuously forgetting and hallucinating, I am increasingly sceptical of Open AI's announcements. (No coding involved, so the context window was likely < 10% full.)
- iandanforth 1y agoo3 failed the first test I gave it. I wanted it to create a bar chart using Python of the first 10 Fibonacci numbers (did this easily), and then use that image as input to generate an info-graphic of the chart with an animal theme. It failed in two ways. It didn't have access to the visual output from python and, when I gave it a screenshot of that output, it failed in standard GenAI fashion by having poor / incomplete text and not adhering exactly to bar heights, which were critical in this case. So one failure that could be resolved with better integration on the back end and then an open problem with image generation in general.
- erikw 1y agoInteresting... I asked o3 for help writing a flake so I could install the latest Webstorm on NixOS (since the one in the package repo is several months old), and it looks like it actually spun up a NixOS VM, downloaded the Webstorm package, wrote the Flake, calculated the SHA hash that NixOS needs, and wrote a test suite. The test suite indicates that it even did GUI testing- not sure whether that is a hallucination or not though. Nevertheless, it one-shotted the installation instructions for me, and I don't see how it could have calculated the package hash without downloading, so I think this indicates some very interesting new capabilities. Highly impressive.
- peterldowns 1y agoIf it can write a nixos flake it's significantly smarter than the average programmer. Certainly smarter than me, one-shotting a flake is not something I'll ever be able to do — usually takes me about thirty shots and a few minutes to cool off from how mad I am at whoever designed this fucking idiotic language. That's awesome.
- deleted 1y ago[deleted]
- ZeroTalent 1y agoI was a major contributor of Flake. What in particular is so idiotic in your opinion?
- yjftsjthsd-h 1y agoFWIW, they said the language was bad, not specifically flakes. IMHO, nix is super easy if you already know Haskell (possibly others in that family). If you don't, it's extremely unintuitive.
- peterldowns 1y agoI use flakes a lot and I think both flakes and the Nix language are beyond comprehension. Try searching duckduckgo or google for “what is nix flakes” or “nix flake schema” and take an honest read at the results. Insanely complicated and confusing answers, multiple different seemingly-canonical sources of information. Then go look at some flakes for common projects; the almost necessary usage of things like flake-compat and flake-util, the many-valid-approaches to devshell and package definitions, the concepts of “apps” in addition to packages. All very complicated and crazy! Thank you for your service, I use your work with great anger (check my github I really do!)
- AcerbicZero 1y agoI can't even get ChatGPT to tell me which chatgpt to use.
- lubitelpospat 1y agoSooo... are any of these (or their distils) getting open-sourced/open-weighted?
- croemer 1y agoI wonder where o3 and o4-mini will land on the LMarena leaderboard. When might we see them there?
- jumploops 1y agoThe big step function here seems to be RL on tool calling. Claude 3.7/3.5 are the only models that seem to be able to handle "pure agent" usecases well (agent in a loop, not in an agentic workflow scaffold[0]). OpenAI has made a bet on reasoning models as the core to a purely agentic loop, but it hasn't worked particularly well yet (in my own tests, though folks have hacked a Claude Code workaround[1]). o3-mini has been better at some technical problems than 3.7/3.5 (particularly refactoring, in my experience), but still struggles with long chains of tool calling. My hunch is that these models were tuned _with_ OpenAI Codex[2], which is presumably what Anthropic was doing internally with Claude Code on 3.5/3.7 tl;dr - GPT-3 launched with completions (predict the next token), then OpenAI fine-tuned that model on "chat completions" which then led GPT-3.5/GPT-4, and ultimately the success of ChatGPT. This new agent paradigm, requires fine-tuning on the LLM interacting with itself (thinking) and with the outside world (tools), sans any human input. [0]https://www.anthropic.com/engineering/building-effective-agents#what-are-agents https://www.anthropic.com/engineering/building-effective-age... [1]https://github.com/1rgs/claude-code-proxy https://github.com/1rgs/claude-code-proxy [2]https://openai.com/index/openai-codex/ https://openai.com/index/openai-codex/
- waltercool 1y ago[dead]
- siva7 1y agoSo what are they selling with the 200 dollar subscription? Only a model that has now caught up with their competitor who sells for 1/10 of their price?
- hybrid_study 1y agoDoesn't achieving AGI mean the beginning of the end of humanity's current economic model? I'm not sure I understand the presumption by many that achieving AGI is just another step in some company's offering.
- BriggyDwiggs42 1y agoNo you see because everyone will become agi engineers actually that makes sense and is going to happen
- JFingleton 1y agoMost days I feel the same. Other days I remember that humans like "handmade" furniture, and live performances, and unique styles, and human contact. Perhaps there's life in us still?
- M4v3R 1y agoOk, I’m a bit underwhelmed. I’ve asked it a fairly technical question, about a very niche topic (Final Fantasy VII reverse engineering): https://chatgpt.com/share/68001766-92c8-8004-908f-fb185b7549d9 https://chatgpt.com/share/68001766-92c8-8004-908f-fb185b7549... With right knowledge and web searches one can answer this question in a matter of minutes at most. The model fumbled around modding forums and other sites and did manage to find some good information but then started to hallucinate some details and used them in the further research. The end result it gave me was incorrect, and the steps it described to get the value were totally fabricated. What’s even worse in the thinking trace it looks like it is aware it does not have an answer and that the 399 is just an estimate. But in the answer itself it confidently states it found the correct value. Essentially, it lied to me that it doesn’t really know and provided me with an estimate without telling me. Now, I’m perfectly aware that this is a very niche topic, but at this point I expect the AI to either find me a good answer or tell me it couldn’t do it. Not to lie me in the face. Edit: Turns out it’s not just me: https://x.com/transluceai/status/1912552046269771985?s=46 https://x.com/transluceai/status/1912552046269771985?s=46
- siva7 1y agoIt can imitate its creator. We reached AGI.
- casinoplayer0 1y agoI wanted to believe. But not now.
- werdnapk 1y agoI've used AI with "niche" programming questions and it's always a total let down. I truly don't understand this "vibe coding" movement unless everyone is building todo apps.
- hatefulmoron 1y agoIt's incredible when I ask Claude 3.7 a question about Typescript/Python and it can generate hundreds of lines of code that are pretty on point (it's usually not exactly correct on first prompt, but it's coherent). I've recently been asking questions about Dafny and Lean -- it's frustrating that it will completely make up syntax and features that don't exist, but still speak to me with the same confidence as when it's talking about Typescript. It's possible that shoving lots of documentation or a book about the language into the context would help (I haven't tried), but I'm not sure if it would make up for the model's lack of "intuition" about the subject.
- highfrequency 1y agoThe benchmarks reference o3-low, medium and high. What is plain “o3”? Is that medium?
- DrNosferatu 1y agoAny good leaderboard where all the very latest models are compared?
- simonw 1y agoHere's a summary of this conversation so far, generated using o3 after 306 comments. This time I ran it like so: llm install llm-openai-plugin llm install llm-hacker-news llm -m openai/o3 -f hn:43707719 -s 'Summarize the themes of the opinions expressed here. For each theme, output a markdown header. Include direct "quotations" (with author attribution) where appropriate. You MUST quote directly from users when crediting them, with double quotes. Fix HTML entities. Output markdown. Go long. Include a section of quotes that illustrate opinions uncommon in the rest of the piece' https://gist.github.com/simonw/a35f39b070978e703d9eb8b1aa7c003f#response https://gist.github.com/simonw/a35f39b070978e703d9eb8b1aa7c0... - cost 2,684 input, 2,452 output (of which 896 were reasoning tokens) which is 12.492 cents. Then again with o4-mini using the exact same content (hence the hash ID for -f): llm -m openai/o4-mini \ -f f16158f09f76ab5cb80febad60a6e9d5b96050bfcf97e972a8898c4006cbd544 \ -s 'Summarize the themes of the opinions expressed here. For each theme, output a markdown header. Include direct "quotations" (with author attribution) where appropriate. You MUST quote directly from users when crediting them, with double quotes. Fix HTML entities. Output markdown. Go long. Include a section of quotes that illustrate opinions uncommon in the rest of the piece' Output: https://gist.github.com/simonw/b11ba0b11e71eea0292fb6adaf9cd1f0 https://gist.github.com/simonw/b11ba0b11e71eea0292fb6adaf9cd... Cost 2,684 input, 2,681 output (of which 1,088 reasoning tokens) = 1.4749 cents The above uses these two plugins: https://github.com/simonw/llm-openai-plugin https://github.com/simonw/llm-openai-plugin and https://github.com/simonw/llm-hacker-news https://github.com/simonw/llm-hacker-news - taking advantage of new -f "fragments" feature I released last week: https://simonwillison.net/2025/Apr/7/long-context-llm/ https://simonwillison.net/2025/Apr/7/long-context-llm/
- siliconc0w 1y agoPlease just give me a best value and a highest performance model.
- deleted 1y ago[deleted]
- throwaway13337 1y agoo4-mini is available on vs code. I've been playing with it for the last couple of hours. It's quite fast for a thinking model. It's also super concise with code. Where claude 3.7 and gemini 2.5 will write a ton, o4-mini will write a tiny portion of it accomplishing the same task. On the flip side, in its conciseness, it's more lazy with implementation than the other leading models missing features. For fixing very complex typescript types, I've previously found that o1 outperformed the others. o4-mini seems to understand things well here. I still think gemini will continue to be my favorite model for code. It's more consistent and follows instructions better. However, openAI's more advanced models have a better shot at providing a solution when gemini and claude are stuck. Maybe there's a win here in having o4-mini or o3 do a first draft for conciseness, revise with gemini to fill in what's missed (but with a base that is not overdone), and then run fixes with o4-mini. Things are still changing quite quickly.
- famouswaffles 1y agoo3 joins gemini-2.5-pro as the only other model that can pace long form creative writing properly when details about the story are provided.
- roskelld 1y agoAfter refreshing the browser I see that the old o3-mini-high has gone now so I continued my coding task conversation with o4-mini-high. In two separate conversations it butchered things in a way that I never saw o3-mini-high do. In one case it rewrote working code without reason, breaking it, in the other it took a function I asked it to apply a code fix to and it instead refactored it with a different and unrelated function that was part of an earlier bit of chat history. I notice too that it employs a different style of code where it often puts assignment on a different line, which looks like it's trying to maintain an ~80 character line limit, but does so in places where the entire line of code is only about 40 characters.
- upbeat_general 1y agoNot saying it’s for sure the case but it might be that the model gets confused by OOD text from the other model whereas it expects its own text to be online from itself (particularly if the CoT is used as context for later conversations).
- caseyy 1y agoThe demo video is very impressive, and it shows what AI could be. Our current models are unreliable in research, but if they were reliable, then what's shown alone would be better than AGI. There are 8 billion+ instances of general intelligence on the planet; there isn't a shortage. I'd rather see AI do data science and applied math at computer speeds. Those are the hard problems, a lot of the AGI problems (to human brains) are easy.
- benoau 1y ago> Downloaded an untouched char.lgp from the current Steam build (1.0.9) to make sure the count reflects the shipping game rather than a modded archive. How?
- feelingsonice 1y agoI’m having very mixed feelings about it. I’m using o3 to help me parse and understand a book about statistics and ML, it’s very dense in math. On one hand the answers became a lot more comprehensive and deep. It’s now able to give me very advanced explanations. On the other hand, it started overloading the answers with information. Entire concepts became single sentence summaries. Complex topics and theorems became acronyms. In a way I’m feeling overwhelmed by the information it’s now throwing at me. I can’t tell if it’s actually smarter or just too complicated for me to understand.
- deleted 1y ago[deleted]
- shanecp 1y agoHere are some notes I made to understand each of these models and when to use them. # OpenAI Models ## Reasoning Models (o-series) - All `oX` (o-series aka `omni`) models are reasoning models. - Use these for complex, multi-step, reasoning tasks. ## Flagship/Core Models - All `x.x` and `Xo` models are the core models. - Use these for one-shot results - Examples: 4o, 4.1 ## Cost Optimized - All `-mini`, `-nano` are cheaper, faster models. - Use these for high-volume, low effort tasks. ## Flagship vs Reasoning (o-series) Models - Latest flagship model = 4.1 - Latest reasoning model = o3 - The flagship models are general purpose, typically with larger context windows. These rely mostly on pattern matching. - The reasoning models are trained with extended chain-of-thought and reinforcement learning models. They work best with tools, code and other multi-step workflows. Because tools are used, the accuracy will be higher. # List of Models ## 4o (omni) - 128K context window - complex multimodal, applications requiring the top level of reliability and nuance ## 4o-mini - 128K context window - Use: multimodal reasoning for math, coding, and structured outputs - Use: Cheaper than `4o`. Use when you can trade off accuracy vs speed/cost. - Dont Use: When high accuracy is needed ## 4.1 - 1M context window - Use: For large context ingest, such as full codebases - Use: For reliable instruction following, comprehension - Dont Use: For high volume/faster tasks ## 4.1-mini - 1M context window - Use: For large context ingest - Use: When a tradeoff can be made with accuracy vs speed ## 4.1-nano - 1M context window - Use: For high-volume, near-instant responses - Dont Use: When accuracy is required - Examples: classification, autocompletion, short-answers ## o3 - 200K context window - Use: for the most challenging reasoning tasks in coding, STEM, and vision that demand deep chain‑of‑thought and tool use - Use: Agentic workflows leveraging web search, Python execution, and image analysis in one coherent loop - Dont Use: For simple tasks, where lighter model will be faster and cheaper. ## o4-mini - 200K context window - Use: High-volume needs where reasoning and cost should be balanced - Use: For high throughput applications - Dont Use: When accuracy is critical ## o4-mini-high - 200K context window - Use: When o4-mini results are not satisfactory, but before moving to o3. - Use: Compex tool-driven reasoning, where o4-mini results are not satisfactory - Dont Use: When accuracy is critical ## o1-pro-mode - 200K context window - Use: Highly specialized science, coding, or reasoning jobs that benefit from extra compute for consistency - Dont Use: For simple tasks ## Models Sorted for Complex Coding Tasks (my opinion) 1. o3 2. Gemini 2.5 Pro 3. Claude 3.7 2. o1-pro-mode 3. o4-mini-high 4. 4.1 5. o4-mini
- andai 1y agoThe most striking difference to me is that o3 and o4 know when the web search tool is unavailable, and will tell you they can't answer a question that requires it. While 4o and (sadly) 4.1 will just make up a bunch of nonsense. I'm simultaneously impressed that they can do that, and also wondering why the heck that's so impressive (isn't "is this tool in this list?" something GPT-3 was able to handle?) and why 4.1 still fails at it too—especially considering it's hyped as the agentic coder model! That's pretty damning for the general intelligence aspect of it, that they apparently had to special-case something so trivial... and I say that as someone who's really optimistic about this stuff! That being said, the new "enhanced" web search seems great so far, and means I can finally delete another stupid 10 line Python script from 2023 that I shouldn't have needed in the first place ;) (...Now if they'd just put 4.1 in the Chat... why the hell do I need to use a 3rd party UI for their best model!)
- mianos 1y agoI have been using o4-mini-high today. Most of the time for a file longer than 100 lines it stops generating randomly and won't complete a file unless I re-prompt it with the end of the missing file. As usual, it's a frustrating experience for anything more complex than the usual problems everyone else does.
- thom 1y agoo4 is doing a better job than o3 on my current project, and while this isn’t really a priority, its personality is somehow far more engaging now.
- immibis 1y agohttps://transluce.org/investigating-o3-truthfulness https://transluce.org/investigating-o3-truthfulness Some interesting hallucinations going on here!
- deleted 1y ago[deleted]
- klasko 1y agoFWIW, o4-mini-high does not feel better o3-mini-high for working on fairly simply econ theory proofs. It does feel faster. And both elementary mistakes.
- Davidzheng 1y agoI find it worse than Gemini 2.5 Pro at math research.
- jdlyga 1y agoAt this point, it's like comparing the iPhone 5s vs the iPhone 6. The upgrades are still noticeable, but it's nowhere the huge jump between GPT 3.5 and GPT 4.
- bloqs 1y agoI'm confused. I typically use o1 for all of my questions. Now it's disappeared. Is o3 a better model?
- euph0ria 1y agoYes, in almost all aspects if you do not use the o1-pro. o3-pro is not available yet.
- kurtis_reed 1y agoI thought they weren't going to release o3 and it would just be bundled into "GPT-5".
- wg0 1y agoIf you download GIMP, Blender etc - every user would have to report exactly the same experience mostly given the hardware is recent. In this thread however - there are varying experiences from amazing to awful. I'm not saying anyone is wrong but all I'm saying is that this wide range of operational accuracy is what will pop the AI bubble eventually in that they can't be reliably deployed almost anywhere with any certainty or guarantees of any sorts.
- momoelz 1y agoI find o4 very bad at coding. I tried to improve a script created by 3.5 mini-high with o4 mini-high and it doesn't return nearly as good results as what i used to get by o3.5
- sks38317 1y agothanks for your information!
- rpgbr 1y agoThis post[1] is highlighted by Techmeme: >I'm obsessed with o3. It's way better than the previous models. It just helped me resolve a psychological/emotional problem I've been dealing with for years in like 3 back-and-forths (one that wasn't socially acceptable to share, and those I shared it with didn't/couldn't help) Genuinely intrigued by what kind of “psychological/emotional problem I've been dealing with for years” could an AI solve in a matter of hours after its release. [1] https://x.com/carmenleelau/status/1912645771955962300 https://x.com/carmenleelau/status/1912645771955962300
- AbuAssar 1y agoI noticed that OpenAI don't compare their models to third party models in their announcement posts, unlike google, meta and the others.
- deleted 1y ago[deleted]