19 ms·
GPT-4.1 in the API
- elias_t 1y agoDoes someone have the benchmarks compared to other models?
- cbg0 1y agoclaude 3.7 no thinking (diff) - 60.4% claude 3.7 32k thinking tokens (diff) - 64.9% GPT-4.1 (diff) - 52.9% (stat is from the blog post) https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/
- i_love_retros 1y agoI feel overwhelmed
- porphyra 1y agopretty wild versioning that GPT 4.1 is newer and better in many regards than GPT 4.5.
- mhh__ 1y agoI think they're doing it deliberately at this point
- hmottestad 1y agoTomorrow they are releasing the open source GPT-1.4 model :P
- mhh__ 1y agoI'm apparently dyslexic enough that I only just noticed the joke 2 days later
- asdev 1y agoit's worse on nearly every benchmark
- brokensegue 1y agono? it's better on AIME '24, Multilingual MMLU, SWE-bench, Aider’s polyglot, MMMU, ComplexFuncBench and it ties on a lot of benchmarks
- asdev 1y agolook at all the graphs in the article
- brokensegue 1y agothe data i posted all came from the graphs/charts in the article
- porphyra 1y agoOpenAI themselves said > One last note: we’ll also begin deprecating GPT-4.5 Preview in the API today as GPT-4.1 offers improved or similar performance on many key capabilities at lower latency and cost. GPT-4.5 in the API will be turned off in three months, on July 14, to allow time to transition (and GPT 4.5 will continue to be available in ChatGPT). https://x.com/OpenAIDevs/status/1911860805810716929 https://x.com/OpenAIDevs/status/1911860805810716929
- exizt88 1y agoFor conversational AI, the most significant part is GPT-4.1 mini being 2x faster than GPT-4o at basically the same reasoning capabilities.
- bakugo 1y ago> We will also begin deprecating GPT‑4.5 Preview in the API, as GPT‑4.1 offers improved or similar performance on many key capabilities at much lower cost and latency. GPT‑4.5 Preview will be turned off in three months, on July 14, 2025, to allow time for developers to transition. Well, that didn't last long.
- WorldPeas 1y agoso we're going back... .4 of a gpt? make it make sense openai..
- huxley 1y agoThink of 4.5 as being the lacklustre major upgrade to a software package, pick one maybe Photoshop or whatever. The 4.0 version is still available and most people are continuing to use it, then suddenly 4.0 gets a small upgrade which makes it considerably better and the vendor starts talking about how the real future is in 5.0. I wish OpenAI had invented this but it’s not that uncommon.
- deleted 1y ago[deleted]
- oidar 1y agoI need an AI to understand the naming conventions that OpenAI is using.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- fusionadvocate 1y agoThey envy the USB committee.
- ZeroCool2u 1y agoNo benchmark comparisons to other models, especially Gemini 2.5 Pro, is telling.
- dmd 1y agoGemini 2.5 Pro gets 64% on SWE-bench verified. Sonnet 3.7 gets 70% They are reporting that GPT-4.1 gets 55%.
- hmottestad 1y agoAre those with «thinking» or without?
- energy123 1y agoWith
- chaos_emergent 1y agobased on their release cadence, I suspect that o4-mini will compete on price, performance, and context length with the rest of these models.
- hecticjeff 1y agoo4-mini, not to be confused with 4o-mini
- sanxiyn 1y agoSonnet 3.7's 70% is without thinking, see https://www.anthropic.com/news/claude-3-7-sonnet https://www.anthropic.com/news/claude-3-7-sonnet
- aledalgrande 1y agoThe thinking tokens (even just 1024) make a massive difference in real world tasks with 3.7 in my experience
- egeozcan 1y ago
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- codingwagie 1y agoGPT-4.1 probably is a distilled version of GPT-4.5 I dont understand the constant complaining about naming conventions. The number system differentiates the models based on capability, any other method would not do that. After ten models with random names like "gemini", "nebula" you would have no idea which is which. Its a low IQ take. You dont name new versions of software as completely different software Also, Yesterday, using v0, I replicated a full nextjs UI copying a major saas player. No backend integration, but the design and UX were stunning, and better than I could do if I tried. I have 15 years of backend experience at FAANG. Software will get automated, and it already is, people just havent figured it out yet
- rvz 1y ago> Yesterday, using v0, I replicated a full nextjs UI copying a major saas player. No backend integration, but the design and UX were stunning, and better than I could do if I tried. Exactly. Those who do frontend or focus on pretty much anything Javascript are, how should I say it? Cooked? > Software will get automated The first to go are those that use JavaScript / TypeScript engineers have already been automated out of a job. It is all over for them.
- codingwagie 1y agoYeah its over for them. Complicated business logic and sprawling systems are what are keeping backend safe for now. But the big front end code bases where individual files (like react components) are largely decoupled from the rest of the code base is why front end is completely cooked
- camdenreslink 1y agoI have a medium-sized typescript personal project I work on. It probably has 20k LOC of well organized typescript (react frontend, express backend). I also have somewhat comprehensive docs and cursor project rules. In general I use Cursor in manual mode asking it to make very well scoped small changes (e.g. “write this function that does this in this exact spot”). Yesterday I needed to make a largely mechanical change (change a concept in the front end, make updates to the corresponding endpoints, update the data access methods, update the database schema). This is something very easy I would expect a junior developer to be able to accomplish. It is simple, largely mechanical, but touches a lot of files. Cursor agent mode puked all over itself using Gemini 2.5. It could summarize what changes would need to be made, but it was totally incapable of making the changes. It would add weird hard coded conditions, define new unrelated files, not follow the conventions of the surrounding code at all. TLDR; I think LLMs right now are good for greenfield development (create this front end from scratch following common patterns), and small scoped changes to a few files. If you have any kind of medium sized refactor on an existing code base forget about it.
- rvz 1y agoThe big change about this announcement is the 1M context window on all models. But the price is what matters.
- croemer 1y agoNothing compared to Llama 4's 7M. What matters is how well it performs with such long context, not what the technical maximum is.
- polytely 1y agoIt seems that OpenAI is really differentiating itself in the AI market by developing the most incomprehensible product names in the history of software.
- croes 1y agoThey learned from the best: Microsoft
- deleted 1y ago[deleted]
- pixl97 1y ago"Hey buddy, want some .Net, oh I mean dotnet"
- nivertech 1y agoGPT 4 Workgroups
- amarcheschi 1y agoGpTeams Classic
- greenavocado 1y agoMicrosoft Neural Language Processing Hyperscale Datacenter Enterprise Edition 4.1 A massive transformer-based language model requiring: - 128 Xeon server-grade CPUs - 25,000MB RAM minimum (40,000MB recommended) - 80GB hard disk space for model weights - Dedicated NVIDIA Quantum Accelerator Cards (minimum 8) - Enterprise-grade cooling solution - Dedicated 30-amp power circuit - Windows NT Advanced Server with Parallel Processing Extensions ~ Features: - Natural language understanding and generation - Context window of 8,192 tokens - Enterprise security compliance module - Custom prompt engineering interface - API gateway for third-party applications *Includes 24/7 on-call Microsoft support team and requires dedicated server room with raised floor cooling
- jmount 1y agoOr Intel.
- yberreby 1y ago> Note that GPT‑4.1 will only be available via the API. In ChatGPT, many of the improvements in instruction following, coding, and intelligence have been gradually incorporated into the latest version (opens in a new window) of GPT‑4o, and we will continue to incorporate more with future releases. The lack of availability in ChatGPT is disappointing, and they're playing on ambiguity here. They are framing this as if it were unnecessary to release 4.1 on ChatGPT, since 4o is apparently great, while simultaneously showing how much better 4.1 is relative to GPT-4o. One wager is that the inference cost is significantly higher for 4.1 than for 4o, and that they expect most ChatGPT users not to notice a marginal difference in output quality. API users, however, will notice. Alternatively, 4o might have been aggressively tuned to be conversational while 4.1 is more "neutral"? I wonder.
- themanmaran 1y agoI disagree. From the average user perspective, it's quite confusing to see half a dozen models to choose from in the UI. In an ideal world, ChatGPT would just abstract away the decision. So I don't need to be an expert in the relatively minor differences between each model to have a good experience. Vs in the API, I want to have very strict versioning of the models I'm using. And so letting me run by own evals and pick the model that works best.
- florakel 1y ago> it's quite confusing to see half a dozen models to choose from in the UI. In an ideal world, ChatGPT would just abstract away the decision Supposedly that’s coming with GPT 5.
- yberreby 1y agoI agree on both naming on stability. However, this wasn't my point. They still have a mess of models in ChatGPT for now, and it doesn't look like this is going to get better immediately (even though for GPT-5, they ostensibly want to unify them). You have to choose among all of them anyway. I'd like to be able to choose 4.1.
- Tiberium 1y ago
- meetpateltech 1y agoGPT-4.1 Pricing (per 1M tokens): gpt-4.1 - Input: $2.00 - Cached Input: $0.50 - Output: $8.00 gpt-4.1-mini - Input: $0.40 - Cached Input: $0.10 - Output: $1.60 gpt-4.1-nano - Input: $0.10 - Cached Input: $0.025 - Output: $0.40
- minimaxir 1y agoThe cached input price is notable here: previously with GPT-4o it was 1/2 the cost of raw input, now it's 1/4th. It's still not as notable as Claude's 1/10th the cost of raw input, but it shows OpenAI's making improvements in this area.
- persedes 1y agoUnless that has changed, anthropics (and gemini) caches are opt-in though if I recall, openai automatically chaches for you.
- glenstein 1y agoAwesome, thank you for posting. As someone who regularly uses 4o mini from the API, any guesses or intuitions about the performance of Nano? I'm not as concerned about nomenclature as other people, which I think is too often reacting to a headline as opposed to the article. But in this case, I'm not sure if I'm supposed to understand nano as categorically different than many in terms of what it means as a variation from a core model.
- pzo 1y agothey share in livestream that 4.1-nano is worse than 4o-mini - so nano is cheaper, faster and have bigger context but worse in intelligence. 4.1mini is smarter but there is price increase.
- twistslider 1y agoThe fact that they're raising the price for the mini models by 166% is pretty notable. gpt-4o-mini for comparison: - Input: $0.15 - Cached Input $0.075 - Output: $0.60
- minimaxir 1y agoIt's not the point of the announcement, but I do like the use of the (abs) subscript to demonstrate the improvement in LLM performance since in these types of benchmark descriptions I never can tell if the percentage increase is absolute or relative.
- croemer 1y agoTesting against unspecified other "leading" models allows for shenanigangs: > Qodo tested GPT‑4.1 head-to-head against other leading models [...] they found that GPT‑4.1 produced the better suggestion in 55% of cases The linked blog post goes 404: https://www.qodo.ai/blog/benchmarked-gpt-4-1/ https://www.qodo.ai/blog/benchmarked-gpt-4-1/
- gs17 1y agoThe post seems to be up now and seems to compare it slightly favorable to Claude 3.7.
- croemer 1y agoRight, now it's up and comparison against Claude 3.7 is better than I feared based on the wording. Though why does the OpenAI announcement talk of comparison against multiple leading models when the Qodo blog post only tests against Claude 3.7...
- runako 1y agoChatGPT currently recommends I use o3-mini-high ("great at coding and logic") when I start a code conversation with 4o. I don't understand why the comparison in the announcement talks so much about comparing with 4o's coding abilities to 4.1. Wouldn't the relevant comparison be to o3-mini-high? 4.1 costs a lot more than o3-mini-high, so this seems like a pertinent thing for them to have addressed here. Maybe I am misunderstanding the relationship between the models?
- zamadatix 1y ago4.1 is a pinned API variant with the improvements from the newer iterations of 4o you're already using in the app, so that's why the comparison focuses between those two. Pricing wise the per token cost of o3-mini is less than 4.1 but keep in mind o3-mini is a reasoning model and you will pay for those tokens too, not just the final output tokens. Also be aware reasoning models can take a long time to return a response... which isn't great if you're trying to use an API for interactive coding.
- ac29 1y ago> I don't understand why the comparison in the announcement talks so much about comparing with 4o's coding abilities to 4.1. Wouldn't the relevant comparison be to o3-mini-high? There are tons of comparisons to o3-mini-high in the linked article.
- Tiberium 1y agoVery important note: >Note that GPT‑4.1 will only be available via the API. In ChatGPT, many of the improvements in instruction following, coding, and intelligence have been gradually incorporated into the latest version If anyone here doesn't know, OpenAI does offer the ChatGPT model version in the API as chatgpt-4o-latest, but it's bad because they continuously update it so businesses can't reliably rely on it being stable, that's why OpenAI made GPT 4.1.
- croemer 1y agoSo you're saying that "ChatGPT-4o-latest (2025-03-26)" in LMarena is 4.1?
- exizt88 1y ago> chatgpt-4o-latest, but it's bad because they continuously update it Version explicitly marked as "latest" being continuously updated it? Crazy.
- sbarre 1y agoNo one's arguing that it's improperly labelled, but if you're going to use it via API, you might want consistency over bleeding edge.
- IanCal 1y agoLots of the other models are checkpoint releases, and latest is a pointer to the latest checkpoint. Something being continuously updated is quite different and worth knowing about.
- flakiness 1y agoBig focus on coding. It feels like a defensive move against Claude (and more recently, Gemini Pro) which became very popular in that regime. I guess they recently figured out some ways to train the model for these "agentic" coding through RL or something - and the finding is too new to apply 4.5 on time.
- modeless 1y agoNumbers for SWE-bench Verified, Aider Polyglot, cost per million output tokens, output tokens per second, and knowledge cutoff month/year: SWE Aider Cost Fast Fresh Claude 3.7 70% 65% $15 77 8/24 Gemini 2.5 64% 69% $10 200 1/25 GPT-4.1 55% 53% $8 169 6/24 DeepSeek R1 49% 57% $2.2 22 7/24 Grok 3 Beta ? 53% $15 ? 11/24 I'm not sure this is really an apples-to-apples comparison as it may involve different test scaffolding and levels of "thinking". Tokens per second numbers are from here: https://artificialanalysis.ai/models/gpt-4o-chatgpt-03-25/providers https://artificialanalysis.ai/models/gpt-4o-chatgpt-03-25/pr... and I'm assuming 4.1 is the speed of 4o given the "latency" graph in the article putting them at the same latency. Is it available in Cursor yet?
- meetpateltech 1y agoYes, it is available in Cursor[1] and Windsurf[2] as well. [1] https://twitter.com/cursor_ai/status/1911835651810738406 https://twitter.com/cursor_ai/status/1911835651810738406 [2] https://twitter.com/windsurf_ai/status/1911833698825286142 https://twitter.com/windsurf_ai/status/1911833698825286142
- cellwebb 1y agoAnd free on windsurf for a week! Vibe time.
- deleted 1y ago[deleted]
- tomjen3 1y agoIts available for free in Windsurf so you can try it out there. Edit: Now also in Cursor
- deleted 1y ago[deleted]
- jsnell 1y agohttps://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/ shows 73% rather than 69% for Gemini 2.5 Pro? Looks like they also added the cost of the benchmark run to the leaderboard, which is quite cool. Cost per output token is no longer representative of the actual cost when the number of tokens can vary by an order of magnitude for the same problem just based on how many thinking tokens the model is told to use.
- msp26 1y agoI was hoping for native image gen in the API but better pricing is always appreciated. Gemini was drastically cheaper for image/video analysis, I'll have to see how 4.1 mini and nano compare.
- oofbaroomf 1y agoI'm not really bullish on OpenAI. Why would they only compare with their own models? The only explanation could be that they aren't as competitive with other labs as they were before.
- greenavocado 1y agoSee figure 1 for up-to-date benchmarks https://github.com/KCORES/kcores-llm-arena https://github.com/KCORES/kcores-llm-arena (Direct Link) https://raw.githubusercontent.com/KCORES/kcores-llm-arena/refs/heads/main/scripts/llm_benchmark_results_normalized.png https://raw.githubusercontent.com/KCORES/kcores-llm-arena/re...
- poormathskills 1y agoGo look at their past blog posts. OpenAI only ever benchmarks against their own models.
- oofbaroomf 1y agoOh, ok. But it's still quite telling of their attitude as an organization.
- rvnx 1y agoIt's the same organization that kept repeating that sharing weights of GPT would be "too dangerous for the world". Eventually DeepSeek thankfully did something like that, though they are supposed to be the evil guys.
- kcatskcolbdi 1y agoI don't mind what they benchmark against as long as, when I use the model, it continues to give me better results than their competition.
- gizmodo59 1y agoApple compares against its own products most of the times.
- asdev 1y agoit's worse than 4.5 on nearly every benchmark. just an incremental improvement. AI is slowing down
- conradkay 1y agoIt's like 30x cheaper though. Probably just distilled 4.5
- GaggiX 1y agoIt's better on AIME '24, Multilingual MMLU, SWE-bench, Aider’s polyglot, MMMU, ComplexFuncBench while being much much cheaper and smaller.
- asdev 1y agoand it's worse on just as many benchmarks by a significant amount. as a consumer I don't care about cheapness, I want the maximum accuracy and performance
- GaggiX 1y agoAs a consumer you care about speed tho, and GPT-4.5 is extremely slow, at this point just use a reasoning model if you want the best of the best.
- simianwords 1y agoSorry what is the source for this?
- Nckpz 1y agoThey don't disclose parameter counts so it's hard to say exactly how far apart they are in terms of size, but based on the pricing it seems like a pretty wild comparison, with one being an attempt at an ultra-massive SOTA model and one being a model scaled down for efficiency and probably distilled from the big one. The way they're presented as version numbers is business nonsense which obscures a lot about what's going on.
- 1y ago
- theturtletalks 1y agoWith these being 1M context size, does that all but confirm that Quasar Alpha and Optimus Alpha were cloaked OpenAI models on OpenRouter?
- atemerev 1y agoYes, confirmed by citing Aider benchmarks: https://openai.com/index/gpt-4-1/ https://openai.com/index/gpt-4-1/ Which means that these models are _absolutely_ not SOTA, and Gemini 2.5 pro is much better, and Sonnet is better, and even R1 is better. Sorry Sam, you are losing the game.
- Tinkeringz 1y agoAren’t all of these reasoning models? Won’t the reasoning models of openAI benchmarked against these be a test of if Sam is losing?
- atemerev 1y agoThere is no OpenAI model better than R1, reasoning or not (as confirmed by the same Aider benchmark; non-coding tests are less objective, but I think it still holds). With Gemini (current SOTA) and Sonnet (great potential, but tends to overengineer/overdo things) it is debatable, they are probably better than R1 (and all OpenAI models by extension).
- vitorgrs 1y agoEven without reasoning, isn't Deepseek V3 from March better?
- maeil 1y agoSonnet 3.7 non-reasoning is better on its own. In fact even Sonnet 3.5-v2 is, and that was released 6 months ago. Now to be fair, they're close enough that there will be usecases - especially non-coding - where 4.1 beats it consistently. Also, 4.1 is quite a lot cheaper and faster. Still, OpenAI is clearly behind.
- 1y ago
- elashri 1y agoAre there any benchmarks or someone who did tests of performance of using this long max token models in scenarios where you actually use more of this token limit? I found from my experience with Gemini models that after ~200k that the quality drops and that it basically doesn't keep track of things. But I don't have any numbers or systematic study of this behavior. I think all providers who announce increased max token limit should address that. Because I don't think it is useful to just say that max allowed tokens are 1M when you basically cannot use anything near that in practice.
- gymbeaux 1y agoI’m not optimistic. It’s the Wild West and comparing models for one’s specific use case is difficult, essentially impossible at scale.
- enginoid 1y agoThere are some benchmarks such as Fiction.LiveBench[0] that give an indication and the new Graphwalks approach looks super interesting. But I'd love to see one specifically for "meaningful coding." Coding has specific properties that are important such as variable tracking (following coreference chains) described in RULER[1]. This paper also cautions against Single-Needle-In-The-Haystack tests which I think the OpenAI one might be. You really need at least Multi-NIAH for it to tell you anything meaningful, which is what they've done for the Gemini models. I think something a bit more interpretable like `pass@1 rate for coding turns at 128k` would so much more useful than "we have 1m context" (with the acknowledgement that good-enough performance is often domain dependant) [0] https://fiction.live/stories/Fiction-liveBench-Mar-25-2025/oQdzQvKHw8JyXbN87 https://fiction.live/stories/Fiction-liveBench-Mar-25-2025/o... [1] https://arxiv.org/pdf/2404.06654 https://arxiv.org/pdf/2404.06654
- jbentley1 1y agohttps://fiction.live/stories/Fiction-liveBench-Mar-25-2025/oQdzQvKHw8JyXbN87 https://fiction.live/stories/Fiction-liveBench-Mar-25-2025/o... IMO this is the best long context benchmark. Hopefully they will run it for the new models soon. Needle-in-a-haystack is useless at this point. Llama-4 had perfect needle in a haystack results but horrible real-world-performance.
- deleted 1y ago[deleted]
- soheil 1y agoMain takeaways: - Coding accuracy improved dramatically - Handles 1M-token context reliably - Much stronger instruction following
- jmkni 1y agoThe increased context length is interesting. It would be incredible to be able to feed an entire codebase into a model and say "add this feature" or "we're having a bug where X is happening, tell me why", but then you are limited by the output token length As others have pointed out too, the more tokens you use, the less accuracy you get and the more it gets confused, I've noticed this too We are a ways away yet from being able to input an entire codebase, and have it give you back an updated version of that codebase.
- impure 1y agoI like how Nano matches Gemini 2.0 Flash's price. That will help drive down prices which will be good for my app. However I don't like how Nano behaves worse than 4o Mini in some benchmarks. Maybe it will be good enough, we'll see.
- chaos_emergent 1y agoTheory here is that 4.1-nano is competing with that tier, 4.1 with flash-thinking (although likely to do significantly worse), and o4-mini or o3-large will compete with 2.5 thinking
- pzo 1y agoyeah and considering that gemini 2.0 flash is much better than 4o-mini. On top of that gemini have also audio input as modality and realtime API for both audio input and output + web search grounding + free tier.
- xnx 1y ago> That will help drive down prices which will be good for my app Why not use Gemini?
- pcwelder 1y agoCan someone explain to me why we should take Aider's polyglot benchmark seriously? All the solutions are already available on the internet on which various models are trained, albeit in various ratios. Any variance could likely be due to the mix of the data.
- meroes 1y agoTo join in the faux rigor?
- philipbjorge 1y agoIf you're looking to test an LLMs ability to solve a coding task without prior knowledge of the task at hand, I don't think their benchmark is super useful. If you care about understanding relative performance between models for solving known problems and producing correct output format, it's pretty useful. - Even for well-known problems, we see a large distribution of quality between models (5 to 75% correctness) - Additionally, we see a large distribution of model's ability to produce responses in formats they were instructed in At the end of the day, benchmarks are pretty fuzzy, but I always welcome a formalized benchmark as a means to understand model performance over vibe checking.
- asdev 1y ago> We will also begin deprecating GPT‑4.5 Preview in the API, as GPT‑4.1 offers improved or similar performance on many key capabilities at much lower cost and latency. why would they deprecate when it's the better model? too expensive?
- ComputerGuru 1y ago> why would they deprecate when it's the better model? too expensive? Too expensive, but not for them - for their customers. The only reason they’d deprecated it is if it wasn’t seeing usage worth keeping it up and that probably stems from it being insanely more expensive and slower than everything else.
- tootyskooty 1y agosits on too many GPUs, they mentioned it during the stream I'm guessing the (API) demand isn't there to saturate them fully
- simianwords 1y agoWhere did you find that 4.5 is a better model? Everything from the video told me that 4.5 was largely a mistake and 4.1 beats 4.5 at everything. There's no point keeping 4.5 at this point.
- rob 1y agoBigger numbers are supposed to mean better. 3.5, 4, 4.5. Going from 4 to 4.5 to 4.1 seems weird to most people. If it's better, it should of been GPT-4.6 or 5.0 or something else, not a downgraded number.
- HDThoreaun 1y agoOpenAI has decided to troll via crappy naming conventions as a sort of in joke. Sam Altman tweets about it pretty often
- taikahessu 1y ago> They feature a refreshed knowledge cutoff of June 2024. As opposed to Gemini 2.5 Pro having cutoff of Jan 2025. Honestly this feels underwhelming and surprising. Especially if you're coding with frameworks with breaking changes, this can hurt you.
- forbiddenvoid 1y agoIt's definitely an issue. Even the simplest use case of "create React app with Vite and Tailwind" is broken with these models right now because they're not up to date.
- asadm 1y agousually enabling "Search" fixes it sometimes as they fetch the newer methods.
- Zambyte 1y agoBy "broken" you mean it doesn't use the latest and greatest hot trend, right? Or does it literally not work?
- dbbk 1y agoPeriodically I keep trying these coding models in Copilot and I have yet to have an experience where it produced working code with a pretty straightforward TypeScript codebase. Specifically, it cannot for the life of it produce working Drizzle code. It will hallucinate methods that don't exist despite throwing bright red type errors. Does it even check for TS errors?
- dalmo3 1y agoNot sure about Copilot, but the Cursor agent runs both eslint and tsc by default and fixes the errors automatically. You can tell it to run tests too, and whatever other tools. I've had a good experience writing drizzle schemas with it.
- 1y ago
- ComputerGuru 1y agoThe benchmarks and charts they have up are frustrating because they don’t include 03-mini(-high) which they’ve been pushing as the low-latency+low-cost smart model to use for coding challenges instead of 4o and 4o-mini. Why won’t they include that in the charts?
- j_maffe 1y agoOAI are so ahead of the competition, they don't need to compare with the competition anymore /s
- neal_ 1y agohahahahaha
- forbiddenvoid 1y agoLots of improvements here (hopefully), but still no image generation updates, which is what I'm most eager for right now.
- taikahessu 1y agoOr text to speech generation ... but I guess that is coming.
- dharmab 1y agoYeah, I tried the 4o models and they severely mispronounced common words and read numbers incorrectly (eg reading 16000 as 1600)
- Tinkeringz 1y agoThey just realised a new image generation a couple of weeks ago, why are you eager for another one so soon?
- nanook 1y agoAre the image generation improvements available via API? Don't think so
- curtisszmania 1y ago[dead]
- marsh_mellow 1y agoFrom OpenAI's announcement: > Qodo tested GPT‑4.1 head-to-head against Claude Sonnet 3.7 on generating high-quality code reviews from GitHub pull requests. Across 200 real-world pull requests with the same prompts and conditions, they found that GPT‑4.1 produced the better suggestion in 55% of cases. Notably, they found that GPT‑4.1 excels at both precision (knowing when not to make suggestions) and comprehensiveness (providing thorough analysis when warranted). https://www.qodo.ai/blog/benchmarked-gpt-4-1/ https://www.qodo.ai/blog/benchmarked-gpt-4-1/
- coldcache 1y agoInteresting link. Worth noting that the pull requests were judged by o3-mini. Further, I'm not sure that 55% vs 45% is a huge difference.
- marsh_mellow 1y agoGood point. They said they validated the results by testing with other models (including Claude), as well as with manual sanity checks. 55% to 45% definitely isn't a blowout but it is meaningful — in terms of ELO it equates to about a 36 point difference. So not in a different league but definitely a clear edge
- elAhmo 1y agoI first read it as 55% better, which sounds significantly higher than ~22% which they report here. Sounds misleading.
- joshgachnang 1y agoMaybe not as much to us, but for people building these tools, 4.1 being significantly cheaper than Clause 3.7 is a huge difference.
- InkCanon 1y ago>4.1 Was better in 55% of cases Um, isn't that just a fancy way of saying it is slightly better >Score of 6.81 against 6.66 So very slightly better
- simianwords 1y agoCould any one guess the reason as to why they didn't ship this in the chat UI?
- KoolKat23 1y agoThe memory thing? More resources intensive?
- simianwords 1y agoAnswering my own question after some research. It looks like OpenAI decided not to introduce 4.1 in ChatGPT UI because 4.1 is not necessarily a better model than 4o because it is not multi modal. Now you can imagine introducing a newer "type" of model like 4.1 that's better at following instructions and better at coding to bring a sort of overhead thats already too much with the given options. OpenAI confirmed somewhere that they have already incorporated the enhancements made in 4.1 to 4o model in ChatGPT UI. I assume they would delegate to 4.1 model if the prompt doesn't require specific 4o capabilities. Also one of the improvements made to 4.1 is following instructions. This type of thing is better suited for agentic use cases that are typically used in the form of an API.
- nikcub 1y agoEasy to miss in the announcement that 4.5 is being shut down > GPT‑4.5 Preview will be turned off in three months, on July 14, 2025
- OxfordOutlander 1y agoJuice not worth the squeeze I imagine. 4.5 is chonky, and having to reserve GPU space for it must not have been worth it. Makes sense to me - I hadn't founding anything it was so much better at that it was worth the incremental cost over Sonnet 3.7 or o3-mini.
- pcwelder 1y agoDid some quick tests. I believe its the same model as Quasar. It struggles with agentic loop [1]. You'd have to force it to do tool calls. Tool use ability feels ability better than gemini-2.5-pro-exp [2] which struggles with JSON schema understanding sometimes. Llama 4 has suprising agentic capabilities, better than both of them [3] but isn't as intelligent as the others. [1] https://github.com/rusiaaman/chat.md/blob/main/samples/4.1/todo.chat.md https://github.com/rusiaaman/chat.md/blob/main/samples/4.1/t... [2] https://github.com/rusiaaman/chat.md/blob/main/samples/gemini-45-pro/todo.chat.md https://github.com/rusiaaman/chat.md/blob/main/samples/gemin... [3] https://github.com/rusiaaman/chat.md/blob/main/samples/llama4/todo.chat.md https://github.com/rusiaaman/chat.md/blob/main/samples/llama...
- ludwik 1y agoCorrect. They've mentioned the name during the live announcement - https://www.youtube.com/live/kA-P9ood-cE?si=GYosi4FtX1YSAujE&t=84 https://www.youtube.com/live/kA-P9ood-cE?si=GYosi4FtX1YSAujE...
- Yoplaid 1y ago[dead]
- simonw 1y agoHere's a summary of this Hacker News thread created by GPT-4.1 (the full sized model) when the conversation hit 164 comments: https://gist.github.com/simonw/93b2a67a54667ac46a247e7c5a2fea2c https://gist.github.com/simonw/93b2a67a54667ac46a247e7c5a2fe... I think it did very well - it's clearly good at instruction following. Total token cost: 11,758 input, 2,743 output = 4.546 cents. Same experiment run with GPT-4.1 mini: https://gist.github.com/simonw/325e6e5e63d449cc5394e92b8f2a3a36 https://gist.github.com/simonw/325e6e5e63d449cc5394e92b8f2a3... (0.8802 cents) And GPT-4.1 nano: https://gist.github.com/simonw/1d19f034edf285a788245b7b08734e17 https://gist.github.com/simonw/1d19f034edf285a788245b7b08734... (0.2018 cents)
- ilrwbwrkhv 1y agoNow try Deepseek V3 and see the magic!
- krat0sprakhar 1y agoHey Simon, I love how you generates these summaries and share them on every model release. Do you have a quick script that allows you to do that? Would love to take a look if possible :)
- jimmySixDOF 1y agoHe has a couple of nifty plugins to the LLM utility [1] so I would guess its something as simple as ```llm -t fabric:some_prompt_template -f hn:1234567890``` and that applies a template (in this case from a fabric library) and then appends a 'fragment' block from HN plugin which gets the comments, strips everything but the author and text, adds an index number (1.2.3.x), and inserts it into the prompt (+ SQLite). [1] https://llm.datasette.io/en/stable/plugins/directory.html#fragments-and-template-loaders https://llm.datasette.io/en/stable/plugins/directory.html#fr...
- simonw 1y agoI use this one: https://til.simonwillison.net/llms/claude-hacker-news-themes https://til.simonwillison.net/llms/claude-hacker-news-themes
- deleted 1y ago[deleted]
- swyx 1y agodon't miss that OAI also published a prompting guide WITH RECEIPTS for GPT 4.1 specifically for those building agents... with a new recommendation for: - telling the model to be persistent (+20%) - dont self-inject/parse toolcalls (+2%) - prompted planning (+4%) - JSON BAD - use XML or arxiv 2406.13121 (GDM format) - put instructions + user query at TOP -and- BOTTOM - bottom-only is VERY BAD - no evidence that ALL CAPS or Bribes or Tips or threats to grandma work source: https://cookbook.openai.com/examples/gpt4-1_prompting_guide#prompting-induced-planning--chain-of-thought https://cookbook.openai.com/examples/gpt4-1_prompting_guide#...
- simonw 1y agoI'm surprised and a little disappointed by the result concerning instructions at the top, because it's incompatible with prompt caching: I would much rather cache the part of the prompt that includes the long document and then swap out the user question at the end.
- swyx 1y agoyep. we address it in the podcast. presumably this is just a recent discovery and can be post-trained away.
- aoeusnth1 1y agoIf you're skimming a text to answer a specific question, you can go a lot faster than if you have to memorize the text well enough to answer an unknown question after the fact.
- zaptrem 1y agoPrompt on bottom is also easier for humans to read as I can have my actual question and the model’s answer on screen at the same time instead of scrolling through 70k tokens of context between them.
- mmoskal 1y agoThe way I understand it: if the instruction are at the top, the KV entries computed for "content" can be influenced by the instructions - the model can "focus" on what you're asking it to do and perform some computation, while it's "reading" the content. Otherwise, you're completely relaying on attention to find the information in the content, leaving it much less token space to "think".
- frognumber 1y agoMarginally on-topic: I'd love if the charts included prior models, including GPT 4 and 3.5. Not all systems upgrade every few months. A major question is when we reach step-improvements in performance warranting a re-eval, redesign of prompts, etc. There's a small bleeding edge, and a much larger number of followers.
- deleted 1y ago[deleted]
- bartkappenburg 1y agoBy leaving out scale or prior models they are effectively manipulating improvement. If from 3 to 4 it was from 10 to 80, and from 4 to 4o it was 80 to 82, leaving out 3 would let us see a steep line instead of steep decrease of growth. Lies, damn lies and statistics ;-)
- growt 1y agoMy theory: they need to move off the 4o version number before releasing o4-mini next week or so.
- kgeist 1y agoThe 'oN' schema was a such strange choice for branding. They had to skip 'o2' because it's already trademarked, and now 'o4' can easily be confused with '4o'.
- neal_ 1y agoThe better the benchmarks, the worse the model is. Subjectively for me the more advanced models dont follow instructions, and are less capable of implementing features or building stuff. I could not tell a difference in blind testing SOTA models gemini, claude, openai, deepseek. There has been no major improvements in the LLM space since the original models gained popularity. Each release claims to be much better the last, and every time i have been disappointed and think this is worse. First it was the models stopped putting in effort and felt lazy, tell it to do something and it will tell you to do it your self. Now its the opposite and the models go ham changing everything they see, instead of changing one line, SOTA models rather rewrite the whole project and still not fix the issue. Two years back I totally thought these models are amazing. I always would test out the newest models and would get hyped up about it. Every problem i had i thought if i just prompt it differently I can get it to solve this. Often times i have spent hours prompting starting new chats, adding more context. Now i realize its kinda useless and its better to just accept the models where they are, rather then try and make them a one stop shop, or try to stretch capabilities. I think this release I won’t even test it out, im not interested anymore. I’ll probably just continue using deepseek free, and gemini free. I canceled my openai subscription like 6 months ago, and canceled claude after 3.7 disappointment.
- T3uZr5Fg 1y ago[dead]
- 999900000999 1y agoHave they implemented "I don't know" yet. I probably spend 100$ a month on AI coding, and it's great at small straightforward tasks. Drop it into a larger codebase and it'll get confused. Even if the same tool built it in the first place due to context limits. Then again, the way things are rapidly improving I suspect I can wait 6 months and they'll have a model that can do what I want.
- cheschire 1y agoI wonder if documentation would help to create an carefully and intentionally tokenized overview of the system. Maximize the amount of routine larger scope information provided in minimal tokens in order to leave room for more immediate context. Similar to the function documentation provides to developers today, I suppose.
- yokto 1y agoIt does, shockingly well in my experience. Check out this blog post outlining such an approach, called Literate Development by the author: https://news.ycombinator.com/item?id=43524673 https://news.ycombinator.com/item?id=43524673
- mianos 1y agoI agree. I use it a lot but there is endless frustration when the C++ code I am working on gets both complex and largish. Once it gets to a certain size and the context gets too long they all pretty much lose the plot and start producing complete rubbish. It would be great for it to give some measure so I know to take over and not have it start injecting random bugs or deleting functional code. It even starts doing things like returning locally allocated pointers lately.
- paradite 1y agoHave you tried using a tool like 16x Prompt to send only relevant code to the model? This helps the model to focus on a subset of codebase thst is relevant to the current task. https://prompt.16x.engineer/ https://prompt.16x.engineer/ (I built it)
- vinhnx 1y ago• Flagship GPT-4.1: top‑tier intelligence, full endpoints & premium features • GPT-4.1-mini: balances performance, speed & cost • GPT-4.1-nano: prioritizes throughput & low cost with streamlined capabilities All share a 1 million‑token context window (vs 120–200k on 4o-o3/o1), excelling in instruction following, tool calls & coding. Benchmarks vs prior models: • AIME ’24: 48.1% vs 13.1% (~3.7× gain) • MMLU: 90.2% vs 85.7% (+4.5 pp) • Video‑MME: 72.0% vs 65.3% (+6.7 pp) • SWE‑bench Verified: 54.6% vs 33.2% (+21.4 pp)
- comex 1y agoSam Altman wrote in February that GPT-4.5 would be "our last non-chain-of-thought model" [1], but GPT-4.1 also does not have internal chain-of-thought [2]. It seems like OpenAI keeps changing its plans. Deprecating GPT-4.5 less than 2 months after introducing it also seems unlikely to be the original plan. Changing plans is necessarily a bad thing, but I wonder why. Did they not expect this model to turn out as well as it did? [1] https://x.com/sama/status/1889755723078443244 https://x.com/sama/status/1889755723078443244 [2] https://github.com/openai/openai-cookbook/blob/6a47d53c967a04ebb56fe9592cf5e0423e31b286/examples/gpt4-1_prompting_guide.ipynb#L57 https://github.com/openai/openai-cookbook/blob/6a47d53c967a0...
- wongarsu 1y agoMaybe that's why they named this model 4.1, despite coming out after 4.5 and supposedly outperforming it. They can pretend GPT-4.5 is the last non-chain-of-thought model by just giving all non-chain-of-thought-models version numbers below 4.5
- chrisweekly 1y agoOk, I know naming things is hard, but 4.1 comes out after 4.5? Just, wat.
- CamperBob2 1y agoFor a long time, you could fool models with questions like "Which is greater, 4.10 or 4.5?" Maybe they're still struggling with that at OpenAI.
- ben_w 1y agoAt this point, I'm just assuming most AI models — not just OpenAI's — name themselves. And that they write their own press releases.
- Cheer2171 1y agoWhy do you expect to believe a single word Sam Altman says?
- gcy 1y ago4.10 > 4.5 — @stevenheidel @sama: underrated tweet Source: https://x.com/stevenheidel/status/1911833398588719274 https://x.com/stevenheidel/status/1911833398588719274
- wongarsu 1y agoToo bad OpenAI named it 4.1 instead of 4.10. You can either claim 4.10 > 4.5 (the dots separate natural numbers) or 4.1 == 4.10 (they are decimal numbers), but you can't have both at once
- stevenheidel 1y agoso true
- furyofantares 1y agoIt's another Daft Punk day. Change a string in your program* and it's better, faster, cheaper: pick 3. *Then fix all your prompts over the next two weeks.
- wongarsu 1y agoIs the version number a retcon of 4.5? On OpenAI's models page the names appear completely reasonable [1]: The o1 and o3 reasoning models, and non-reasoning there is 3.5, 4, 4o and 4.1 (let's pretend 4o makes sense). But that is only reasonable as long as we pretend 4.5 never happened, which the models page apparently does 1: https://platform.openai.com/docs/models https://platform.openai.com/docs/models
- esafak 1y agoMore information here: https://platform.openai.com/docs/models/gpt-4.1 https://platform.openai.com/docs/models/gpt-4.1-mini https://platform.openai.com/docs/models/gpt-4.1-nano
- archeantus 1y ago“GPT‑4.1 scores 54.6% on SWE-bench Verified, improving by 21.4%abs over GPT‑4o and 26.6%abs over GPT‑4.5—making it a leading model for coding.” 4.1 is 26.6% better at coding than 4.5. Got it. Also…see the em dash
- drexlspivey 1y agoShould have named it 4.10
- clbrmbr 1y agoBut it’s so much weaker than 4.5 in broader tasks… maybe more optimized against benchmarks but it’s just no replacement for a huge model.
- pdabbadabba 1y agoWhat's wrong with the em-dash? That's just...the typographically correct dash AFAIK.
- clbrmbr 1y agoMaybe a reference to the OpenAI models loving to output em-dashes?
- LeicaLatte 1y agoi've recently set claude 3.7 as the default option for customers when they start new chats in my app. this was a recent change, and i'm feeling good about it. supporting multiple providers can be a nightmare for customer service, especially when it comes to billing and handling response quality queries. with so many choices from just one provider, it simplifies things significantly. curious about how openai manages customer service internally.
- bbstats 1y agook.
- XCSme 1y agoI tried 4.1-mini and 4.1-nano. The response are a lot faster, but for my use-case they seem to be a lot worse than 4o-mini(they fail to complete the task when 4o-mini could do it). Maybe I have to update my prompts...
- XCSme 1y agoEven after updating my prompts, 4o-mini still seems to do better than 4.1-mini or 4.1-nano for a data-processing task.
- BOOSTERHIDROGEN 1y agoMind sharing your system prompt?
- XCSme 1y agoIt's quite complex, but the task is to parse some HTML content, or to choose from a list of URLs which one is the best. I will check again the prompt, maybe 4o-mini ignores some instructions that 4.1 doesn't (instructions which might result in the LLM returning zero data).
- jjani 1y agoThat sounds incredibly disappointing given how high their benchmarks are, indicating they might be overtuned for those, similar to Llama4.
- XCSme 1y agoYeah, I think so too. They seemed to be better at specific tasks, but worse overall, at broader tasks.
- pbmango 1y agoI think an under appreciated reality is that all of the large AI labs and OpenAI in particular are fighting multiple market battles at once. This is coming across in both the number of products and the packaging. 1, to win consumer growth they have continued to benefit on hyper viral moments, lately that was was image generation in 4o, which likely was technically possible a long time before launched. 2, for enterprise workloads and large API use, they seem to have focused less lately but the pricing of 4.1 is clearly an answer to Gemini which has been winning on ultra high volume and consistency. 3, for full frontier benchmarks they pushed out 4.5 to stay SOTA and attract the best researchers. 4, on top of all they they had to, and did, quickly answer the reasoning promise and DeepSeek threat with faster and cheaper o models. They are still winning many of these battles but history highlights how hard multi front warfare is, at least for teams of humans.
- spiderfarmer 1y agoOn that note, I want to see benchmarks for which LLM's are best at translating between languages. To me, it's an entire product category.
- pbmango 1y agoThere are probably many more small battles being fought or emerging. I think voice and PDF parsing are growing battles too.
- oezi 1y agoI would love to see a stackexchange-like site where humans ask questions and we get to vote on the reply by various LLMs.
- anotherengineer 1y agois this like what you're thinking of? https://lmarena.ai https://lmarena.ai
- oezi 1y ago
- pastureofplenty 1y agoThe plagiarism machine got an update! Yay!
- sharkjacobs 1y ago> You're eligible for free daily usage on traffic shared with OpenAI through April 30, 2025. > Up to 1 million tokens per day across gpt-4.5-preview, gpt-4.1, gpt-4o and o1 > Up to 10 million tokens per day across gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, o1-mini and o3-mini > Usage beyond these limits, as well as usage for other models, will be billed at standard rates. Some limitations apply. I just found this option in https://platform.openai.com/settings/organization/data-controls/sharing https://platform.openai.com/settings/organization/data-contr... Is just this something I haven't noticed before? Or is this new?
- XCSme 1y agoSo, that's like $10/day to give all your data/prompts?
- bangaladore 1y agoIIRC 4.5 was 75$/1M input and 150$/1M output. O1 is 15$ in 60$ out. So you could easily get 75+$ per day free from this.
- sacrosaunt 1y agoNot new, launched in December 2024. https://community.openai.com/t/free-tokens-on-traffic-shared-with-openai-extended-through-april-30-2025/1129643 https://community.openai.com/t/free-tokens-on-traffic-shared...
- __mharrison__ 1y agoI know this is somewhat off topic, but can someone explain the naming convention used by OpenAI? Number vs "mini" vs "o" vs "turbo" vs "chat"?
- iteratethis 1y agoMini means the size of the model (less parameters) "o" means "omni", which means its multimodal.
- kristianp 1y agoLooks like the Quasar and Optimus stealth models on Openrouter were in fact GPT-4.1. This is what I get when I try to access the openrouter/optimus-alpha model now: {"error": {"message":"Quasar and Optimus were stealth models, and revealed on April 14th as early testing versions of GPT 4.1. Check it out: https://openrouter.ai/openai/gpt-4.1","code":404}
- lxgr 1y agoAs a ChatGPT user, I'm weirdly happy that it's not available there yet. I already have to make a conscious choice between - 4o (can search the web, use Canvas, evaluate Python server-side, generate images, but has no chain of thought) - o3-mini (web search, CoT, canvas, but no image generation) - o1 (CoT, maybe better than o3, but no canvas or web search and also no images) - Deep Research (very powerful, but I have only 10 attempts per month, so I end up using roughly zero) - 4.5 (better in creative writing, and probably warmer sound thanks to being vinyl based and using analog tube amplifiers, but slower and request limited, and I don't even know which of the other features it supports) - 4o "with scheduled tasks" (why on earth is that a model and not a tool that the other models can use!?) Why do I have to figure all of this out myself?
- fragmede 1y agowhat's hilarious to me is that I asked ChatGPT about the model names and approachs and it did a better job than they have.
- resters 1y agoI use them as follows: o1-pro: anything important involving accuracy or reasoning. Does the best at accomplishing things correctly in one go even with lots of context. deepseek R1: anything where I want high quality non-academic prose or poetry. Hands down the best model for these. Also very solid for fast and interesting analytical takes. I love bouncing ideas around with R1 and Grok-3 bc of their fast responses and reasoning. I think R1 is the most creative yet also the best at mimicking prose styles and tone. I've speculated that Grok-3 is R1 with mods and think it's reasonably likely. 4o: image generation, occasionally something else but never for code or analysis. Can't wait till it can generate accurate technical diagrams from text. o3-mini-high and grok-3: code or analysis that I don't want to wait for o1-pro to complete. claude 3.7: occasionally for code if the other models are making lots of errors. Sometimes models will anchor to outdated information in spite of being informed of newer information. gemini models: occasionally I test to see if they are competitive, so far not really, though I sense they are good at certain things. Excited to try 2.5 Deep Research more, as it seems promising. Perplexity: discontinued subscription once the search functionality in other models improved. I'm really looking forward to o3-pro. Let's hope it's available soon as there are some things I'm working on that are on hold waiting for it.
- htrp 1y agoanyone want to guess parameter sizes here for GPT‑4.1, GPT‑4.1 mini GPT‑4.1 nano I'll start with 800 bn MoE (probably 120 bn activated), 200 bn MoE (33 bn activated), and 7bn parameter for nano
- Ninjinka 1y agoI've been using it in Cursor for the past few hours and prefer it to Sonnet 3.7. It's much faster and doesn't seem to make the sort of stupid mistakes Sonnet has been making recently.
- omneity 1y agoI have been trying GPT-4.1 for a few hours by now through Cursor on a fairly complicated code base. For reference, my gold standard for a coding agent is Claude Sonnet 3.7 despite its tendency to diverge and lose focus. My take aways: - This is the first model from OpenAI that feels relatively agentic to me (o3-mini sucks at tool use, 4o just sucks). It seems to be able to piece together several tools to reach the desired goal and follows a roughly coherent plan. - There is still more work to do here. Despite OpenAI's cookbook[0] and some prompt engineering on my side, GPT-4.1 stops quickly to ask questions, getting into a quite useless "convo mode". Its tool calls fails way too often as well in my opinion. - It's also able to handle significantly less complexity than Claude, resulting in some comical failures. Where Claude would create server endpoints, frontend components and routes and connect the two, GPT-4.1 creates simplistic UI that calls a mock API despite explicit instructions. When prompted to fix it, it went haywire and couldn't handle the multiple scopes involved in that test app. - With that said, within all these parameters, it's much less unnerving than Claude and it sticks to the request, as long as the request is not too complex. My conclusion: I like it, and totally see where it shines, narrow targeted work, adding to Claude 3.7 - for creative work, and Gemini 2.5 Pro for deep complex tasks. GPT-4.1 does feel like a smaller model compared to these last two, but maybe I just need to use it for longer. 0: https://cookbook.openai.com/examples/gpt4-1_prompting_guide https://cookbook.openai.com/examples/gpt4-1_prompting_guide
- ttul 1y agoI feel the same way about these models as you conclude. Gemini 2.5 is where I paste whole projects for major refactoring efforts or building big new bits of functionality. Claude 3.7 is great for most day to day edits. And 4.1 okay for small things. I hope they release a distillation of 4.5 that uses the same training approach; that might be a pretty decent model.
- sreeptkid 1y agoI completely agree. On initial takeaway I find 3.7 sonnet to still be the superior coding model. I'm suspicious now of how they decide these benchmarks...
- feelingsonice 1y agoDoes this mean that the o1 and o3-mini models are also using 4.1 as the base now?
- osigurdson 1y agoSam made a strange statement imo in a recent Ted Talk. He said (something like) models come and go but they want to be the best platform. For me, it was jaw dropping. Perhaps he didn't mean it the way it sounded, but seemed like a major shift to me.
- mvkel 1y agoOpenAI has been a product company ever since ChatGPT launched. Their value is firmly rooted in how they wrap ux around models.
- mrieck 1y agoBefore everyone caught up: We are in a race to make a new God, and the company that wins the race will have omnipotent power beyond our comprehension. After everyone else caught up: The models come and go, some are SOTA in evals and some not. What matters is our platform and market share.
- nsoonhui 1y agoWe will also begin deprecating GPT‑4.5 Preview in the API, as GPT‑4.1 offers improved or similar performance on many key capabilities at much lower cost and latency. GPT‑4.5 Preview will be turned off in three months Here's something I just don't understand, how can ChatGPT 4.5 be worse than 4.1? Or the only thing bad is that the OpenAI naming ability?
- chr15m 1y agoThey tried something and it didn't work well. Branching paths of experimentation is not compatible with number-goes-up versioning.
- yieldcrv 1y agoMore season 4’s than attack on titan
- thund 1y agoHey OpenAI if you ever need a Version Engineer, I’m available.
- starchild3001 1y agoI feel there's some "benchmark-hacking" is going on with GPT4.1 model as its metrics on livebench.com aren't all that exciting. - It's basically GPT4o level on average. - More optimized for coding, but slightly inferior in other areas. It seems to be a better model than 4o for coding tasks, but I'm not sure if it will replace the current leaders -- Gemini 2.5 Pro, o3-mini / o1, Claude 3.7/3.5.
- sandspar 1y agoIs this correct: OpenAI will sequester 4.1 in the API permanently? And, since November 2024, they've already wrapped much of 4.1's features into ChatGPT 4o?
- clbrmbr 1y agoThe deprecation of GPT-4.5 makes me sad. It's an amazing model with great world-knowledge and subtly. It KNOWS THINGS that, on a quick experiment, 4.1 just does not. 4.5 could tell me what I would see from a random street corner in New Jersey, or how to use minor features of my niche API (well, almost), and it could write remarkably. But 4.1 doesn't hold a candle to it. Please, continue to charge me $150/1M tokens. Sometimes you need a Big Model. Tells me it was costing more than $150/1M to serve (!).
- tdehnke 1y agoI just wish they would start using human friendly names for them, and use a YY.rev version number so it's easier to know how new/old something is. Broad Knowledge 25.1 Coder: Larger Problems 25.1 Coder: Line focused 25.1
- p1dda 1y agoLLMs are not intelligent
- aitchnyu 1y agoI'm using models which scored at least 50% in Aider leaderboard but I'm micromanaging 50 line changes instead of being more vibe. Is it worth experimenting with a model that didnt crack 10%?
- lich-001 1y agoI wish they would deprecate all existing ones when they bake a new model instead of aiming for pointless model diversity.
- elAhmo 1y agoCompany worth hundreds of billions of dollars, on paper at least, has one of the worst naming schemes for their products in the recent history. Sam acknowledged this a few months ago, but with another release not really bringing any clarity, this is getting ridiculous now.
- intended 1y agoIf reasoning models are any good, then can they figure out overpowered builds for poe2? Wait, wouldn’t this be a decent test for reasoning ? Every patch changes things, and there’s massive complexity with the various interactions between items, uniques, runes, and more.
- miki123211 1y agoMost of the improvements in this model, basically everything except the longer context, image understanding and better pricing, are basically things that reinforcement learning (without human feedback) should be good at. Getting better at code is something you can verify automatically, same for diff formats and custom response formats. Instruction following is also either automatically verifiable, or can be verified via LLM as a judge. I strongly suspect that this model is a GPT-4.5 (or GPT-5???) distill, with the traditional pretrain -> SFT -> RLHF pipeline augmented with an RLVR stage, as described in Lambert et al[1], and a bunch of boring technical infrastructure improvements sprinkled on top. [1] https://arxiv.org/abs/2411.15124 https://arxiv.org/abs/2411.15124
- clbrmbr 1y agoIf so, the loss of fidelity versus 4.5 is really noticeable and a loss for numerous applications. (Finding a vegan restaurant in a random city neighborhood, for example.)
- weird-eye-issue 1y agoIn your example the LLM should not be responsible for that directly. It should be calling out to an API or search results to get accurate and up-to-date information (relatively speaking) and then use that context to generate a response
- clbrmbr 1y agoYou should actually try it. The really big models (4 and 4.5, sadly not 4o) have truly breathtaking ability to dig up hidden gems that have a really low profile on the internet. The recommendations also seem to cut through all the SEO and review manipulation and deliver quality recommendations. It really all can be in one massive model.
- muzani 1y agoThe real news for me is GPT 4.5 being deprecated and the creativity is being brought to "future models" and not 4.1. 4.5 was okay in many ways but it was absolutely a genius in production for creative writing. 4o writes like a skilled human, but 4.5 can actually write a 10 minute scene that gives me goosebumps. I think it's the context window that allows for it to actually build up scenes to hammer it down much later.
- oezi 1y agoCool to hear that you got something out of it, but for most users 4.5 might have just felt less capable on their solution-oriented questions. I guess this why they are deprecating it. It is just such a big failure of OpenAI not to include smart routing on each question and hide the complexity of choosing a model from users.
- user14159265 1y agoAnd it is available at https://t3.chat/ https://t3.chat/ (as well as claude, grok, gemini etc) for 8usd/month
- sschueller 1y ago> These Terms and your use of T3 Chat will be governed by and construed in accordance with the laws of the jurisdiction where T3 Tools Inc. is incorporated, without regard to its conflict of law provisions. Any disputes arising out of or in connection with these Terms will be resolved exclusively in the courts located in that jurisdiction, unless otherwise required by applicable law. Would be nice if there was at least some hint as to where T3 Tools Inc. is located and what jurisdiction applies.
- lsaferite 1y agoIs there an API endpoint at OpenAI that gives the information on this page as structured data? https://platform.openai.com/docs/models/gpt-4.1 https://platform.openai.com/docs/models/gpt-4.1 As far as I can tell there's no way to discover the details of a model via the API right now. Given the announced adoption of MCP and MCP's ability to perform model selection for Sampling based on a ranking for speed and intelligence, it would be great to have a model discovery endpoint that came with all the details on that page.
- sc077y 1y agoI'm wondering if one of the big reasons that OpenAI is making gpt-4.5 deprecated is not only because it's not cost-effective to host but because they don't want their parent model being used to train competitors' models (like deepseek).
- NewUser76312 1y agoAs a user I'm getting so confused as to what's the "best" for various categories. I don't have time/want to dig into benchmarks for different categories, look into the example data to see which best maps onto my current problems. The graphs presented don't even show a clear winner across all categories. The one with the biggest "number", GPT-4.5, isn't even in the best in most categories, actually it's like 3rd in a lot of them. This is quite confusing as a user. Otherwise big fan of OAI products thus far. I keep paying $20/mo, they keep improving across the board.
- nebben64 1y agoI think "best" is slightly subjective / user. But I understand your gripe. I think the only way is using them iteratively, settling on the one that best fits you / your use-case, whilst reading other peoples' experiences and getting a general vibe
- vzaliva 1y agoThey continue to baffle users with their version numbering. Intiutively 4.5 is newer/better than 4.1 and perhaps 4o but of course this is not the case.
- composableaide 1y agoExcited to see 4.1 in the API. The Nano model pricing is comparable to Gemini Flash but not where we would like it to be: https://composableai.de/openai-veroeffentlicht-4-1-nano-als-antwort-auf-gemini-2-0-flash/ https://composableai.de/openai-veroeffentlicht-4-1-nano-als-...
- Aeroi 1y agoThe user shoudn't have to research which model is the best for them. OpenAI needs to do a better job in UX and putting the best model forward in chatgpt.