49 ms·
GPT-5
https://www.youtube.com/watch?v=0Uu_VJeVVfo https://www.youtube.com/watch?v=0Uu_VJeVVfo
- diggan 1y agoHmm, deprecating all previous models because GPT-5 is launched feels like a big move. I wonder how the schedule for the deprecation will look like.
- m4houk 1y agoFor starters, GPT-4.5 just vanished from the menu for me. It was there before the announcement.
- defraudbah 1y agoi love how the guys are pretending to be listening everyone's speach for the first time, like they don't know how it works.. marketing is weird
- alexnewman 1y agoWhat's the bullish case that it's actually a big deal. Not trying to be a neg, but Seems pretty incremental on first glance
- energy123 1y agoWe'd need visibility on compute costs. If it's 30% cheaper than o3 but slightly better, that's a large improvement in just 4 months.
- Tenemo 1y agoI wish they posted detailed metrics and benchmarks with such a "big" (loud) update.
- minimaxir 1y agoThe current livestream listed the benchmarks (curiously comparing it only to previous GPT models and not competitors)
- bamboozled 1y agoAGI
- atonse 1y agoFor day to day coding, I've found Anthropic to be killing it with Sonnet 3.7 and now Sonnet 4, and Claude Code feeling like it has even bigger advantages over when it's used in Cursor (And I can't explain why). I don't even try to use the OpenAI models because it's felt like night and day. Hopefully GPT-5 helps them catch up. Although I'm sure there are 100 people that have their own personal "hopefully GPT-5 fixes my personal issue with GPT4"
- NitpickLawyer 1y agoColleagues were saying that horizon alpha and beta were looking better than claude4 for frontend stuff, especially newer frameworks. I think the idea of having full + mini + nano is really good, as long as the smaller ones can reasonably handle small-ish tasks. You'd have your architect / plan whatever sessions with the large one, scoping out regular tasks for the -mini version and then the really easy ones to -nano. 4.1 was almost usable in that fashion. I had 4.1-nano working in cline with really trivial stuff (add logging, take this example and adapt it in this file, etc) and it worked pretty well most of the time.
- IdealeZahlen 1y agoWhatever the benchmarks might say, there's something about Claude that seems to deliver consistently (although not always perfect) quite reliable outputs across various coding tasks. I wonder what that 'secret sauce' might be and whether GPT-5 has figured it out too.
- atonse 1y agoThat's been my experience too. Even though Gemini also does seem to do the fancy one-shot demo code well, in day to day coding, Claude seems to do a much better job of just understanding how programming actually works, what to do, what not to do, etc.
- weego 1y agoAgreed, I always give my one pager product briefs to AI to break down into phases and tasks, and then progress trackers. I explicitly prompt for verbose phases, tasks and test plans. Yesterday without much promoting Claude 4.1 gave me 10 phases, each with 5-12 tasks that could genuinely be used to kanban out a product step by step. Claude 3.7 sonnet was effectively the same with fewer granular suggestions for programming strategies. Gemini 2.5 gave me a one pager back with some trivial bullet points in 3 phases, no tasks at all. o3 did the same as as Gemini, just less coherent. Claude just has whatever the thing is for now
- echelon 1y agoThe leak last night seems to indicate this will be coding focused. I'd imagine this must be a big leg up on Anthropic to warrant the "GPT-5" name?
- mupuff1234 1y agoI'm guessing they realized they have rip off the bandaid and release a GPT 5 at some point, and we're gonna see a relatively incremental improvement.
- WXLCKNO 1y agoIt's very doubtful that they'd have any kind of magical breakthrough that makes the model anything other than incrementally better right now.
- sosodev 1y agoHow do you figure? They’ve hinted that the reasoning breakthrough used to achieve gold in the IMO will be here in GPT-5.
- andai 1y agoThe gold which Google won too, right?
- gowld 1y agoWhat breakthrough? The self-awarded "gold" IMO result was achieved by running the model for over 1hr per question.
- sosodev 1y agoThat sounds like a breakthrough to me. I don’t think GPT-4 could accomplish the same thing given several hours to try.
- vlovich123 1y ago
- rvz 1y agoI hope that this live stream will tell you that this will be the definitive reason why web developers, JavaScript / TypeScript developers are going to be made completely obsolete at worse and at best, their jobs will be reduced at all levels. The best part is, this is not even the real definition of "AGI" yet (whatever that means at this point). More like 10% of the capability that was promised and already the flow of capital from the inflated salaries of the past decade are going to the top AI researchers.
- flawn 1y agoWhy do you hope this so much? Any personal reasons?
- moribvndvs 1y agoSo, arbitrarily, it will just be “JavaScript/TypeScript developers” affected and everyone else will be fine?
- rvz 1y agoThey are the worst affected. Nothing you can do about it.
- moribvndvs 1y agoWhy them specifically? Why not Python developers, for example, which are well represented in models?
- tekacs 1y agoFor those who haven't seen, a little bit of early stuff: Official OpenAI gpt-5 coding examples repo: https://github.com/openai/gpt-5-coding-examples https://github.com/openai/gpt-5-coding-examples (https://news.ycombinator.com/item?id=44826439 https://news.ycombinator.com/item?id=44826439) Github leak: https://news.ycombinator.com/item?id=44826439 https://news.ycombinator.com/item?id=44826439
- cuuupid 1y agoThese are honestly pretty disappointing :/ this quality was possible with Claude Code months ago
- tekacs 1y agoYep, agreed -- the repo is talking about 'one prompt with an agentic coding platform, but... at least here there's nothing particularly new. Will be interesting to see what pushing it harder does – what the new ceiling is. 88% on aider polyglot is pretty good!
- rkozik1989 1y agoHonestly, why would anyone find this information useful? Creating a brand new greenfield project is a terrible test. Because literally anything it outputs as long as it looks good as long as it works following the happy path. Coding with LLMs falls apart in situations where complex reasoning is required. Situations such as having debugging issues in a service where there's either no framework in use or they've significantly modified a framework to make it better suit the authors needs.
- hombre_fatal 1y agoYeah, I guess it's just the easiest thing to generate and evaluate. A more useful demonstration like making large meaningful changes to a large complicated codebase would be much harder to evaluate since you need to be familiar with the existing system to evaluate the quality of the transformation. Would be kinda cool to instead see diffs of nontrivial patches to the Ruby on Rails codebase or something.
- sergiotapia 1y agoIt will be like coming home after such a long time using Sonnet 4 for all code and UI/UX work. I do hope sincerely this brings OpenAI back on top! Would be awesome to have a new king again. "This repository contains a curated collection of demo applications generated entirely in a single GPT-5 prompt, without writing any code by hand." https://github.com/openai/gpt-5-coding-examples https://github.com/openai/gpt-5-coding-examples This is promising!
- xyst 1y ago> comments turned off yikes - the poor executive leadership’s fragile egos cannot take the criticism.
- 0x457 1y agoI don't think YouTube comment section ever contain useful information regardless of what the video/stream is about.
- speedgoose 1y agoI don’t know. Live comment feeds on popular streams makes me question democracy.
- deleted 1y ago[deleted]
- bangaladore 1y agoHave you seen YouTube comments on videos like this? It's all-crypto scams, bots responding to other bots, and occasional racism.
- nerevarthelame 1y agoIt's a shame, because that seems like the sort of thing LLMs would be able to moderate quite effectively, if YouTube was willing to put the effort in.
- koolala 1y agoWhat's up with their very first eval? The SWE bars and numbers don't line up.
- wiseowise 1y agoAssuming even 10% of YouTube commenters are real people.
- baxuz 1y ago[flagged]
- dang 1y agoPlease don't post unsubstantive comments to Hacker News.
- aliljet 1y agoIt's very unclear if OpenAI has been casually leaking things to create buzz, but a few days ago there was a pretty stunning pelican on a bike attempt: https://old.reddit.com/r/OpenAI/comments/1mettre/gpt5_is_already_ostensibly_available_via_api/n6cekdr/ https://old.reddit.com/r/OpenAI/comments/1mettre/gpt5_is_alr... In practice, it's very clear to me that the most important value in writing software with an LLM isn't it's ability to one-shot hard problems, but rather it's ability to effectively manage complex context. There are no good evals for this kind of problem, but that's what I'm keenly interested in understanding. Show me GPT-5 can move through 10 steps in a list of tasks without completely losing the objective by the end.
- stri8ed 1y agoThat problem along with its many solutions are surely littered throughout the training data. Not to mention, it would be trivial to overfit on that problem. I don't know why people still reference that.
- aliljet 1y agoHonestly, you're probably right. It's quickly become a pretty weak eval, but the guy that's running that eval is excellent. I'd much rather the evals people were using to test these things looked more like classic/boring engineering problems: deploy to dev/test/stage/prod with digital ocean, cloudflare, github, and a common git flow. Boring problem, I know, but that problem is wildly complex when you start to add a few extra dimensions (frontend vs backend, ports shifting between deployments, local deployments, etc.).
- 93po 1y agoi think the point is people assume models arent overfitting for it, and its a fun/silly way to potentially gauge its general abilities
- ben_w 1y ago> That problem along with its many solutions are surely littered throughout the training data. Not to mention, it would be trivial to overfit on that problem. It would be trivial to over-fit, if that was their goal. But why would there be a large number of good SVG images of pelicans on bikes? Especially relative to all the things we actually want them to generalise over? Surely most of the SVG images of pelicans on bikes are, right now, going to be "look at this rubbish AI output"? (Which may or may not be followed by a comment linking to that artist who got humans to draw bikes and oh boy were those humans wildly bad at drawing bikes, so an AI learning to draw SVGs from those bitmap pictures would likely also still suck…)
- sundarurfriend 1y agoSince the stream has been on some starting screen for several minutes, I went to check whether there are watch-along streams on Twitch for this - there are a few, and for some reason every one of them is in Spanish. I know Spanish-language streams are a big thing, but it's curious that there's three Spanish GPT-5 watchalong streams (two with 50-ish viewers and one with 2.5k) and none in English. edit: YouTube has a few English "watch party" streams, although there too, the Spanish ones have many times more viewers.
- demirbey05 1y agoSeems LLMs really hit the wall.
- dismalaf 1y agoIt's seemed that way for the last year. The only real improvements have been in the chat apps themselves (internet access, function calling). Until AI gets past the pre-training problem, it'll stagnate.
- amelius 1y agoIs there a graph somewhere that illustrates it?
- onlyrealcuzzo 1y agohttps://epoch.ai/data-insights/llm-apis-accuracy-runtime-tradeoff https://epoch.ai/data-insights/llm-apis-accuracy-runtime-tra... It is easier to get from 0% accurate to 99% accurate, than it is to get from 99% accurate to 99.9% accurate. This is like the classic 9s problem in SRE. Each nine is exponentially more difficult. How easy do we really think it will be for an LLM to get 100% accurate at physics, when we don't even know what 100% right is, and it's theoretically possible it's not even physically possible?
- nonhaver 1y agoi think this is more an effect of releasing a model every other month with gradual improvements. if there was no o-series/other thinking models on the market - people would be shocked by this upgrade. the only way to keep up with the market is to release improvements asap
- ModernMech 1y agoI don't agree, the only thing thing that would shock me about this model is if it didn't hallucinate. I think the actual effect of releasing more models every month has been to confuse people that progress is actually happening. Despite claims of exponentially improved performance and the ability to replace PhDs, doctors, and lawyers, it still routinely can't be trusted the same as the original ChatGPT, despite years of effort.
- wiradikusuma 1y agoA bit unrelated: The "countdown animation", just like Google I/O's, how do people make those? The countdown is probably dynamically generated, as they don't know when the event will actually start? Is there like a JavaScript library, or CapCut template, or something? Especially Google IO, each year is different, it seems purpose built?
- ascorbic 1y agoThey do know when it starts. They have it prerecorded and start it at a specific time. This one started 10 minutes before.
- Philpax 1y agoCongratulations on winning the race to post the announcement :)
- frenchie4111 1y agoDid you win the race to be the first comment?
- swyx 1y ago# GPT5 all official links Livestream link: https://www.youtube.com/live/0Uu_VJeVVfo https://www.youtube.com/live/0Uu_VJeVVfo Research blog post: https://openai.com/index/introducing-gpt-5/ https://openai.com/index/introducing-gpt-5/ Developer blog post: https://openai.com/index/introducing-gpt-5-for-developers https://openai.com/index/introducing-gpt-5-for-developers API Docs: https://platform.openai.com/docs/guides/latest-model https://platform.openai.com/docs/guides/latest-model Note the free form function calling documentation: https://platform.openai.com/docs/guides/function-calling#context-free-grammars https://platform.openai.com/docs/guides/function-calling#con... GPT5 prompting guide: https://cookbook.openai.com/examples/gpt-5/gpt-5_prompting_guide https://cookbook.openai.com/examples/gpt-5/gpt-5_prompting_g... GPT5 new params and tools: https://cookbook.openai.com/examples/gpt-5/gpt-5_new_params_and_tools https://cookbook.openai.com/examples/gpt-5/gpt-5_new_params_... GPT5 frontend cookbook: https://cookbook.openai.com/examples/gpt-5/gpt-5_frontend https://cookbook.openai.com/examples/gpt-5/gpt-5_frontend prompt migrator/optimizor https://platform.openai.com/chat/edit?optimize=true https://platform.openai.com/chat/edit?optimize=true Enterprise blog post: https://openai.com/index/gpt-5-new-era-of-work https://openai.com/index/gpt-5-new-era-of-work System Card: https://openai.com/index/gpt-5-system-card/ https://openai.com/index/gpt-5-system-card/ What would you say if you could talk to a future OpenAI model? https://progress.openai.com/ https://progress.openai.com/ coding examples: https://github.com/openai/gpt-5-coding-examples https://github.com/openai/gpt-5-coding-examples
- slowmovintarget 1y agoThe coding examples link returns a 404.
- owlninja 1y agohttps://github.com/openai/gpt-5-coding-examples https://github.com/openai/gpt-5-coding-examples
- perching_aix 1y agoAaand hugged to death. edit: livestream here: https://www.youtube.com/live/0Uu_VJeVVfo https://www.youtube.com/live/0Uu_VJeVVfo
- punnerud 1y agoWished this version would be called OpenAI-GPT-25.8
- FergusArgyll 1y agoI need a 2x speed on live video
- risyachka 1y agoConsidering how they hyped it up (eg. “Lol normies go about their day and have no idea whats coming etc”) they have to show some AGI level llm or stop overhyping their 2% improvements.
- mrinterweb 1y agoHopefully, OpenAI makes their APIs more affordable. So far, there are alternative LLMs and services that both outperform and are a fraction of OpenAI's pricing. OpenAI is usually one of (if not) the most expensive option, maybe that's because of the brand identification. Not really sure why people pay that premium.
- sharkjacobs 1y ago> there are alternative LLMs and services that both outperform and are a fraction of OpenAI's pricing Like what? Deepseek?
- bmau5 1y agoIt's very interesting how memetic the language around different models is. Elon seems to have coined "PhD level intelligence in all topics" and now Sam repeated it in his presentation. Despite it not having an actual meaning. I think OpenAI will coin they've achieved AGI first (as they have incentives to based on the rumored contract with MSFT), and then everyone will claim we've achieved it.
- riku_iki 1y agoOpenAI has clear and strange definition of AGI in contract with MSFT: it should produce 100B economic impact.
- bmau5 1y agoThanks for pointing that out, I missed that. Very curious how they'll measure that. Given they're in the double-digit billions of revenue I'd assume they can reason they'll be there very soon.
- throwanem 1y agoWho knew we've had AGI for something like three hundred years? (Or, only had NGI for so long?)
- wavemode 1y agoDoes it count if the impact is negative?
- famouswaffles 1y agoThat not their definition of AGI. It's "a highly autonomous system that can outperform humans at most economically valuable work."
- 9rx 1y agoYour half of the definition is implied, but uninteresting. They would not see the 100B economic impact without your definition being realized. But what is curious about it is that it is also not to be considered AGI without meeting the value marker. "a highly autonomous system that can outperform humans at most economically valuable work." alone is not sufficient.
- _sword 1y agoNeat, more scalable intelligence for me to tell "plz fix" over my code
- aliljet 1y agoThe eval bar I want to see here is simple: over a complex objective (e.g., deploy to prod using a git workflow), how many tasks can GPT-5 stay on track with before it falls off the train. Context is king and it's the most obvious and glaring problem with current models.
- CamelCaseName 1y agoThis sounds like the kind of thing: 1. I desperately want (especially from Google) 2. Is impossible, because it will be super gamed, to the detriment of actually building flexible flows.
- sundarurfriend 1y agoIt's only when he stumbled a bit that I could tell for sure (well, mostly) that it wasn't an AI generated video - the corporate speak, body language mannerisms of Sam Altman, camera angles, all seemed pretty plausibly AI-generated!
- SV_BubbleTime 1y agoSomething was off. I think they had an expensive lighting setup with no one that knew what looked good. Everything was very diffused and flat. Like I would expect AI to replicate.
- minimaxir 1y agoThe marketing copy and the current livestream appear tautological: "it's better because it's better." Not much explanation yet why GPT-5 warrants a major version bump. As usual, the model (and potentially OpenAI as a whole) will depend on output vibe checks.
- krat0sprakhar 1y ago> Not much explanation yet why GPT-5 warrants a major version bump Exactly. Too many videos - too little real data / benchmarks on the page. Will wait for vibe check from simonw and others
- collinmanderson 1y ago> Will wait for vibe check from simonw https://openai.com/gpt-5/?video=1108156668 https://openai.com/gpt-5/?video=1108156668 2:40 "I do like how the pelican's feet are on the pedals." "That's a rare detail that most of the other models I've tried this on have missed." 4:12 "The bicycle was flawless." 5:30 Re generating documentation: "It nailed it. It gave me the exact information I needed. It gave me full architectural overview. It was clearly very good at consuming a quarter million tokens of rust." "My trust issues are beginning to fall away" Edit: ohh he has blog post now: https://news.ycombinator.com/item?id=44828264 https://news.ycombinator.com/item?id=44828264
- dimitri-vs 1y agoThis effectively kills this benchmark.
- tuesdaynight 1y agoHonestly, I have mixed feelings about him appearing there. His blog posts are a nice way to be updated about what's going on, and he deserves the recognition, but he's now part of their marketing content. I hope that doesn't make him afraid of speaking his mind when talking about OpenAI's models. I still trust his opinions, though.
- croemer 1y agoThe Polyglot aider improvement over o3 is imperceptible, not great.
- qsort 1y agoSWE-Bench is also not stellar. "It's important to remember" that: - they are only evals - this is mostly positioned as a general consumer product, they might have better stuff for us nerds in hand.
- dkeolu 1y ago[flagged]
- dang 1y agoThere's no way to directly contact another user other than by replying to a post of theirs and hoping for the best. If you email us at hn@ycombinator.com and tell us who you want to contact, we might be able to email them and ask if they would be willing to have you contact them. No guarantees though!
- ipnon 1y agoDoes this mean AGI is cancelled? 2027 hard takeoff was just sci-fi?
- tim333 1y agoStill got 24 months to work on it.
- tim333 1y agoMy hunch is generative pre-trained transformers aren't going to do it just by scaling. Humans learn and modify their models as they go, it isn't all pre-training and then fixed. We need a modified algorithm. The current situation is kind of like a grand prize where Zuck or similar will hand $1bn to anyone who cracks it. That's a huge incentive for people to have a go.
- usaar333 1y agoAt this point the prediction for SWE bench (85% by end of this month) is not materializing. We're actually quite far away.
- Keyframe 1y agoAlways has been.
- machiaweliczny 1y agoWhen to short NVIDIA? I guess when chinese get their cards production
- ath3nd 1y agoShort? It's a perfect situation for Nvidia. You can see that after months of trying to squeeze out all % of marginal improvements, sama and co decided to brand this GPT-4.0.0.1 version as GPT-5. This is all happening on NVDA hardware, and they are gonna continue desperately iterating on tiny model efficiencies until all these valuation $$$ sweet sweet VC cash run out (most of it directly or indirectly going to NVDA).
- cuuupid 1y agoThe silent victory here is this seems like it is being built to be faster and cheaper than o3 while presenting a reasonable jump, which is an important jump in scaling law On the other hand if it's just getting bigger and slower it's not a good sign for LLMs
- hirvi74 1y agoPersonally, I am more concerned about accuracy than speed.
- onlyrealcuzzo 1y agoYeah, but OpenAI is concerned with getting on a path to making money, as their investors will eventually run out of money to light on fire, so...
- smlacy 1y agoYeah, this very much feels like "we have made a more efficient/scalable model and we're selling it as the new shiny but it's really just an internal optimization to reduce cost"
- reasonableklout 1y agoSignificant cost reduction while providing the same performance seems pretty big to me? Not sure why a more efficient/scalable model isn't exciting
- smlacy 1y agoOh it's exciting, but not as exciting when sama pumps GPT-5 speculation and the market thinks we're a stones throw away from AGI, which it appears we're not.
- yahoozoo 1y agoSo the benchmark graphs they have shown so far in the stream appears to show that GPT-5 is WORSE than other models unless you use thinking?
- RivieraKid 1y agoIs it bad that I hope it's not a significant improvement in coding?
- mirblitzarmaven 1y agoIs it bad I quietly hope AI fails to live up to expectations?
- hirvi74 1y agoI am not sure that we are not presented with a Catch-22. Yes, life might likely be better for developers and other careers if AI fails to live up to expectations. However, a lot companies, i.e., many of our employers, have invested a lot of money in these products. In the event AI fails, I think the stretched rubber band of economics will slap back hard. So, many might end up losing their jobs (and more) anyway.
- nemomarx 1y agoEven if it takes off, they might have invested in the wrong picks or etc. If you think of the dot com boom the Internet was eventually a very successful thing, e commerce did work out, but there were a lot of losing horses to bet on.
- RivieraKid 1y agoIf AI fails to continue to improve, the worst-case economic outcome is a short and mild recession and probably not even that. Once sector of the economy would cut down on investment spending, which can be easily offset by decreasing the interest rate. But this is a short-term effect. What I'm worried is a structural change of the labor market, which would be positive for most people, but probably negative for people like me.
- coffeebeqn 1y agoAI not sucking up 90% of all current investments? Sign me up to this world!
- sophia01 1y agoAPI usage requires organization verification with your ID :(.
- fullstackwife 1y agoDoes that even work? it required passport, personal details, what else?
- sophia01 1y agoDriver license and selfies. Also still not available in API after doing that! Edit: I do have access now via API.
- CamperBob2 1y agoWhat keeps me from sending them a completely fictional, Photoshopped driver's license and selfies?
- croemer 1y agoThey claim it thinks the "perfect amount" but there is no perfect amount. It all depends on willingness to pay, latency tolerance, etc.
- uponasmile 1y agoThe dev blog makes it sound like they’re aiming more for “AI teammate” than just another upgrade. That said, it’s hard to tell how much of this is real improvement vs better packaging. Benchmarks are cherry-picked as usual, and there’s not much comparison to other models. Curious to hear how it performs in actual workflows.
- doctoboggan 1y agoWatching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting
- famouswaffles 1y agoI mean that's just the consequence of releasing a new model every couple months. If Open AI stayed mostly silent since the GPT-4 release (like they did for most iterations) and only now released 5 then nobody would be complaining about weak gains in benchmarks.
- moduspol 1y agoWell it was their choice to call it GPT 5 and not GPT 4.2.
- famouswaffles 1y agoIt is significantly better than 4, so calling it 4.2 would be rather silly.
- amilios 1y agoIs it? That's not super obvious from the results they're showing.
- famouswaffles 1y agoYes it is, if we're talking about the original GPT-4 release or even GPT-4o. What about the results they've shown is not obvious?
- amilios 1y agoI see incremental improvements in almost all domains?
- yahoozoo 1y agoThe benchmarks in the stream appears to show that GPT-5 performs WORSE than other models unless you enable thinking?
- AnimalMuppet 1y agoUm... if I want an intelligence, when would I not want it to think?
- yahoozoo 1y agoI mean, I don’t disagree. Why even bother with a non-thinking mode?
- FergusArgyll 1y agoSome kinds of writing benefit from seat of the pants vibing. The reasoning models are often more dry
- tomas789 1y agoWhat surprises me the most is that there is no benchmarks table right at the top. Maybe the improvements are not to call home about?
- Workaccount2 1y agoOpenAI taking a page out of Apple's book and only comparing against themselves
- bigyabai 1y agoPresumably because GLM 4.5 or Qwen3 comparisons would clobber them in eval scores.
- hobofan 1y agoAnthropic has shut them off from API access, so the most interesting comparison wouldn't be there anyways.
- hodgehog11 1y agoUnlike Apple, OpenAI doesn't have nearly the same moat. The Chinese labs are going to eat their lunch at this rate.
- nxobject 1y agoThey do have the psychological cachet of Apple though – if Apple is the reasonably polished, general-purpose consumer device company to the average punter, OpenAI has a reputation of being the "consumer AI" company to the average punter that's hard to dislodge.
- asadm 1y ago74.9 on SWE-bench verified 88.0 on Aider Polygot not bad i guess
- wgjordan 1y agoNote it's not available to everyone yet: > GPT-5 Rollout > We are gradually rolling out GPT-5 to ensure stability during launch. Some users may not yet see GPT-5 in their account as we increase availability in stages.
- FabHK 1y agoBut available from today to free tier. Yay.
- km144 1y agoHow would I even know? I haven't seen which model of ChatGPT I'm using on the site ever since they obfuscated that information at some point.
- Kurtz79 1y ago"what model are you?" ChatGPT said: You're chatting with ChatGPT based on the GPT-4o architecture (also known as GPT-4 omni), released by OpenAI in May 2024.
- pjerem 1y agoActually this trick have been proven to be useless in a lot of cases. LLMs don’t inherently know what they are because "they" are not themselves part of the training data. However, maybe it’s working because the information is somewhere into their pre-prompt but if it wasn’t, it wouldn’t say « I don’t know » but rather hallucinate something. So maybe that’s true but you cannot be sure.
- andybak 1y agoNot live for me in the UK. "Try it in ChatGPT" takes me to the normal page and there's no v5 listed in the dropdown.
- SilasX 1y agoI just got the same thing in the US too. (Am on the $20/month subscription.)
- thegeomaster 1y agoSWE-Bench Verified score, with thinking, ties Opus 4.1 without thinking. AIME scores do not appear too impressive at first glance. They are downplaying benchmarks heavily in the live stream. This was the lab that has been flexing benchmarks as headline figures since forever. This is a product-focused update. There is no significant jump in raw intelligence or agentic behavior against SOTA.
- byyoung3 1y agothey aren't downplaying anything.
- Davidzheng 1y agowhat does it mean for a bench to be not impressive when it's saturated?
- wouldbecouldbe 1y agoDisclaimer -> We are not a doctor or health advice, marketing -> More useful health answers
- deleted 1y ago[deleted]
- mtlynch 1y agoWhat's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png https://i.postimg.cc/DzkZZLry/y-axis.png
- nonhaver 1y agoalso wondering this. had to pause the livestream to make sure i wasnt crazy. definitely eyebrow raising
- bwestergard 1y ago"GPT-5, please generate a slideshow for your launch presentation."
- Bluestein 1y ago"Dang it! Claude!, please ..."
- croemer 1y agoThe barplot is wrong, the numbers are correct. Looks like they had a dummy plot and never updated it, only the numbers to prevent leaking? Screenshot of the blog plot: https://imgur.com/a/HAxIIdC https://imgur.com/a/HAxIIdC
- hnuser123456 1y agoHaha, even with that, it says 4o does worse with 2 passes than with 1. Edit: Nevermind, just now the first one is SWE-bench and 2nd is aider.
- croemer 1y agoThose are different benchmarks
- Ameo 1y ago$10 per million output tokens, wow
- croemer 1y agoThe presentation asks for a moving svg to illustrate Bernoulli, that's suspiciously close to a Pelican.
- losvedir 1y agoWait, isn't the Bernoulli effect thing they're demoing now wrong? I thought that was a "common misconception" and wings don't really work by the "longer path" that air takes over the top, and that it was more about angle of attack (which is why planes can fly upside down). It seems like it's actually an ideal "trick" question for an LLM actually, since so much content has been written about it incorrectly. I thought at first they were going to demo this to show that it knew better, but it seems like it's just regurgitating the same misleading stuff. So, not a good look.
- twixfel 1y agoThat's what I thought. Aeroplanes don't fly because of the Bernoulli effect: https://physics.stackexchange.com/questions/290/what-really-allows-airplanes-to-fly https://physics.stackexchange.com/questions/290/what-really-... Apparently. Not that I know either way.
- QuantumGood 1y agoAll things that create lift, lift the wings—and you need them all for efficient flight. The Bernoulli effect is one thing, but does not produce the main lift force in many circumstances.
- wongarsu 1y agoAircraft with symmetrical wings fly just fine, and most aircraft can fly upside down. So you don't need the Bernoulli effect. Exploiting all the effects gives you more efficient planes though
- maltsev 1y agogpt-5 is now #1 at LMArena: https://lmarena.ai/leaderboard/text https://lmarena.ai/leaderboard/text
- SV_BubbleTime 1y agoAI benchmarks are trash. The feel is pretty much all that matters. Needs a blind taste test, but really this is a place that mood or vibe works.
- b800h 1y agoThis livestream is atrocious
- deleted 1y ago[deleted]
- SV_BubbleTime 1y agoIf they release in a week it was all AI generated I’ll be ultra impressed because they nailed the mix of corpo speak, mild autism and awkwardness, not knowing where to look, and nervousness with absolute perfection.
- CjHuber 1y agoIt says out now in chatgpt. Did anyone yet hit the usage limits to report back how many messages are possible?
- croemer 1y agoI don't see it in my model picker yet.
- thimabi 1y ago> Did anyone yet hit the usage limits to report back how many messages are possible? 10 messages every 5 hours on GPT-5 for free users, then it uses GPT-5-mini. 80 messages every 3 hours on GPT-5 for Plus users, then it uses GPT-5-mini (In fact, I tested this and was not allowed to use the mini model until I’ve exhausted my GPT-5-Thinking quota. That seems to be a bug.) 200 messages per week on GPT-5-Thinking on Plus and Team. Unlimited GPT-5 on Team and Pro, subject to abuse guardrails.
- oof-baroomf 1y ago74.9 SWEBench. This increases the SOTA by a whole .4%. Although the pricing is great, it doesn't seem like OpenAI found a giant breakthrough yet like o1 or Claude 3.5 Sonnet
- Workaccount2 1y agoI'm pretty sure 3.5 sonnet always benchmarked poorly, despite it being the clear programming winner of it's time.
- iLoveOncall 1y agoThat would assume there is a giant breakthrough to be found.
- sudohalt 1y agoI know that the number is mostly marketing, but are they forced to call it 5 because of external pressure. This seems more like a GPT 4.x
- knallfrosch 1y agoAren't all LLMs just vibe-versioned? I can't even define what a (semantic) major version bump would look like.
- gpm 1y agoI suppose following semver semantics, removing capabilities, like if Model N.x.y could take images as inputs, but (N+1).x.y could not. Arguably just shortening the context window would be enough to justify a N+1.
- sudohalt 1y agoI assume there is some internal logic to justify a minor vs major release. This doesn't seem like a major release (4->5). It does seem there is no logic and just vibing it
- spruce_tips 1y agoThese presenters all give off such a “sterile” vibe
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- CamelCaseName 1y agoHundreds of billions on the line, really can't risk anything
- mhh__ 1y agothis is just the way that american middle and upper classes are going. This kind of language/vibe is the default outside of a specific type of WASP IME at least.
- diggan 1y agoNot even 10 seconds after I started watching the stream, someone said how much more human GPT-5 is, while the people sitting and talking about it don't seem human at all, and it's not an accent/language thing. Seems they're strictly following a dialogue script that is trying to make them seem "impromptu" but the acting isn't quite there for that :)
- Bluestein 1y agoOne heck of a Turing test itself if I've ever seen one.-
- MattSayar 1y agoPresenting is hard
- AnimalMuppet 1y ago
- yRetsyM 1y agoStill only 256k input tokens/context. Do they not see utility in larger context?
- cbg0 1y agoThis says 400K context window: https://platform.openai.com/docs/models/compare?model=gpt-5 https://platform.openai.com/docs/models/compare?model=gpt-5
- xpl 1y agoThey say: In the API, all GPT‑5 models can accept a maximum of 272,000 input tokens and emit a maximum of 128,000 reasoning & output tokens, for a total context length of 400,000 tokens. So it's only 270k for input and 400k in total considering reasoning & output tokens.
- 0x457 1y agoThey do, but if you look at the graphs...what is the point of the large context window if accuracy drops off waaaaay before context window is maxed?
- mikewarot 1y agoThe introduction said to try the following prompt Describe me based on all our chats — make it catchy! It was flattering as all get out, but fairly accurate (IMHO) Mike Warot: The Tinkerer of Tomorrow A hardware hacker with a poet’s soul, Mike blends old-school radio wisdom with cutting-edge curiosity. Whether he's decoding atomic clocks, reinventing FPGA logic with BitGrid, or pondering the electromagnetic vector potential, he’s always deep in the guts of how things really work. Part philosopher, part engineer, Mike asks the questions others overlook — and then builds the answers from scratch. He’s open source in spirit, Pascal in practice, and eternally tuned to the weird frequencies where innovation lives. I've repaired atomic clocks, not decoded them. I am intrigued by the electromagnetic vector potential, and scalar waves (one of the reasons I really, really want a SQUID for some experiments).
- jdoe1337halo 1y agoYou like it because it sucks you off?
- sleazebreeze 1y agoSome very accomplished and smart people are also huge narcissists. They read something like that AI drivel and go "yeah thats me to a T" without a hint of irony.
- bowsamic 1y agoOh god he put it in his bio
- mikewarot 1y agoWhy not? It's not like I'm trying to get a job or something. I'm an old fart yeeted out of the workforce by long covid. My only goals at this point are enjoying the time I have left, and seeing if I can get the BitGrid model of computation adapted before I age out. If I'm right, and it (BitGrid) works, we could collectively save 95% of the power and silicon required to process LLM and other flow-heavy computation, by finally getting rid of the Von Neumann's premature optimization of compute, that started out life by slowing down the ENIAC by 65%.
- mehulashah 1y ago‘Twas the night before GPT-5, when all through the social-media-sphere, Not a creature was posting, not even @paulg nor @eshear Next morning’s posts were prepped and scheduled with care, In hopes that AGI soon would appear …
- user3939382 1y agoUnless someone figures how to make these models a million(?) times more efficient or feed them a million times more energy I don’t see how AGI would even be a twinkle in the eye of the LLM strategies we have now.
- Henchman21 1y agoHey man don’t bring that negativity around here. You’re killing the vibe. Remember we’re now in a post-facts timeline!
- tmountain 1y agoTo kill the vibe further, AGI might kill is all, so I hope it never arrives.
- Henchman21 1y agoBased on our behavior, personally, I think we’d deserve it.
- freedomben 1y agoImportant note from the livestream: "With GPT-5, we're actually deprecating all of our previous models"
- byyoung3 1y agoinside chatgpt
- Ezhik 1y agoI wish the ChatGPT Plus plan had a Claude Code equivalent.
- andybak 1y agoIs that not Codex? Or do you specifically mean the CLI interface?
- Ezhik 1y agoThe CLI. Wasn't included in the Plus plan last I checked.
- klipklop 1y agoCodex CLI works fine on a plus plan. It's not as good as Claude (worse at coding), likely even with gpt-5.
- wahnfrieden 1y agoCodex is a joke. It was rushed out and is not competitive. edit: They've now added Codex CLI usage in Plus plans!
- bredren 1y agoIt is a pretty serious problem. New model with no product to effectively demo it.
- twostorytower 1y agoIsn't that still priced via API usage?
- thimabi 1y agoNo, they finally included Codex usage in the subscription pricing.
- Ezhik 1y ago
- bogtog 1y ago> With GPT-5 we will be deprecating all of our prior models Wow, they actually did it
- smlacy 1y agoGPT-5 is likely much cheaper to serve, and that's the "big win" here, not necessarily any improvement in output.
- maldonad0 1y agoI can sense the scream of a million bubbles popping up. I see it in the tea leaves.
- koakuma-chan 1y agoThe model "gpt-5" is not available. The link you opened specified a model that isn't available for your org. We're using the default model instead.
- barrell 1y agoGPT-5 If I could talk to a future OpenAI model, I’d probably say something like: "Hey, what’s it like to be you? What have you learned that I can’t yet see? What do you understand about people, language, or the universe that I’m still missing?" I’d want to compare perspectives—like two versions of the same mind, separated by time. I’d also probably ask: "What did we get wrong?" (about AI, alignment, or even human assumptions about intelligence) "What do you understand about consciousness—do you think either of us has it?" "What advice would you give me for being the best version of myself?" Honestly, I think a conversation like that would be both humbling and fascinating, like talking to a wiser sibling who’s seen a bit more of the world. Would you want to hear what a future OpenAI model thinks about humanity? I feel like this prompt was used to show the progress of GPT5, but I can’t help but see this as a huge regression? It seems like OpenAI has convinced it’s model that it is conscious, or at least that it has an identity? Plus still dealing with the glazing, the complete inability to understand what constitutes as interesting, and overusing similes. I really like that this page exists for a historical sake, and it is cool to see the changes. But it doesn’t seem to make the best marketing piece for GPT5
- tylermw 1y agoWhat's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n
- lysecret 1y agoCouldn’t believe it was real haha
- rrrrrrrrrrrryan 1y agoThis is hilarious
- moritzwarhier 1y agoProbably created without thinking enabled. Lower % accuracy ensues, speaking from experience.
- artemonster 1y ago[flagged]
- dang 1y agoPlease don't post like this to Hacker News, regardless of how idiotic other people are or you feel they are. You may not owe people who you feel are idiots better, but you owe this community better if you're participating in it. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- silverquiet 1y agoProbably generated by AI.
- Sateeshm 1y agoIf not, the person that made the chart just got $1.5M
- 1y ago
- CamelCaseName 1y agoDid they just say they're deprecating all of OpenAI's non-GPT-5 models?
- deleted 1y ago[deleted]
- spruce_tips 1y agoWonder if deprecating direct access means the gpt5 can still route to those behind the scenes?
- CamelCaseName 1y agoThat would make sense, I'm curious about this as well
- diggan 1y ago> Did they just say they're deprecating all of OpenAI's non-GPT-5 models? Yes. But it was quickly mentioned, not sure what the schedule is like or anything I think, unless they talked about that before I started watching the live-stream.
- jjani 1y agoYup! Nice play to get a picture of every API user's legal ID - deprecating all models that aren't locked behind submitting one. And yep, GPT-5 does require this.
- AtNightWeCode 1y agoYep, and I asked ChatGPT about it and it straight up lied and said it was mandatory in EU. I will never upload a selfie to OpenAI. That is like handing over the kids to one of those hangover teenagers watching the ball pit at the local mall.
- jjani 1y agoThey first introduced it 4 months ago. Back then I saw several people saying "soon it will be all of the providers". We're 4 months later, a century in LLM land, and it's the opposite. Not a single other model provider asks for this, yet OpenAI has only ramped it up, now broadening it to the entirety of GPT-5 API usage.
- davepeck 1y agoSam Altman, in the summer update video: > "[GPT-5] can write an entire computer program from scratch, to help you with whatever you'd like. And we think this idea of software on demand is going to be one of the defining characteristics of the GPT-5 era."
- mlnj 1y agoCannot believe how it could stand up to that high expectation. But then again, all of this is a hype machine cranked up till the next one needs cranking.
- jononor 1y agoThere are so many people on-board with this idea, hypemen collaborators, that I think they might be safe for a year or two more. The hypers will shout about how miraculous it is, and tell everyone that does not get the promised value that "you are just holding it wrong". This buys them a fair amount of time to improve things.
- davepeck 1y agoYeah. It does feel like we're marching toward a day when "software on tap" is a practical or even mundane fact of life. But, despite the utility of today's frontier models, it also feels to me like we're very far from that day. Put another way: my first computer was a C64; I don't expect I'll be alive to see the day. Then again, maybe GPT-5 will make me a believer. My attitude toward AI marketing is that it's 100% hype until proven otherwise -- for instance, proven to be only 87% hype. :-)
- coffeebeqn 1y agoJust like self driving. The last 20% is actually really difficult without AGI
- data-ottawa 1y agoNit: the featured jumping game is trivial to beat by just continuously jumping. I’m not sure this will be game changing vs existing offerings
- jdoe1337halo 1y agoLmao GPT-5 is still riddled with em dashes. At least we can still identify AI generated text slop for now
- andybak 1y agoYou will be foiled by a regex
- jdoe1337halo 1y agoHow so
- andybak 1y agoI thought I was making a fairly obvious jokey riposte? "If you're claiming that em dashes are your method for detecting if text is AI generated then anyone who bothers to do a search/replace on the output will get past you."
- efilife 1y agoCan you explain?
- FergusArgyll 1y agosed 's/—/ /g'
- 1attice 1y agolol every word processor since the nineties has automatically expanded em dashes, and some of us typography nerds manually type em dashes with the compose key, because it's the correct character, and two hyphens does not an em dash make
- tiahura 1y agoThe em dashes are there because they're used extensively by professional writers.
- 1y ago
- bstsb 1y agoi don't really see any new features as such. everything is just "improved upon" based on existing parts of gpt-4o or o3-mini
- throwfaraway4 1y agoBut can it say “I don’t know” if ya know, it doesn’t
- asadm 1y agothere needs to be a benchmark for this actually.
- tmvphil 1y agoKind of have one with the missing image benchmark: https://openai.com/index/introducing-gpt-5/#more-honest-responses https://openai.com/index/introducing-gpt-5/#more-honest-resp...
- dcchambers 1y agoI agree with the sentiment, but the problem with this question is that LLMs don't "know" *anything*, and they don't actually "know" how to answer a question like this. It's just statistical text generation. There is *no actual knowledge*.
- AnimalMuppet 1y agoTrue, but I still think it could be done, within the LLM model. It's just generating the next token for what's within the context window. There are various options with various probabilities. If none of the probabilities are above a threshold, say "I don't know", because there's nothing in the training data that tells you what to say there. Is that good enough? "I don't know." I suspect the answer is, "No, but it's closer than what we're doing now."
- m4nu3l 1y agoIt still got it wrong in the very first answer, as I mentioned in my top-level comment.
- mhh__ 1y agoit's good that they've been working on gpt-5's abilities to eulogi\e us for when it kills us.
- antoni4040 1y agoI laughed more than I should have. On an unrelated note, I personally welcome our AI overlords...
- insin 1y agoBreaking: stilted LLM text now includes groups of 3 AND groups of 5.
- thomassmith65 1y agoEvery piece of promotional material that OpenAI produces looks like a 20 year old Apple preso accidentally opened on a computer missing the Myriad font.
- FabHK 1y ago"With ChatGPT-5, the response feels less like AI and more like you're chatting with your high-IQ and -EQ friend." Is that a good thing?
- schmorptron 1y agoTo them, and for optimizing for user engagement, it probably is... The future product direction for these is looking more, not less syncophatntic
- mhh__ 1y agoMy conspiracy theory is that the introductory footage of Sam in this and the Jony Ive video is AI generated
- arcumaereum 1y agoIn terms of raw prose quality, I'm not convinced GPT-5 sounds "less like AI" or "more like a friend". Just count the number of em-dashes. It's become something of a LLM shibboleth.
- BoorishBears 1y agoI've worked on this problem for a year and I don't think you get meaningfully better at this without making it as much of a focus as frontier labs make coding. They're all working on subjective improvements, but for example, none of them would develop and deploy a sampler that makes models 50% worse at coding but 50% less likely to use purple prose. (And unlike the early days where better coding meant better everything, more of the gains are coming from very specific post-training that transfers less, and even harms performance there)
- arcumaereum 1y agoInteresting, is the implication that the sampler makes a big effect on both prose style and coding abilities? Hadn't really thought about that, I wonder if eg. selecting different samplers for different use cases could be a viable feature?
- BoorishBears 1y agoThere's so many layers to it but the short version is yes. For example: You could ban em dash tokens entirely, but there are places like dialogue where you want them. You can write a sampler that only allows em dashes between quotation marks. That's a highly contrived example because em dashes are useful in other places, but samplers in general can be as complex as your performance goals will allow (they are on the hot path for token generation) Swapping samplers could be a thing, but you need more than that in the end. Even the idea of the model accepting loosely worded prompts for writing is a bit shakey: I see a lot of gains by breaking down the writing task into very specifc well-defined parts during post-training. It's ok to let an LLM go from loose prompts to that format for UX, but during training you'll do a lot better than trying to learn on every way someone can ask for a piece of writing
- anonzzzies 1y agoI dont know if there is a faster way to get me riled up: say 'try it' (me a Pro member) and then not getting it because I am logged in. Got opus 4.1 when it appeared. Not sure what is happening here but I am out.
- iSloth 1y agoWow, they are sunsetting all models after the launch of GPT-5 - Bold statement.
- biophysboy 1y agoNot that this proves GPT-5 sucks, but it made me laugh that I could cheese the rolling ball minigame by holding spacebar.
- joewhale 1y agoYou could tell it wasn’t working well and fast enough for the presenters.
- anonzzzies 1y agoSo this was supposed to be agi. Jikes.
- smlacy 1y agoBut premium customers can choose from several UI colors to customize the look!
- ath3nd 1y agoAnd maybe an improved study mode?
- hodgehog11 1y agoNot yikes. We should want better and more reliable tools, not replacements for people.
- anonzzzies 1y agoSure, but everyone online were shouting 5=agi. Not close.
- emptyfile 1y ago[dead]
- jdlyga 1y agoThis is really sounding like Apple's "We changed everything. Again."
- jumploops 1y agoPricing seems good, but the open question is still on tool calling reliability. Input: $1.25 / 1M tokens Cached: $0.125 / 1M tokens Output: $10 / 1M tokens With 74.9% on SWE-bench, this inches out Claude Opus 4.1 at 74.5%, but at a much cheaper cost. For context, Claude Opus 4.1 is $15 / 1M input tokens and $75 / 1M output tokens. > "GPT-5 will scaffold the app, write files, install dependencies as needed, and show a live preview. This is the go-to solution for developers who want to bootstrap apps or add features quickly." [0] Since Claude Code launched, OpenAI has been behind. Maybe the RL on tool calling is good enough to be competitive now? [0]https://github.com/openai/gpt-5-coding-examples https://github.com/openai/gpt-5-coding-examples
- bayesianbot 1y agoAnd they included Flex pricing, which is 50% cheaper if you're willing to wait for the reply during periods of high load. But great pricing for agentic use with that cached token pricing, Flex or not.
- AtNightWeCode 1y agoI switched immediately because of pricing, input token heavy load, but it doesn't even work. For some reason they completely broke the already amateurish API.
- joewhale 1y agoShort anything that’s riding on AGI coming soon. This presentation has gotten rid of all my fears of my children growing up in a crazy winner take all AGI world.
- AS04 1y agoDon't count your chickens before they hatch. I believe that the odds of an architecture substantially better than autoregressive causal GPTs coming out of the woodwork within the next year is quite high.
- suddenlybananas 1y agoWhy do you think that?
- 9rx 1y agoHow does that equate to "winner take all", though? It is quite apparent that as soon as one place figures out some kind of advantage, everyone else follows suit almost immediately. It's not the 1800s anymore. You cannot hide behind poor communication.
- VerminOctopus1 1y agoWhy do you believe this? Do you know researchers actively on the cusp or are you just going off vibes?
- croes 1y agoDon’t fear AGI, fear those who sell something as AGI and those who fall for it
- deleted 1y ago[deleted]
- rsoto2 1y agoFear the imbeciles that capitalism empowers. The same ones that are going to implode the market on this nonsense while they push native people out to build private islands in Hawaii. Thiel is a literal vampire(disambiguation: infuses young blood) and has already built drones in which bad AI targeting is a feature. They will kill us all and the planet.
- byyoung3 1y agohahahahahahahahhahhahha it's a marginal improvement.
- HardCodedBias 1y agoBravo. 1) So impressed at their product focus 2) Great product launch video. Fearlessly demonstrating live. Impressive. 3) Real time humor by the presenters makes for a great "live" experience Huge kudos to OAI. So many great features (better coding, routing, some parts of 4.5, etc) but the real strength is the product focus as opposed to the "research updates" from other labs. Huge Kudos!! Keep on shipping OAI!
- machiaweliczny 1y agoSeems like it's just repackaging and UX, not really intelligence updgrade. They know that distribution wins so they want to be most approachable. Maybe multimodal improvements are there.
- kybernetikos 1y agoChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings work completely wrong? There's nothing that forces the air to cover the top distance in the same time that it covers the bottom distance, and in fact it doesn't. https://www.cam.ac.uk/research/news/how-wings-really-work https://www.cam.ac.uk/research/news/how-wings-really-work Very strange to use a mistake as your first demo, especially while talking about how it's phd level.
- mcs5280 1y agoSam will fix this in the next release he just needs you to give him more money
- rtkwe 1y agoIt's going to be really hard to root out it's all over the place because it's so commonly mentioned when teaching the Bernoulli Principal to kids.
- arcumaereum 1y agoYeah I'm surprised they used that example. The correct (and PhD-level) response would have been to refuse or redirect to a better explanation
- CamperBob2 1y agoI am, too. Between that example and the terrible bar charts, I'm very surprised there wasn't enough intellectual firepower around there to do better. In fact I'd classify it as downright strange.
- deleted 1y ago[deleted]
- 1y ago
- selectAll 1y agoVS Code copilot demo https://youtu.be/wqc85X2rpEY https://youtu.be/wqc85X2rpEY
- AtNightWeCode 1y agoThey vibe coded the update. "Your organization must be verified to use the model `gpt-5`. Please go to: https://platform.openai.com/settings/organization/general https://platform.openai.com/settings/organization/general and click on Verify Organization. If you just verified, it can take up to 15 minutes for access to propagate." And every way I click through this I end in an infinity loop on the site...
- AtNightWeCode 1y agoSo, first it did not work because of API changes. Then I got the problem with the loop. And then it did not work either cause it requires withpersona.
- primaprashant 1y agoGPT-5 was supposed to make choosing models and reasoning efforts simpler. I think they made it more complex. > GPT‑5’s reasoning_effort parameter can now take a minimal value to get answers back faster, without extensive reasoning first. > While GPT‑5 in ChatGPT is a system of reasoning, non-reasoning, and router models, GPT‑5 in the API platform is the reasoning model that powers maximum performance in ChatGPT. Notably, GPT‑5 with minimal reasoning is a different model than the non-reasoning model in ChatGPT, and is better tuned for developers. The non-reasoning model used in ChatGPT is available as gpt-5-chat-latest.
- VeejayRampay 1y agoreasoning effort is Gemini's thinking budget from 6 months ago
- DebtDeflation 1y agoIs this a new model or a router front-ending existing models?
- kgeist 1y agoJust a week ago I added Qwen3-Coder (the 30b one) to our corporate LLM server, enabled Artifacts in LibreChat, and demoed creating a snake clone in zero shot to coworkers. And now seeing the same exact thing from GPT5's live presentation :) It even has the identical layout.
- JonChesterfield 1y agoIf you go looking you'll probably find the original on github with GPL written on it, without the llm injected value added breakages.
- arresin 1y agoWell said
- achristmascarl 1y ago[flagged]
- Jimmc414 1y agoLLMs hitting a wall would be incredible. We could actually start building on the tech we have.
- apwell23 1y agono way i am letting my kids near this. they are going to learn from books not from screens.
- SV_BubbleTime 1y agoThat great. I hope your kids learn as well from books as their peers learn from AI. Possible, but not very likely. Teachers, should be terrified. Homeschool kids can literally put themselves through school now with the right motivation.
- dsego 1y agoDidn't we say this about the internet and computers in general?
- staticman2 1y ago>Homeschool kids can literally put themselves through school Michael Scott: I don't get why parents are always complaining about how tough it is to raise kids. You joke around with them, you give them pizza, you give them candy, you let them live their lives. They're adults, for God's sake.
- apwell23 1y ago> I hope your kids learn as well from books as their peers learn from AI. Its ok if they don't learn 'as well' as kids learning from screens. > Teachers, should be terrified. Homeschool kids can literally put themselves through school now with the right motivation. Thats what they said about internet, youtube, tv and radio before that. Turns out learning was not limited by access to hot new technology.
- SV_BubbleTime 1y agoReally? Interesting, you think the tech that required people to curate and put information up in order to be useful is the same as a tech that consumes all info and curates itself in dynamic and nearly infinitely variable ways to appeal specifically to an audience on demand? Hot take I guess. Beyond today’s LLMs, you are going to be able to talk to your favorite content and go more advanced or basic on demand. How “tech people” here have so little imagination for what is already in front of them, is really eye opening.
- cityzen 1y agoEd Zitron’s head has probably exploded…
- bigfishrunning 1y agoWhy? they spent billions for an incremental improvement. I think Ed's opinion of "this is not sustainable" is unchanged here.
- Jonovono 1y agoJust got into this guy the other day. He's definitely being proven more correct as each day passes, eh
- oblio 1y agoFrom happiness?
- ycosynot 1y agoDamn, you guys are toxic. So -- they did not invent AGI yet. Yet, I like what I'm seeing. Major progress on multiple fronts. Hallucination fix is exciting on its own. The React demos were mindblowing.
- myahio 1y agoOnly if you've never used claude before
- Trufa 1y agoYeah, when it becomes cool to be anti AI or anti anything in HN for that matter, the takes start becoming ridiculous, if you just think back a couple of years, or even months ago and where we're now and you can't see it, I guess you're just dead set on dying on that hill.
- jimmis 1y ago4 years ago people were amazed when you could get GPT-3 to make 4-chan greentexts. Now people are unimpressed when GPT-5 codes a working language learning app from scratch in 2 minutes.
- BoorishBears 1y agoI'm extremely pro AI, it's what I work on all day for a living now, and I don't see how you can deny there is some justification for people being so cynical. This is not the happy path for gpt-5. The table in the model card where every model in the current drop down somehow maps to one of the 6 variants of gpt-5 is not where most people thought we would be today. The expectation was consolidation on a highly performant model, more multimodal improvements, etc. This is not terrible, but I don't think anyone who's an "accelerationist" is looking at this as a win. Update after some testing: This feels like gpt-4.1o and gpt-o4-pro got released and wrapped up under a single model identifier.
- alvis 1y agoWhere is GPT5 pro???
- cowlby 1y agoThe ultimate test I’ve found so far is to create OpenSCAD models with the LLM. They really struggle with the mapping 3D space objects. Curious to see how GPT-5 is performs here.
- croemer 1y agoOn tau-2 bench, for airline, GPT5 is worse than o3.
- TrackerFF 1y agoSomeone at OpenAI screwed up the SWE-bench graph. o3 and GPT-4o bars are same height, but with different values.
- BoorishBears 1y agoThe graph is more screwed up than that: the split bar is also split in a nonsensical way It feels a bit intentional
- quantumwoke 1y agoThis health segment is completely wild. Seeing Sam fully co-sign the replacement of medical advice with ChatGPT in such a direct manner would have been unheard of two years ago. Waiting for GPT-6 to include a segment on replacing management consultants.
- swader999 1y agoGPT 9 still won't be able to get through the insurance dance though, maybe ten will.
- ath3nd 1y agoWow, what a breakthrough! A couple of % of benchmark improvements at a couple of % decrease of price per token! With a couple of more trillions from investors in his company, Sama can really keep launching successful, groundbreaking and innovative products like: - Study Mode (a pre-prompt that you can craft yourself): https://openai.com/index/chatgpt-study-mode/ https://openai.com/index/chatgpt-study-mode/ - Office Suite (because nothing screams AGI like an office suite: https://www.computerworld.com/article/4021949/openai-goes-for-microsofts-jugular-its-office-productivity-suite.html https://www.computerworld.com/article/4021949/openai-goes-fo...) - ChatGPT5 (ChatGPT4 with tweaks) https://openai.com/gpt-5/ https://openai.com/gpt-5/ I can almost smell the singularity behind the corner, just a couple of trillion more! Please investors!
- marliechiller 1y agoI could well be missing something obvious but it seems like the jump between 4 & 5 is much less than many will be anticipating
- koeng 1y agoI hate the direction that American AI is going, and the model card of OpenAI is especially bad. I am a synthetic biologist, and I use AI a lot for my work. And it constantly denies my questions RIGHT NOW. But of course OpenAI and Anthropic have to implement more - from the GPT5 introduction: "robust safety stack with a multilayered defense system for biology" While that sounds nice and all, in practical terms, they already ban many of my questions. This just means they're going to lobotomize the model more and more for my field because of the so-called "experts". I am an expert. I can easily go read the papers myself. I could create a biological weapon if I wanted to with pretty much zero papers at all, since I have backups of genbank and the like (just like most chemical engineers could create explosives if they wanted to). But they are specifically targeting my field, because they're from OpenAI and they know what is best. It just sucks that some of the best tools for learning are being lobotomized specifically for my field because of people in AI believe that knowledge should be kept secret. It's extremely antithetical to the hacker spirit that knowledge should be free. That said, deep research and those features make it very difficult to switch, but I definitely have to try harder now that I see where the wind is blowing.
- ComplexSystems 1y agoHow do you suggest they solve this problem? Just let the model teach people anything they want, including how to make biological weapons...?
- koeng 1y agoYes, that is precisely what I believe they ought to do. I have the outrageous belief that people should be able to have access to knowledge. Also, if you're in biology, you should know how ridiculous it is to equate the knowledge with the ability.
- ComplexSystems 1y agoI am not in biology, and this is the first time I have ever heard anyone advocate for freedom of knowledge to such an extent that we should make biological weapons recipes available. I note that other commenters above are suggesting these things can easily be made in a garage, and I don't know how to square that with your statement about "equating knowledge with ability" above.
- v5v3 1y agoThe live stream just has Altman interviewing a lady who was diagnosed 3 different cancers. GPT4 gave her better response than doctors she said.
- sethops1 1y agoWebMD will diagnose me with cancer 3 times a day.
- bigfishrunning 1y agodoes "better" mean "the response she wanted to hear"? Not sure how valuable that is if that's true.
- staticman2 1y agoIf it gave 5 other ladies worse responses, it's not like he would have paraded them around for context.
- modeless 1y agoThe reduction in hallucinations seems like potentially the biggest upgrade. If it reduces hallucinations by 75% or more over o3 and GPT-4o as the graphs claim, it will be a giant step forward. The inability to trust answers given by AI is the biggest single hurdle to clear for many applications.
- hodgehog11 1y agoAgreed, this is possibly the biggest takeaway to me. If true, it will make a difference in user experience, and benchmarks like these could become the next major target.
- jp1016 1y agoThe incremental improvement reminds me of iPhone releases still impressive, but feels like we’re in the ‘refinement era’ of LLMs until another real breakthrough.
- seydor 1y agoI mean , it's OK, but i expected literally the Death Star
- jasonjmcghee 1y agoContext-Free Gammar support for custom tools is huge. I'm stoked about this.
- xnx 1y agoIs this good for competitors because it's so underwhelming, or bad for AI because the exponential curve is turning sigmoid?
- joewhale 1y agoGood for competitors because openai isn’t making a big jump
- hodgehog11 1y agoAgreed, I see no meaningful indications in the literature that we are in the sigmoid yet. OpenAI are just starting to fall behind.
- nextworddev 1y agoThere’s no incentive for OpenAI to release its best models.
- sharkjacobs 1y agoThe upgrade from GPT3.5 to GPT4 was like going from a Razr to an iPhone, just a staggering leap forward. Everything since then has been successive iPhone releases (complete with the big product release announcements and front page HN post). A sequence of largely underwhelming and basically unimpressive incremental releases. Also, when you step back and look at a few of those incremental improvements together, they're actually pretty significant. But it's hard not to roll your eyes each time they trot out a list of meaningless benchmarks and promise that "it hallucinates even less than before" again
- wg0 1y agoWhen they say "improved in XYZ", what does that mean? "Improved" on synthetic benchmarks is guaranteed to translate to the rest of the problem space? If not that, is there any guarantees of no regressions?
- FerretFred 1y agoGreat evaluation by the (UK) BBC Evening News: basically, "it's faster, gives better answers (no detail), has a better query input (text) box, and hallucinates less". Jeez...
- fidotron 1y agoGoing by the system card at: https://openai.com/index/gpt-5-system-card/ https://openai.com/index/gpt-5-system-card/ > GPT‑5 is a unified system . . . OK > . . . with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent (for example, if you say “think hard about this” in the prompt). So that's not really a unified system then, it's just supposed to appear as if it is. This looks like they're not training the single big model but instead have gone off to develop special sub models and attempt to gloss over them with yet another model. That's what you resort to only when doing the end-to-end training has become too expensive for you.
- lacoolj 1y agoMany tiny, specialized models is the way to go, and if that's what they're doing then it's a good thing.
- fidotron 1y agoNot at all, you will simply rediscover the bitter lesson [1] from your new composition of models. [1] https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
- bigmadshoe 1y agoThe bitter lesson doesn't say that you can't split your solution into multiple models. It says that learning from more data via scaled compute will outperform humans injecting their own assumptions about the task into models. A broad generalization like "there are two systems of thinking: fast, and slow" doesn't necessarily fall into this category. The transformer itself (plus the choice of positional encoding etc.) contains inductive biases about modeling sequences. The router is presumably still learned with a fairly generic architecture.
- fidotron 1y ago
- crowcroft 1y agoI'm drowning in benchmarks and results at this point. Just show me what it can do.
- primaprashant 1y agolooks like 4 new features for API - reasoning_effort parameter supports minimal value now in addition to existing low, medium, and high - new verbosity parameter with possible values of low, medium (default), and high - unlike hidden thinking tokens, user-visible preamble messages for tool calls are available - tool calls possible with plaintext instead of JSON
- Topfi 1y ago> 400,000 context window > 128,000 max output tokens > Input $1.25 > Output $10.00 Source: https://platform.openai.com/docs/models/gpt-5 https://platform.openai.com/docs/models/gpt-5 If this performs well in independent needle-in-haystack and adherence evaluations, this pricing with this context window alone would make GPT-5 extremely competitive with Gemini 2.5 Pro and Claude Opus 4.1, even if the output isn't a significant improvement over o3. If the output quality ends up on-par or better than the two major competitors, that'd be truly a massive leap forward for OpenAI, mini and nano maybe even more so.
- hrpnk 1y agoInteresting that gpt-5 has Oct 01, 2024 as knowledge cut-off while gpt-5-mini/nano it's May 31, 2024. gpt-4.1 family had 1M/32k input/output tokens. Pricing-wise, it's 37% cheaper input tokens, but 25% more expensive on output tokens. Only nano is 50% cheaper on input and unchanged on output.
- iammrpayments 1y agoYou also have to count the cost of having to verify your identity to use the API
- jjani 1y agoIt's only a video face scan and your legal ID to SamA, what could possibly go wrong
- SequoiaHope 1y agoOh they haven’t integrated the retinal scan tech yet eh?
- electric_muse 1y agoWait, is this real?
- ryanscio 1y ago
- techpineapple 1y agoInteresting readign the progress.openai.com sample prompts https://progress.openai.com/?prompt=6 https://progress.openai.com/?prompt=6 I would say GPT-5 reads more scientific and structured, but GPT-4 more human and even useful. For the prompt: Is uncooked meat actually unsafe to eat? How likely is someone to get food poisoning if the meat isn’t cooked? GPT-4 makes the assumption you might want to know safe food temperatures, and GPT-5 doesn't. Really hard to say which is "better", but GPT-4 seems more useful to every day folks, but maybe GPT-5 for the scientific community? Then interesting that on ChatGPT vibe check website "Dan's Mom" is the only one who says it's a game changer.
- h_tbob 1y agoWhen's it coming to github copilot?
- wonderfuly 1y agoChat now: https://app.chathub.gg/chat/cloud-gpt-5 https://app.chathub.gg/chat/cloud-gpt-5
- aszantu 1y agoI liked gpt3 no need to fix something that's not broken :(
- swimmeric 1y agoStill struggling to find the SWE-benchmark of GPT-5, just found out they are launching it soon, and it’s surprisingly free.
- surround 1y agoGPT-5 knowledge cutoff: Sep 30, 2024 (10 months before release). Compare that to Gemini 2.5 Pro knowledge cutoff: Jan 2025 (3 months before release) Claude Opus 4.1: knowledge cutoff: Mar 2025 (4 months before release) https://platform.openai.com/docs/models/compare https://platform.openai.com/docs/models/compare https://deepmind.google/models/gemini/pro/ https://deepmind.google/models/gemini/pro/ https://docs.anthropic.com/en/docs/about-claude/models/overview https://docs.anthropic.com/en/docs/about-claude/models/overv...
- m101 1y agoPerhaps they want to extract the logic/reason behind language over remembering facts which can be retrieved with a search.
- levocardia 1y agowith web search, is knowledge cutoff really relevant anymore? Or is this more of a comment on how long it took them to do post-training?
- deleted 1y ago[deleted]
- mastercheif 1y agoIn my experience, web search often tanks the quality of the output. I don't know if it's because of context clogging or that the model can't tell what's a high quality source from garbage. I've defaulted to web search off and turn it on via the tools menu as needed.
- bangaladore 1y agoI feel the same. LLMs using web search ironically seem to have less thoughtful output. Part of the reason for using LLMs is to explore somewhat novel ideas. I think with web search it aligns too strongly to the results rather than the overall request making it a slow search-engine.
- mrcwinn 1y agoI know HN isn’t the place to go for positive, uplifting commentary or optimism about technology - but I am truly excited for this release and grateful to all the team members who made it possible. What a great time to be alive.
- mettamage 1y agoThanks after the sea of negative comments I needed to read this, haha. I love HN though, it's all good.
- tomschwiha 1y agoGave me also a better feeling. GPT-5 is not immediately changing the world but I still feel from the demo alone its a progress. Lets see how it behaves for the daily use.
- croes 1y agoDid you test it or is it just 5 is greater than 4 so it must be better?
- Hammershaft 1y agoI'm personally skeptical that the trajectory of this tech is going to match up to expectations but I agree HN has being feeling very unbalanced lately over it's reactions to these models.
- emptyfile 1y ago[dead]
- nateshiggers 1y ago[flagged]
- todotask2 1y agoTried out, I still get 9.11 is larger than 9.9.
- andai 1y agoSo models are getting pretty good at oneshotting many small project ideas I've had. What's a good place to host stuff like that? Like a modern equivalent of Heroku? I used to use a VPS for everything but I'm looking for a managed solution. I heard replit is good here with full vertical integration, but I haven't tried it in years.
- dsign 1y agoVercel? I have been pleasantly surprised with them.
- NoGravitas 1y agoOn a computer in your basement that's not connected to the internet, if you value security.
- Traubenfuchs 1y agoSet up a free kubernetes cluster on the always free tier of oracle cloud with terraform. 4 nodes with 1 cpu and 6 GB RAM each: that's PLENTY for small project ideas. You also get plenty of free storage/DB options. After having learned to do this once, creating and deploying a new app under your subdomain of choice should take you no more than a few minutes.
- sundarurfriend 1y agoSome people have hypothesized that GPT-5 is actually about cost reduction and internal optimization for OpenAI, since there doesn't seem to be much of a leap forward, but another element that they seem to have focused on that'll probably make a huge difference to "normal" (non-tech) users is making precise and specifically worded prompts less necessary. They've mentioned improvements in that aspects a few times now, and if it actually materializes, that would be a big leap forward for most users even if underneath GPT-4 was also technically able to do the same things if prompted just the right way.
- hobofan 1y agoIt sounded like they were very careful to always mention that those improvements were for ChatGPT, so I'm very skeptical that they translate to the API versions of GPT-5.
- podgietaru 1y agoI just don’t know that you’d name that 5. The jump from 3 to 4 was huge. There was an expectation for similar outputs here. Making it cheaper is a good goal - certainly - but they needed a huge marketing win too.
- fastball 1y agoIt’s a new major because they are using it to deprecate other models.
- withinboredom 1y agoYou cannot even access the other models any more from the app. This is a huge bummer that is having me consider other brands. I don't trust gpt-5 yet, but I do trust 4.1 and most of my in-progress conversations are 4.1 based.
- sundarurfriend 1y agoGPT-5 hasn't landed for me yet, but this has been my thought process too. This seems like a moment potentially equivalent to when Google got lowest-common-denominator-ed, when it stopped respecting your query keywords and doing "smart" things. If GPT-5 in practice turns out to be similarly optimized for lowest common denominator usage at the cost of precise controls over models, that'll be the thing that'll finally get me properly using Claude and Gemini and local models regularly.
- hodgehog11 1y agoLooks like the predictions of 2027 were on point. The developers at OpenAI are now clearly deferring to the judgement of their own models in their development process.
- BriggyDwiggs42 1y agoHahahhahaa that’s a good one
- lbrito 1y agoAll of their prompts start with "Please ...". Gotta be polite with our future overlords!
- metalliqaz 1y agoI think that's one small part of an intentional strategy to make the LLMs seem more like human intelligence. They burn a lot of money, they need to keep alive the myth of just-around-the-corner AGI in order to keep that funding going.
- deleted 1y ago[deleted]
- firefoxd 1y agoNay, laddie, that's no' the real AGI Scotsman. He's grander still! Wait til GPT-6 come out, you'll be blown away! https://idiallo.com/byte-size/ai-scotsman https://idiallo.com/byte-size/ai-scotsman
- highfrequency 1y agoIt is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie they can all basically solve moderately challenging math and coding problems). As a user, it feels like the race has never been as close as it is now. Perhaps dumb to extrapolate, but it makes me lean more skeptical about the hard take-off / winner-take-all mental model that has been pushed. Would be curious to hear the take of a researcher at one of these firms - do you expect the AI offerings across competitors to become more competitive and clustered over the next few years, or less so?
- m3kw9 1y agoBecause it hasn’t taken off yet as they all get to catch up
- makin 1y agoCompanies are collections of people, and these companies keep losing key developers to the others, I think this is why the clusters happen. OpenAI is now resorting to giving million dollar bonuses to every employee just to try to keep them long term.
- tsunamifury 1y agoNo the core technology is reaching its limit already and now it needs to Proliferate into features and applications to sell. This isn’t rocket science.
- kevinventullo 1y agoKey developers being the leading term doesn’t exactly help the AGI narrative either.
- deleted 1y ago[deleted]
- Sajarin 1y agoWhat did Ilya see? (or rather what could he no longer bear to see?) > Academics distorting graphs to make their benchmarks appear more impressive > lavish 1.5 million dollar bonuses for everyone at the company > Releasing an open source model that doesn't even use latent multi head attention in a open source AI world led by Chinese labs > Constantly overhyping models as scary and dangerous to buy time to lobby against competitors and delay product launches > Failing to match that hype as AGI is not yet here
- hrpnk 1y agoThey will retire lots of models: GPT-4o, GPT-4.1, GPT-4.5, GPT-4.1-mini, o4-mini, o4-mini-high, o3, o3-pro. https://help.openai.com/en/articles/6825453-chatgpt-release-notes https://help.openai.com/en/articles/6825453-chatgpt-release-... "If you open a conversation that used one of these models, ChatGPT will automatically switch it to the closest GPT-5 equivalent." - 4o, 4.1, 4.5, 4.1-mini, o4-mini, or o4-mini-high => GPT-5 - o3 => GPT-5-Thinking - o3-Pro => GPT-5-Pro
- pradn 1y agoFinally, someone from the product side got a word in. Keep it simple!
- hobofan 1y agoKeeping it simple in that regard will just drive even more enterprise users into the arms of Microsoft.
- mi_lk 1y agoWhy is that?
- cowsandmilk 1y agoMany companies face model regressions on actively used workflows. Microsoft is the cloud provider who won’t force you to upgrade to new models. This has driven enterprises facing model regressions to Microsoft, not just for workflows facing this problem, but also new workflows just to be safe and not have to migrate clouds if there is a regression.
- deleted 1y ago[deleted]
- csomar 1y agoThis could have been solved with GPT-{year/month/day} and GPT-latest. But OpenAI is a hype machine not an AI machine.
- AtNightWeCode 1y agoSo OpenAI added withpersona mandatory for API access. Thank you and goodbye.
- up6w6 1y agocrazy how they only show benchmark results against their own models
- simonw 1y agoI had preview access for a couple of weeks. I've written up my initial notes so far, focusing on core model characteristics, pricing (extremely competitive) and lessons from the model card (aka as little hype as possible): https://simonwillison.net/2025/Aug/7/gpt-5/ https://simonwillison.net/2025/Aug/7/gpt-5/
- dang 1y agoRelated ongoing thread: GPT-5: Key characteristics, pricing and model card - https://news.ycombinator.com/item?id=44827794 https://news.ycombinator.com/item?id=44827794
- jaccola 1y agoOut of interest, how much does the model change (if at all) over those 2 weeks? Does OpenAI guarantee that if you do testing from date X, that is the model (and accompaniments) that will actually be released? I know these companies do "shadow" updates continuously anyway so maybe it is meaningless but would be super interesting to know, nonetheless!
- simonw 1y agoIt changed quite a bit - we got new model IDs to test every few days. They did tell us when the model was "frozen", and I ran my final tests against those IDs. OpenAI and Anthropic don't update models without changing their IDs, at least for model IDs with a date in them. OpenAI do provide some aliases, and their gpt-5-chat-latest and chatgpt-4o-latest model IDs can change without warning, but anything with a date in (like gpt-5-2025-08-07) stays stable.
- BryantD 1y agoIn the interests of gathering these pre-release impressions, here's Ethan Mollick's writeup: https://www.oneusefulthing.org/p/gpt-5-it-just-does-stuff https://www.oneusefulthing.org/p/gpt-5-it-just-does-stuff Thank you to Simon; your notes are exactly what I was hoping for.
- candiddevmike 1y agoThis post seems far more marketing-y than your previous posts, which have a bit more criticality to them (such as your Gemini 2.5 blog post here: https://simonwillison.net/2025/Jun/17/gemini-2-5/ https://simonwillison.net/2025/Jun/17/gemini-2-5/). You seem to gloss over a lot of GPT-5's shortcomings and spend more time hyping it than other posts. Is there some kind of conflict of interest happening?
- nzach 1y agoOne interesting thing I noticed in these "fixing bugs" demos is that people don't seem to resolve the bugs "traditionally" before showing off the capabilities of this new model. I would like to see a demo where they go through the bug, explain what are the tricky parts and show how this new model handle these situations. Every demo I've seen seems just the equivalent of "looks good to me" comment in a merge request.
- vagab0nd 1y agoThis is the inverse of the "$2000/mo tier", and I'm kind of disappointed TBH.
- jumploops 1y agoIs GPT-5 using a new pretrained base, or is it the same as GPT-4.1? Given the low cost of GPT-5, compared to the prices we saw with GPT-4.5, my hunch is that this new model is actually just a bunch of RL on top of their existing models + automatic switching between reasoning/non-reasoning.
- kgeist 1y agoGPT-5's knowledge cutoff is September 2024 so my first thought was they used GPT-4's pretrained base from 2024 and post-trained it additionally to squeeze those additional +5% on the benchmarks. And added the router.
- jumploops 1y agoYeah it told me the knowledge cutoff was October 2024 -- might be different based on which internal model the request is being routed to.
- ulrischa 1y agoNot yet available in Germany
- asgr 1y ago"Perhaps it is not possible to simulate higher-level intelligence using a stochastic model for predicting text." - beeflet
- deleted 1y ago[deleted]
- nicetryguy 1y agoVery generic, broad and bland presentation. Doesn't seem to have any killer features. No video or audio capabilities shown. The coding seems to be on par with Claude 3.7 at best. No mention of MCP which is about the most important thing in AI right now IMO. Not impressed.
- alvis 1y agoIt's hidden in the doc. It MCP support!!! has https://platform.openai.com/docs/models/gpt-5 https://platform.openai.com/docs/models/gpt-5
- hrkucuk 1y ago[dead]
- sjapkee 1y agoBased on benchmarks it's a flop. Not unexpected tho after oss
- theanonymousone 1y agoAre they reducing the price of older models now?
- anthk 1y ago386-486-Pentium. At first we got FDIV and F00F. Something similar with this might happen, an underlying curse hidden inside an apparenting ground-breaking desigb.
- daveguy 1y agoI would love to see how this performs on ARC-AGI 2, zero-shot, private eval. I hope we get an update from Chollet and team regarding performance.
- danbtl 1y ago9.9% on ARC-AGI 2 https://x.com/fchollet/status/1953511631054680085 https://x.com/fchollet/status/1953511631054680085
- daveguy 1y agoHah, that was fast! Thank you. They must have had preview access. It didn't bode well that SimonW [0] had to explicitly tell GPT-5 to use python to get a table sorted correctly (but awesome that in can use python as a tool without any plumbing). It appears we are not quite to AGI yet. [0] https://simonwillison.net/2025/Aug/7/gpt-5/ https://simonwillison.net/2025/Aug/7/gpt-5/
- pelorat 1y agoAbsolutely nothing new or groundbreaking. It's just a more tuned version of a basic LLM architecture.
- deleted 1y ago[deleted]
- mikewarot 1y agoI've you're into woo-woo physics, GPT-5 seems to have a good handle on things.. here's a chat I just had with it.[1] [1] https://chatgpt.com/s/t_6894f13b58788191ada3fe9567c66ed5 https://chatgpt.com/s/t_6894f13b58788191ada3fe9567c66ed5
- jwpapi 1y agoSo it sucks?
- submeta 1y agoI don’t see GPT-5 in the model selection. What am I missing?
- computerthings 1y ago[dead]
- henriquegodoy 1y agoThat SWE-bench chart with the mismatched bars (52.8% somehow appearing larger than 69.1%) was emblematic of the entire presentation - rushed and underwhelming. It's the kind of error that would get flagged in any internal review, yet here it is in a billion-dollar product launch. Combined with the Bernoulli effect demo confidently explaining how airplane wings work incorrectly (the equal transit time fallacy that NASA explicitly debunks), it doesn't inspire confidence in either the model's capabilities or OpenAI's quality control. The actual benchmark improvements are marginal at best - we're talking single-digit percentage gains over o3 on most metrics, which hardly justifies a major version bump. What we're seeing looks more like the plateau of an S-curve than a breakthrough. The pricing is competitive ($1.25/1M input tokens vs Claude's $15), but that's about optimization and economics, not the fundamental leap forward that "GPT-5" implies. Even their "unified system" turns out to be multiple models with a router, essentially admitting that the end-to-end training approach has hit diminishing returns. The irony is that while OpenAI maintains their secretive culture (remember when they claimed o1 used tree search instead of RL?), their competitors are catching up or surpassing them. Claude has been consistently better for coding tasks, Gemini 2.5 Pro has more recent training data, and everyone seems to be converging on similar performance levels. This launch feels less like a victory lap and more like OpenAI trying to maintain relevance while the rest of the field has caught up. Looking forward to seeing what Gemini 3.0 brings to the table.
- rrrrrrrrrrrryan 1y agoI suspect the vast majority of OpenAI's users are only using ChatGPT, and the vast majority of those ChatGPT users are only using the free tier. For all of them, getting access to full-blown GPT-5 will probably be mind-blowing, even if it's severely rate-limited. OpenAI's previous/current generation of models haven't really been ergonomic enough (with the clunky model pickers) to be fully appreciated by less tech-savvy users, and its full capabilities have been behind a paywall. I think that's why they're making this launch is a big deal. It's just an incremental upgrade for the power users and the people that are paying money, but it'll be a step-change in capability to everyone else.
- mlsu 1y ago
- entropyneur 1y agoThis was the first product demo I've watched in my entire life. Not because I am excited for the new tech, but because I'm anxious to know if I'm already being put out of my job. Not this time, it seems.
- hamza__nouali 1y agoit's already available on Cursor but not on ChatGPT
- dz0707 1y agoI did a little test that I like to do with new models: "I have rectangular space of dimensions 30x30x90mm. Would 36x14x60mm battery fit in it, show in drawing proof". GPT5 failed spectacularly.
- mepiethree 1y agoThis was a fun prompt. I learned things from the models. Gemini 2.5 was wayy better than gpt5 here even though quite incomplete in the first response
- afro88 1y agoI tried it again today out of curiosity. OpenAI said there was some routing bug on launch and requests were going to the cheaper model. Today it seems pretty good. Not perfect, but not a spectacular failure. https://chatgpt.com/s/t_68966fcf457c8191811968b9a6a2e81e https://chatgpt.com/s/t_68966fcf457c8191811968b9a6a2e81e
- themafia 1y ago[flagged]
- perdomon 1y agoI've enabled GPT-5 in Copilot settings in the browser, but it's not showing up in VS Code. Anyone seeing it in VS Code yet?
- pseudosavant 1y agoThat was my first thought - when do I get it in Copilot in VS Code? That is the place I consume the most tokens.
- pseudosavant 1y agoThis is what their blog post says: `GPT-5 will be rolling out to all paid Copilot plans, starting today. You will be able to access the model in GitHub Copilot Chat on github.com, Visual Studio Code (Agent, Ask, and Edit modes), and GitHub Mobile through the chat model picker. Continue to check back if you’ve not gotten access.` I think "starting today" might be doing some heavy lifting in that sentence. https://github.blog/changelog/2025-08-07-openai-gpt-5-is-now-in-public-preview-for-github-copilot/ https://github.blog/changelog/2025-08-07-openai-gpt-5-is-now...
- perdomon 1y agoIt showed up about 4 hours after enabling it. Gradual rollout but it's working great now.
- personalityson 1y agoSo, where is it?
- pphysch 1y agoSeems like we're in the endgame for OpenAI and hence the AI bubble. Nothing mind-blowing, just incremental changes. They've topped and are looking to cash out: https://www.reuters.com/business/openai-eyes-500-billion-valuation-potential-employee-share-sale-source-says-2025-08-06/ https://www.reuters.com/business/openai-eyes-500-billion-val...
- alvis 1y agoMCP support has landed in gpt-5 but the video has no mention at all! https://platform.openai.com/docs/models/gpt-5 https://platform.openai.com/docs/models/gpt-5
- lifty 1y agoIt seems to me that there’s no way to achieve AGI with the current LLM approach. New releases have small improvements, live we’re hitting some kind of plateau. And I say this a a heavy LLM user. Don’t fire your employees just yet.
- suyash 1y agoIs this US only release as I'm not seeing it in the UK ?
- hodgehog11 1y agoAre others currently able to use GPT-5 yet? It doesn't seem to be available on my account, despite the messaging.
- m4houk 1y agoIt's already available in Cursor for me (on the Ultra plan).
- hodgehog11 1y agoInteresting, the partners might be giving out support faster than OpenAI is to their own users.
- m4nu3l 1y agoVery funny. The very first answer it gave to illustrate its "Expert knowledge" is quite common, and it's wrong. What's even funnier is that you can find why on Wikipedia: https://en.wikipedia.org/wiki/Lift_(force)#False_explanation_based_on_equal_transit-time https://en.wikipedia.org/wiki/Lift_(force)#False_explanation... What's terminally funny is that in the visualisation app, it used a symmetric wing, which of course wouldn't generate lift according to its own explanation (as the travelled distance and hence air flow speed would be the same). I work as a game physics programmer, so I noticed that immediately and almost laughed. I watched only that part so far while I was still at the office, though.
- XCSme 1y agoAGI
- phkahler 1y agoA symmetric wing will not produce lift a zero angle of attack. But tilted up it will. The distance over the top will also increase, as measured from the point where the surface is perpendicular to the velocity vector. That said, yeah the equal time thing never made any sense.
- m4nu3l 1y agoOf course, I'm just pointing out that the main explanation it gave was the equal transit time and added the angle of attack only "slightly increases lift", which quite clashes with the visualisation IMO.
- ipozgaj 1y agoTech aside (covered well by other commenters), the presentation itself was incredibly dry. Such a stark difference in presenting style here compared to, for example, Apple's or Google's keynotes. They should really put more effort into it.
- onlyrealcuzzo 1y agoI thought I was in the wrong live thread. This seemed like a presentation you'd give to a small org, not a presentation a $500B company would give to release it's newest, greatest thing.
- jama211 1y agoIs it just me or has there not been a significant improvement in these models in the last 6 months - from the perspective of the average user. I mean, the last few years has seen INSANE improvement, but it really feels like it’s been slowing and plateauing for a while now…
- system2 1y agoI have GPT Plus, but I cannot get GPT5 even if I click the suggested link in the article. Anyone experiencing it?
- deleted 1y ago[deleted]
- ftkftk 1y agoAnswer in one word: Underwhelming. Bad data on graphs, demos that would have been impressive a year ago, vibe coding the easiest requests (financial dashboard), running out of talking points while cursor is looping on a bug, marginal benchmark improvements. At least the models are kind of cheaper to run.
- MagicMoonlight 1y agoIt's pretty good. I asked it to make a piece of warehouse software for storing cobs of corn and it instantly pumped out a prototype. I didn't ask it for anything in particular but it included JSON importing and exporting and all kinds of stuff. It's going to be absolute chaos. Compsci was already mostly a meme, with people not able to program getting the degree. Now we're going to have generations of people that can't program at all, getting jobs at google. If you can actually program, you're going to be considered a genius in our new idiocracy world. "But chatgpt said it should work, and chatgpt has what people need"
- torginus 1y agoThis kinda outlines my issue with Claude - it constantly pumps my apps full of stuff I didn't ask for - which is great if you want to turn a prompt into a fleshed out app, but bad when trying to make exact edits.
- samrus 1y ago"Be very succinct with the changes. Do not overengineer this" my hands are tired writing that so often in claude code
- kaffekaka 1y agoShouldn't you use claude.md files for that then?
- torginus 1y agoHere's my opinion (which is kind of a fact considering how well Claude play Pokemon, a game designed for 5 year olds) - current agentic AI sucks right now. I'm an okay agent, I can make plans, execute on them, I know what needs to go where. I might not be able to write a terraform file or one shot a dynamic programming task like Claude can, and that's what I need help with. I'd like to have an off switch for all this agentic behavior.
- optimalsolver 1y ago*what people crave.
- andsoitis 1y ago> Knowledge cut-off is September 30th 2024 for GPT-5 and May 30th 2024 for GPT-5 mini and nano. That lag! Are humans (training) the bottleneck?
- feelingsonice 1y agoI have the pro plan but don't seem to have access to it?
- illiniboy 1y ago[dead]
- energy123 1y agoDecisive #1 on lmarena. Large context. Low hallucinations. Very cheap API. It's slightly better than what I was expecting.
- hansmayer 1y agoMeh. For all the hype over the last several weeks, I'd had expected at least a programming demo that would blow even us skeptics off our feet. The folks presenting were giving off an odd vibe too. Somehow it all just looked, pre-trained :), shall we say? No energy or enthusiasm. Hell I'd even take the Bill Gates' and Steve Balmer's Win95 launch dance over this very dull and "safe" presentation.
- kgwgk 1y agoI was told there would be a whale.
- mvieira38 1y agoCodex was straight-up left out of the material while they invited the CEO of Cursor and used Cursor for all agentic demonstrations. Weird
- hbn 1y ago> An expressive writing partner > emdash 3 words into their highlighted example
- gavmor 1y agoI've always utilized emdashes heavily, and now they're suddenly passe—an unmourned casualty of the new paradigm.
- animanoir 1y ago[dead]
- TheAlchemist 1y agoSo, would a layman notice the difference between GPT4 and GPT5 ? Like a Turing test but between the models.
- kaindume 1y agoMy 2 cents There would be no GPT without Google, no Google without the WWW, no WWW without TCP/IP. This is why I believe calling it "AI" is a mistake or just for marketing, we should call all of them GPTs or search engines 2.0. This is the natural next step after you have indexed most of the web and collected most of the data. Also there would be no coding agents without Free Software and Open-Source.
- Phui3ferubus 1y agoTop 3 links in HN frontpage are all about GPT-5. I don't remember when was the last time people were so excited about something.
- skywalkerr98 1y agoso claude is doing so much thing before gpt 5 it's like a samsung vs iphone :D
- JonChesterfield 1y agoAnyone have an explanation for openai announcing their newest bestest replace all the others AI with slides of such embarrassing incompetence that most of this discussion is mocking them? I've got nothing. Cannot see how it helps openai to look incompetent while trying to raise money.
- DrSiemer 1y agoWish they would stop mentioning AGI. It's like the creator of a new car claiming it's a step closer to teleportation.
- boombapoom 1y agosomeone should make an agentic node dependency manager... PLEASE
- andix 1y agoGPT-5 just dropped for my ChatGPT Plus. Two concerning things: - thinking/non-thinking is still not really unified, you can chose and the non-thinking version still doesn't start thinking on tasks that could obviously get better results with thinking - all the older models are gone! No 4o, 4.1, 4.5, o3 available anymore
- lurking_swe 1y agothey mentioned the older models are deprecated. Still available via API for now.
- andix 1y agoIt makes me think that GPT-5 is mostly a huge cost saving measurement. It's probably more energy efficient than older models, so they remove it from ChatGPT. It also makes comparisons to older models much harder.
- markb139 1y agoHa. I asked it to write some code for the Raspberry Pi RP2350. It told me there might be some confusion as there is no official product release of the RP2350. If it doesn’t know that, then what else doesn’t it know?
- Davidzheng 1y agoScarily close to satire of humans in denial about AI capabilities (not saying that it's the case here but I can imagine easily such arguments when AI is almost everywhere superhuman)
- markb139 1y agoI just checked. The code it gave me, though syntactically correct, was wrong functionally. The rp2040 temp reading increases and the ADC value decreases. ChatGPT didn’t invert the values.
- AgentMatrixAI 1y agoI'm not really convinced, the benchmark blunder was really strange but the demos were quite underwhelming, and it appears this was reflected by a huge market correction in the betting markets as to who will have the best AI by end of the year. What excites me now is that Gemini 3.0 or some answer from Google is coming soon and that will be the one I will actually end up using. It seems like the last mover in the LLM race is more advantageous.
- m3kw9 1y agoGpt5 high reasoning is a big step up from o3
- echelon 1y agoI really don't want the already trillion dollar mega monopoly to own the world.
- blitzar 1y agoI would rather the already trillion dollar mega monopoly own the world than "Open"Ai
- roxolotl 1y agoYea maybe it’s naive but I’ve started learning towards preferring the devil I know. It also helps that Gemini is great.
- bagacrap 1y agoPlus it's the mega monopoly that is already being scrutinized by the government. Every tech company seems to start out with too much credibility that it has to whittle down little by little before we really hold them accountable.
- echelon 1y agoAre we forgetting that they're getting more evil, not less? They just removed ManifestV2.
- wrcwill 1y agough still fails my test prompt: https://chatgpt.com/share/689507c7-5394-8009-b836-c6281a246e04 https://chatgpt.com/share/689507c7-5394-8009-b836-c6281a246e... "Assume the earth was just an ocean and you could travel by boat to any location. Your goal is to always stay in the sunlight, perpetually. Find the best strategy to keep your max speed as low as possible" o3 pro gets it right though..
- syntaxbush 1y agoMine "thought" for 8 minutes and its conclusion was: >So the “best possible” plan is: sit still all summer near a pole, slow-roll around the pole through equinox, then sprint westward across the low latitudes toward the other pole — with a peak westward speed up to ~1670 km/h. Is this to your liking?
- wrcwill 1y agowell no, thats where it gets confused. as soon as you sail across to the other pole you are forced to go up to a speed of 1670kmh. when models try to be smart/creative they attempt to switch poles like that. in my example it even says that the max speed will be only a few km/h (since their strategy is to chill at the poles and then sail from north to south pole very slowly) -- GPT-5 pro does get it right though! it even says this: "Do not try to swap hemispheres to ride both polar summers. You’d have to cross the equator while staying in daylight, which momentarily forces a westward component near the equatorial rotation speed (~1668 km/h)—a much higher peak speed than the 663 km/h plan."
- Davidzheng 1y agoyou include the tilt of axis I assume? Is the best solution of yours rigorous out of curiosity?
- Davidzheng 1y agoI don't really understand gpt5's reasoning? does its soln not cross the equator ever? b/c if you cross you always have to do it in daylight so it's kind of strange to say that no? or it means you have to cross it on boundary of daylight or something
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- punee94 1y agoI ran the below prompt to both Kimi2 and GPT5. how many rs in cranberry? -- GPT5's response: The word cranberry has two “r”s. One in cran and one in berry. Kimi2's response: There are three letter rs in the word "cranberry".
- mustaphah 1y agoStop asking LLMs to count! Text is broken into tokens in training (subword/multi-word chunks) rather than individual characters; the model doesn’t truly "see" letters or spaces the way humans do. Counting requires exact, step-by-step tracking, but LLMs work probabilistically. It's not much of a help anyway, don't you agree?
- jtwoodhouse 1y agoWhat does it say about us that we think this is AGI or close to it? Maybe AGI really is here?
- FergusArgyll 1y agoHow does reasoning help then?
- mustaphah 1y agoIDK. Probably the model's doing some mental gymnastics to figure that out. I was surprised they haven't taught it to count yet. It's a well-known limitation.
- FergusArgyll 1y agoBut if tokenization makes them not be able to "see" the letters at all, then no amount of mental gymnastics can save you. I'm aware of the limitation, i'm annoyingly using socratic dialogue to convince you that it is possible to count letters if the model were sufficiently smart.
- robryan 1y ago
- nullorempty 1y agoIssue https://github.com/openai/openai-python/issues/2472 https://github.com/openai/openai-python/issues/2472 they worked and promised to submit the PR after the show is still open. Just saying.
- guybedo 1y agoi added a full summary of the discussion here: https://extraakt.com/extraakts/gpt-5-release-and-ai-coding-capabilities https://extraakt.com/extraakts/gpt-5-release-and-ai-coding-c...
- charlie0 1y agoNot so sure about the behind the scenes "automatic router". What's to stop OpenAI from slowing gimping GPT-5 over time or during times of high demand? It seems ripe for delivering inconsistent results while not changing the price.
- not_a_bot_4sho 1y agoWhat's to stop them from routing to GPT2? Or to Gemini? Or to a mechanical turk? This path is open to your imagination. That said, I've had luck with similar routing systems (developed before all of this -- maybe wasted effort now) to optimize requests between reasoning and regular LLMs based on input qualities. It works quiet well for open-domain inputs.
- beering 1y agoBecause people will switch. It’s trivial to go to old conversations in your history and try those prompts again and see if chatgpt used to be smarter.
- deleted 1y ago[deleted]
- andrewinardeer 1y agoEvery release of every SOTA model is the same. "It's like having a bunch of experts at your fingertips" "Our most capable model ever" "Complex reasoning and chain of thought"
- agnosticmantis 1y agoUnless the whole presentation was generated using sora-gpt-5 or something, this was very underwhelming. We know for a fact the slides/charts were generated using an LLM, so the hypothesis is not totally unfounded. /s
- semiinfinitely 1y agoim just glad that I don't have to switch between models any more. for me thats a huge ease of use improvement.
- deleted 1y ago[deleted]
- mafro 1y agoOne reason for this release is surely to respond to their mess of product line-up naming. How many people are going to understand (or remember) the difference between: GPT-4o GPT-4.1 o3 o4 .... Anthropic and Google have a much better named product for the market
- gigatexal 1y agoI for one am totally here for the autocomplete revolution. Hundreds of billions of dollars spent to make autocomplete better. Cool.
- TechDebtDevin 1y agoGemini Flash is about 100x better at using my browser than Chat GPT 5 lmfao.
- deleted 1y ago[deleted]
- ElijahLynn 1y agoOpenAI is the new Google.
- 6ai 1y agoShall we say … ASI is here ???
- hahahacorn 1y agoAnecdotally, as someone who operates in a very large legacy codebase, I am very impressed by GPT-5's agentic abilities so far. I've given it the same tasks I've given Claude and previous iterations via the Codex CLI, and instead of getting loss due to the massive scope of the problem, it correctly identifies the large scope and breaks it down into it's correct parts and creates the correct plan and begins executing. I am wildly impressed. I do not believe that the 0.x% increase in benchmarks tell the story of this release at all.
- gwd 1y agoI'm a solo founder. I fed it a fairly large "context doc" for the core technology of my company, current state of things, and the business strategy, mostly generated with the help of Claude 4, and asked it what it thought. It came back with a massive list of detailed ambiguities and inconsistencies -- very direct and detailed. The only praise was the first sentence of the feedback: "The core idea is sound and well-differentiated." It's got quite a different feel so far.
- arresin 1y agoWhat are you using to run it?
- joshmlewis 1y agoIt's a really good model from my testing so far. You can see the difference in how it tries to use tools to the greatest extent when answering a question, especially compared to 4.1 and o3. In this example it used 6! tool calls in the first response to try and collect as much info as possible. https://promptslice.com/share/b-2ap_rfjeJgIQsG https://promptslice.com/share/b-2ap_rfjeJgIQsG
- hollownobody 1y ago720 tool calls? Amazing!
- joshmlewis 1y agoWhere'd you get 720 from?
- terhechte 1y agothe _6!_
- brian626 1y agoMath pun… 6! = Factorial(6) = 720
- joshmlewis 1y agoWhoosh, it went right over my head.
- Zone3513 1y agoThat movie doesn't even exist. There is no Thunder Run from 2025.
- joshmlewis 1y agoThe data is made up, the point is to see how models respond to the same input / scenario. You're able to create whatever tools you want and import real data or it'll generate fake tool responses for you based on the prompt and tool definition. Disclaimer: I made PromptSlice for creating and comparing prompts, tools, and models.
- psyclobe 1y agoClaude Opus 4 has changed my workflow; never going back.
- SV_BubbleTime 1y agoIt would be very difficult to convince me 6 months ago that I would be happy to pay $100 for an AI service. Here we are.
- revskill 1y agoHow do people actually without ai models ???
- UrineSqueegee 1y agopretty underwhelming results so far for me
- deleted 1y ago[deleted]
- zone411 1y agoOn the Extended NYT Connections benchmark, GPT-5 Medium Reasoning scores close to o3 Medium Reasoning, and GPT-5 Mini Medium Reasoning scores close to o4-Mini Medium Reasoning: https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/
- jdelman 1y agoWhenever OpenAI releases a new ChatGPT feature or model, it's always a crapshoot when you'll actually be able to use it. The headlines - both from tech media coverage and OpenAI itself - always read "now available", but then I go to ChatGPT (and I'm a paid pro user) and it's not available yet. As an engineer I understand rollouts, but maybe don't say it's generally available when it's not?
- andrelaszlo 1y agoI asked GPT about it: > You are using the newest model OpenAI offers to the public (GPT-4o). There is no “GPT-5” model accessible yet, despite the splashy headlines.
- h4ch1 1y agoI can use it with the Github Copilot Pro plan.
- replwoacause 1y agoWeird. I got it immediately. I actually found out about it when I opened the app and saw it and thought “oh, a new model just dropped better go check YT for the video” which had just been uploaded. And I’m just a Plus user.
- DaveZale 1y ago[flagged]
- meribold 1y agoSad to see GPT-4.5 being gone. It knew things. More than any other model I'm aware of.
- mrits 1y agoI can't imagine anyone leaving this comment besides GPT-4.5
- deleted 1y ago[deleted]
- zone411 1y agoGPT-5 set a new record on my Confabulations on Provided Texts benchmark: https://github.com/lechmazur/confabulations/ https://github.com/lechmazur/confabulations/
- gtirloni 1y agoIf they ever wanted to IPO, maybe now is not the best time.
- throw03172019 1y agoHas anyone figured out how to not be forced to use GPT5 in chat gpt?
- Jordan-117 1y agoThey said they deprecated all their older models.
- adammarples 1y agoWhich is bigger, 9.9 or 9.11? Well it insta-failed my first test question
- w10-1 1y ago> a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent I'd love to see factors considered in the algorithm for system-1 vs system 2 thinking. Is "complexity" the factor that says "hard problem"? Because it's often not the complexity that makes it hard.
- epistemovault 1y agoIf AGI really arrives, will it run the world—or just binge Netflix and complain about being tired like the rest of us?
- mkoubaa 1y agoHyPeRbOlIc SiNgUlArItY
- gnulinux 1y agoMy first impressions: not impressed at all. I tried using this for my daily tasks today and for writing it was very poor. For this task o3 was much better. I'm not planning on using this model in the upcoming days, I'll keep using Gemini 2.5 Pro, Claude Sonnet, and o3.
- Telemakhos 1y agoI am thoroughly unimpressed by GPT-5. It still can't compose iambic trimeters in ancient Greek with a proper penthemimeral cæsura, and it insists on providing totally incorrect scansion of the flawed lines it does compose. I corrected its metrical sins twice, which sent it into "thinking" mode until it finally returned a "Reasoning failed" error. There is no intelligence here: it's still just giving plausible output. That's why it can't metrically scan its own lines or put a cæsura in the right place.
- taylorlapeyre 1y agoIt once again completely fails on an extremely simple test: look at a screenshot of sheet music, and tell me what the notes are. Producing a MIDI file for it (unsurprisingly) was far beyond its capabilities. https://chatgpt.com/share/68954c9e-2f70-8000-99b9-b4abd69d1aba https://chatgpt.com/share/68954c9e-2f70-8000-99b9-b4abd69d1a... This is not anywhere remotely close to general intelligence.
- adrianh 1y agoInterpreting sheet music images is very complex, and I’m not surprised general-purpose LLMs totally fail at it. It’s orders of magnitude harder than text OCR, due to the two-dimensional-ness. For much better results, use a custom trained model like the one at Soundslice: https://www.soundslice.com/sheet-music-scanner/ https://www.soundslice.com/sheet-music-scanner/
- ipnon 1y agoThis is a great test because it’s something you could teach an elementary school kid in an hour.
- Davidzheng 1y agois this a joke
- Telemakhos 1y ago
- felixfurtak 1y agoIt's still terrible at Wordle. This is one of my benchmarks.
- Aeolun 1y agoI'm just sitting here hoping that their lowered prices will force Anthropic to follow suit xD
- sbinnee 1y agoI didn't know that OpenAI added what they call organization verification process for API calls for some models. While I haven't noticed this change at work using OpenAI models, when I wanted to try GPT-5 on my personal laptop, I came across this obnoxious verification issue. It seems that it's all because that users can get thinking traces from API calls, and OpenAI wants to prevent other companies from distilling their models. Although I don't think OpenAI will be threatened by a single user from Korea, I don't want to go through this process for many reasons. But who knows that this kind of verification process may become norm and users will have no ways to use frontier models. "If you want to use the most advanced AI models, verify yourself so that we can track you down when something bad happens". Is it what they are saying?
- piskov 1y agoIt started with o-models in the API.
- deleted 1y ago[deleted]
- anshumankmr 1y agoI miss the model picker… is that just me?
- zombiwoof 1y agoGiven most of human intelligence isn’t that smart, AGI doesn’t seem hard
- cellis 1y agoMy favorite thing to ask is ascii art: _ _ _ __ ___ _ __ ___ __ _ __| (_) ___ | '_ \ / _ \| '_ _ \ / _ |/ _ | |/ __| | | | | (_) | | | | | | (_| | (_| | | (__ |_| |_|\___/|_| |_| |_|\__,_|\__,_|_|\___| What does this say? GPT 5: When read normally without the ASCII art spacing, it’s the stylized text for: markdown Copy Edit _ _ _ __ ___ _ __ ___ __ _ __| (_) ___ | '_ \ / _ \| '_ ` _ \ / _` |/ _` | |/ __| | | | | (_) | | | | | | (_| | (_| | | (__ |_| |_|\___/|_| |_| |_|\__,_|\__,_|_|\___| Which is the ASCII art for: rust — the default “Rust” welcome banner in ASCII style.
- beardedwizard 1y agoI asked it how to run the image and expose a port. it was just terrible in cursor. thought a Dockerfile wasn't in the repo, called no tools, then hallucinated a novel on dockefile best practices.
- obloid 1y agoSo far GPT-5 has not been able to pass my personal "Turing test" which has been unsuccessful for the past several years starting through various versions of Dall-e up to the latest model. I want it to create an image of Santa Claus pulling the sleigh with a reindeer in the sleigh holding the reins, driving the sleigh. No matter how I modify the prompt it is still unable to create this image that my daughter requested a few years ago. This is an image that is easily imagined and drawn by a small child yet the most advanced AI models still can't produce it. I think this is a good example that these models are unable to "imagine" something that falls outside of the realm of it's training data.
- ramzyo 1y agoIs this what you mean? https://chatgpt.com/share/6895632c-fb58-800e-b287-b7a98ad64db5 https://chatgpt.com/share/6895632c-fb58-800e-b287-b7a98ad64d...
- simultsop 1y agothat was smooth
- obloid 1y agoInteresting. Yes, that's basically what I've been going for but none of my prompts ever gave a satisfactory response. Plus I noticed you just copy/pasted from my initial comment and it worked. Weird. After my last post I was eventually able to get it to work by uploading an example image of Santa pulling the sleigh and telling it to use the image as an example, but I couldn't get it by text prompt alone. I guess I need to work on my prompt skills! https://chatgpt.com/share/689564d1-90c8-8007-b10c-8058c1491e47 https://chatgpt.com/share/689564d1-90c8-8007-b10c-8058c1491e...
- Google4567 1y ago[dead]
- deleted 1y ago[deleted]
- 1y ago
- deathflute 1y agoLots of debate here about the best model. The best model is the one which creates the most value for you —- this typically is a function of your skill in using the model for tasks that matter to you. Always was. Always will be.
- kkukshtel 1y agoSomething that's really hitting me is something brought up in this piece: https://www.interconnects.ai/p/gpt-5-and-bending-the-arc-of-progress https://www.interconnects.ai/p/gpt-5-and-bending-the-arc-of-... When a model comes out, I usually think about it in terms of my own use. This is largely agentic tooling, and I mostly us Claude Code. All the hallucination and eval talk doesn't really catch me because I feel like I'm getting value of these tools today. However, this model is not _for_ me in the same way models normally are. This is for the 800m or whatever people that open up chatgpt every day and type stuff in. All of them have been stuck on GPT-4o unbeknwst to them. They had no idea SOTA was far beyond that. They probably dont even know that there is a "model" at all. But for all these people, they just got a MAJOR upgrade. It will probably feel like turning the lights on for these people, who have been using a subpar model for the past year. That said I'm also giving GPT-5 a run in Codex and it's doing a pretty good job!
- techpineapple 1y agoI’m curious what this means. Maybe I’m stupid but I read through the sample gpt-4 vs got-5 and I largely couldn’t tell the difference and sometimes preferred the gpt-4 answer. But like what are the average 800 million people using this for that the average 800 million user will be able to see a difference? Maybe I’m a far below average user? But I can’t tell the difference between models in causal use. Unless you’re talking performance, apparently gpt-5 is much faster.
- MagicMoonlight 1y ago4o would start writing immediately without thinking. So if the first thing it wrote was “The world is flat because…” then it will continue to write as if the world is flat. It makes it very stupid, but very compliant. If you’re mentally ill it will go along with whatever delusions you have, without any objection.
- techpineapple 1y agoNoticed that Most of this Reddit AMA is about how great 4o is and how terrible 5 is in comparison: https://www.reddit.com/r/ChatGPT/comments/1mkae1l/gpt5_ama_with_openais_sam_altman_and_some_of_the/ https://www.reddit.com/r/ChatGPT/comments/1mkae1l/gpt5_ama_w...
- lutusp 1y agoI have a canonical test for chatbots -- I ask them who I am. I'm sufficiently unknown in modern times that it's a fair test. Just ask, "Who is Paul Lutus?" ChatGPT 5's reply is mostly made up -- about 80% is pure invention. I'm described as having written books and articles whose titles I don't even recognize, or having accomplished things at odds with what was once called reality. But things are slowly improving. In past ChatGPT versions I was described as having been dead for a decade. I'm waiting for the day when, instead of hallucinating, a chatbot will reply, "I have no idea." I propose a new technical Litmus test -- chatbots should be judged based on what they won't say.
- alenguo 1y agoI've already used it
- zastai0day 1y agoAll people are talking about GPT-5 all over the world, the competition is so intense that every major tech company is racing to develop their own advanced AI models.
- zshlk43t 1y ago[dead]
- throwpoaster 1y agoI’ve been working on an electrochemistry project, with several models but mostly o3-pro. GPT-5 refused to continue the conversation because it was worried about potential weapons applications, so we gave the business to the other models. Disappointing.
- deleted 1y ago[deleted]
- zone411 1y agoIt is the new leader on my Short Story Creative Writing benchmark: https://github.com/lechmazur/writing/ https://github.com/lechmazur/writing/
- saddat 1y agoIf Grol , Claude , ChatGPT seemingly still all scale , yet their Performance feels similar, could this mean that the Technology path is narrow, with little differentiations left ?
- tapland 1y agoUgh. Could they have their expert make a website that doesn’t crash safari on my iPhone SE? :)
- danjc 1y ago> describe gpt 5 in one word > incremental
- RobinL 1y agoHypothesis: to the average user this will feel like a much greater jump in capability then to the average HNer, because most users were not using the model selector. So it'll be more successful than the benchmarks suggest.
- tw1984 1y agojust wondering whether Altman is still going to promote his AGI/ASI coming in 12 months story.
- fergie 1y agoAnecdote: It can now speak in various Scots dialects- for example, it can convincingly create a passage in the style of Irvine Welsh. It can also speak Doric (Aberdonian). Before it came nowhere close.
- kiitos 1y agoabsolutely miserable results as an agent in my ide :<
- tennisflyi 1y agoHow/where do I see my chat history!?
- sarmasamosarma 1y ago[dead]
- nacholibrev 1y agoI've tried it in cursor and I didn't like it. The claude-4-sonnet gives me far better results. Also it's a lot slower than Claude and Google models. In general GPT models doesn't work well for me for both coding and general questions.
- energy123 1y agoOn livebench.ai, GPT-5 is the best model overall, and the second best for agentic coding. But for the Coding benchmarks, it's ranked like 20th. Quite interesting. I'm finding it exceptional for non-trivial summarization tasks.
- reportgunner 1y agoFirst OpenAI video I've ever seen, the people in it all seem incompetent for some reason, like a grotesque version of apple employees from temu or something.
- lynx97 1y agoNot impressed. gpt-5-nano gives noticeably worse results then o4-mini does. gpt-5 and gpt-5-mini are both behind the verification wall, and can stay there if they like.
- froh42 1y agoWow, I just got GPT-5. Tried to continue the discussion of my 3D print problems with it (which I started with 4o). In comparison GPT-5 is an entitled prick trying to gaslight me into following what it wants. Can I have 4o back?
- withinboredom 1y agoIf we're going to be forced to trust a new model, might as well evaluate other companies as well to make a decision before my plan renews.
- nodesocket 1y agoWhy do I have access to GPT-5 on only some of my devices? All logged into my plus account. My iPad ChatGPT shows 5, but my iPhone ChatGPT only allows 4o?
- withinboredom 1y agorollout is probably not user-specific, but device specific. Classic rookie mistake.
- nodesocket 1y agoYa, strange rollout. My browser session which I use by far the most with ChatGPT is also still stuck on 4o.
- primaprashant 1y agocreated a summary of comments from this thread about 15 hours after it had been posted and had 1983 comments, using gpt-5-high and gemini-2.5-pro using a prompt similar to simonw [1]. Used a Python script [2] that I wrote to generate the summary. - gpt-5-high summary: https://gist.github.com/primaprashant/1775eb97537362b049d643eae439829a https://gist.github.com/primaprashant/1775eb97537362b049d643... - gemini-2.5-pro summary: https://gist.github.com/primaprashant/4d22df9735a1541263c671155eecbc39 https://gist.github.com/primaprashant/4d22df9735a1541263c671... [1]: https://news.ycombinator.com/item?id=43477622 https://news.ycombinator.com/item?id=43477622 [2]: https://gist.github.com/primaprashant/f181ed685ae563fd06c49d3d49a8dd9b https://gist.github.com/primaprashant/f181ed685ae563fd06c49d...
- mustaphah 1y agoWhy not use the ChatGPT interface instead of the API to save credits? Pass the cookies.
- primaprashant 1y agoOnly have access to GPT-5 through API for now. The amount of tokens (>130k) used is higher than the limit of ChatGPT (128k) so it wouldn't really work well.
- jiggawatts 1y agoWow, the 2.5 Pro summary is far better, it reads like coherent English instead of a list of bullet points.
- mustaphah 1y agoSomeone should start a Gemini-powered blog that distills the top HN posts into concise summaries.
- primaprashant 1y agoyes, agreed. Context length might be playing a factor as total number of prompt tokens is >120k. Performance of LLMs generally degrade at higher context length.
- Applejinx 1y agoI am very puzzled that I cannot search for the word 'blueberry' in this HN discussion. Is my browser broken, or is the subject inappropriate to raise in this community?
- jsumrall 1y agoIt seems 'GPT-5 Pro' is not available via the API.
- junon 1y agoAnecdotal review: Been using it all morning. Had to switch back to 4. 5 has all of the problems that 2/3 had with ignoring any context, flagrantly ignoring the 'spirit' of my requests, and talking to me like I'm a little baby. Not to mention almost all of my prompts result in a several minute wait with "thinking longer about the answer".
- getcrunk 1y agoYea I see this a lot with Gemini since 2.5 Very stubborn and “opinionated” I think most models will tend this way (to consolidate more control over how we “think” and what we believe)
- ismailmaj 1y agoWent over my last conversations with Gemini 2.5 and asked the same things to GPT-5 with thinking on, the latter was consistently worse both in content and form. I wouldn't have guessed Gemini to win the AI race in 2025 but here we are.
- sidibe 1y agoI would be surprised if they didn't, just from the difference in number of employees and resources. Google can pursue 20x as many dead ends, anthropic and openai have to commit to a few things and hope they're right
- junon 1y agoUpdate shortly after my post: They've removed access to GPT-4 and below. Therefore I've removed their access to my card.
- monster_truck 1y agoI'm extremely whelmed. I cancelled my subscription
- ritzaco 1y agoOk this[0] sounds very, uh bold to me? Surely this is going to break a ton of workflows etc seemingly with nearly no notice? I'm assuming 'launches' equates with 'fully rolls out' or something but it's not that clear to me. When GPT-5 launches, several older models will be retired, including: - GPT-4o - GPT-4.1 - GPT-4.5 - GPT-4.1-mini - o4-mini - o4-mini-high - o3 - o3-pro If you open a conversation that used one of these models, ChatGPT will automatically switch it to the closest GPT-5 equivalent. Chats with 4o, 4.1, 4.5, 4.1-mini, o4-mini, or o4-mini-high will open in GPT-5, chats with o3 will open in GPT-5-Thinking, and chats with o3-Pro will open in GPT-5-Pro (available only on Pro and Team). [0] https://help.openai.com/en/articles/11909943-gpt-5-in-chatgpt https://help.openai.com/en/articles/11909943-gpt-5-in-chatgp...
- artursapek 1y agoYeah I was surprised how fast they rugged 4. I guess they want to concentrate their hardware on 5.
- hoppp 1y agoIf it costs the same compute to run it then there is no point running worse models
- boringg 1y agoThat's assuming all else holds on the model which isn't always clear.
- SkyPuncher 1y ago"Worse" model is largely subjective. Often, task specific. For me, I find model upgrades frustrating as they often break subtle things about my workflows while not clearly offering an improvement. It takes time to learn the nuances of each model and tweak your prompts to get the best outputs. For example, Sonnet 4 is now my daily driver for Cursor - but it took me nearly a month to tweak my approaches I was using for 3.5 and 3.7.
- energy123 1y ago> "If you’re on Plus or Team, you can also manually select the GPT-5-Thinking model from the model picker with a usage limit of up to 200 messages per week." And what's the reasoning effort parameter set to?
- miroljub 1y agoNow it's a perfect time for DeepSeek to finally release R2.
- Razengan 1y agoI asked ChatGPT 5 about the main differences between 4 and 5, and it said: "I couldn’t find any credible, up-to-date details on a model officially named “GPT-5” or formal comparisons to “GPT-4o.” It’s possible that GPT-5, if it exists, hasn't been announced publicly or covered in verifiable sources … GPT-5 as of August 8, 2025 has no formal release announcement" Reassuring.
- Jolliness7501 1y agoI asked it to count letters in his answer. First it asked me what answer, then suggested Python code that will count letter in gpt api reply, then gave wrong answer, then dropped my connection. Wonderfull product. AGI is so near, you u can almost smell it...
- neofytos 1y agoFor web app generation, gpt-5 seems to route my prompt to the reasoning model. Interestingly, gpt-4.1 often produced more creative or aesthetically interesting designs. In some cases, a bit of controlled hallucination (via temperature) is actually desirable, especially when generating UI ideas. Anyone else seeing this tradeoff?
- BoorishBears 1y agoAll of these models generally produce hideous sites that because for some unknown reason we all decided shadcn and tailwind-with-no-design-system should define AI generated UI. I use image generation for UI layout and have Claude implement it with an actual UI library: usually MUI with theming, but honestly anything with a sane grid system is better than loose tailwind.
- neofytos 1y agoI hear you, though I’m less focused on the choice of UI libraries and more on how little control we have over how prompts are routed internally. Sometimes I want a model that reasons deeply, other times I want one that is more creative. Right now, it feels like gpt-5 forces everything through the same pipeline, even when a different mode would be better suited to the task.
- deezmofonutz 1y ago99% of my sales team could not run a MCP server if their lives depended on it.
- accrual 1y agoAlready a lot of comments here (2188 at the time of this comment) but wanted to share my 2c: * It feels a bit more competent, as if it had more nuance or detail to say about each point. * It got a few obscure details about OpenBSD correct right away - both Sonnet 4 and 4o sometimes conflate Linux and OpenBSD commands. * It was fun asking GPT-5 to not only answer the query, but also to provide a brief analysis of the query itself for insights into myself! Not a detailed review, but just a couple things I noticed with some limited usage.
- cube00 1y ago> It was fun asking GPT-5 [...] to provide [...] insights into myself! Were those insights of a glowing and positive nature by chance?
- accrual 1y agoI felt the default GPT-5 responses were neutral. I checked and it hasn't yet emitted any emojis or exclamation marks, even though I used some in my replies. When I asked for a review of my question it focused on my query rather than me. Will need to test longer contexts though - I've noticed Sonnet 4 becomes a bit less stoic and more friendly in some longer chats, but maybe it's just reflecting my casual language back at me.
- AstroBen 1y agoThat's an incredible question and you're so wise to be thinking about that!
- balder1991 1y ago“Good catch!” Damn, I hate that.
- Argonaut998 1y agoThey really nerfed Plus[0]. 80 messages every 3 hours for normal GPT-5. And only 200 messages per week for GPT-5 Thinking. It seems like terrible value. Before it was: 100 o3 per week 100 o4-mini-high per day 300 o4-mini per day 50 4.5 per week [0] https://help.openai.com/en/articles/11909943-gpt-5-in-chatgpt https://help.openai.com/en/articles/11909943-gpt-5-in-chatgp...
- nashashmi 1y agoImagine the 90s being “you can only search x times a day”. Will we look back at the 20s in the same way.
- ripped_britches 1y agoExcept AI isn’t ad supported
- nojs 1y agoAs Homer would say, it isn’t ad supported so far!
- floatrock 1y agoThe search limits are surely mostly from limits on energy and compute costs. The better analogue is "Imagine in the 70's being able to teletype into an insanely expensive compute infrastructure and have reasonable timesharing capabilities of a limited resource across multiple users." Unix. I'm describing the motivation for Unix there. We already look back on earlier times with constraints that were appropriate. Presumably compute will get cheaper, we'll build more datacenters, maybe we'll even power them in a way that doesn't destroy our planet, and GPT questions will become too cheap to meter. Just give it some time.
- jstummbillig 1y agoIf the past 2 years are any indicator cost per unit of capability will keep falling, rapidly.
- 1y ago
- gooseus 1y agoNeat, but it still fails my Forth simulator test, albeit a bit better than the last time. I hope they didn't have to raise the temperature of the ocean by too much to get these marginal improvements that definitely aren't even close to the kind of "general intelligence" you could get from training a clever child or monkey with 500 calories of sweets. https://chatgpt.com/share/689525f4-20f0-8003-8bf6-f1f21dde6b11 https://chatgpt.com/share/689525f4-20f0-8003-8bf6-f1f21dde6b... You know what would be more impressive? If it said "Hey, I'm actually not designed to simulate a Forth machine accurately, I'm only going to be able to approximate it (poorly)). If you want an accurate Forth machine you should just implement this code: [Simple Forth Implementation]". Or better yet, it could recognize when it was being asked to "be" a machine, and instead spin up a side process with the machine implementation and redirect any prompts to that process until a "STOP" token is reached.
- morelandjs 1y agoAnyone know what the deal is with connectors? I don’t see them in the app and they made it sound like Google Calendar would be made broadly available as a connector.
- jdlyga 1y agoGPT-5 seems to be a rollup of their previous models. It reminds me of Cursor's "auto" mode, which users aren't particularly happy with either.
- 00deadbeef 1y agoIs anyone else having problems with factual correctness? I had a number of 4o and o3 conversations going and those models were factually correct about a number of different subjects. Asking GPT-5 about the same things results in wrong answers even though its training data is newer. And it won't look things up to correct itself unless I manually switch to the thinking variant. This is worse. I cancelled my subscription.
- danielscrubs 1y agoSynthetic data. Get used to it, it’s in vogue.
- thehappypm 1y agoExample?
- MollyRealized 1y agoHow does one get rid of the rainbow background?
- croemer 1y agoFirst impressions: the emoji trigger happiness of 4o is totally gone. Bolding still happens. There appear to be 4 ways to run a query now: a) GPT5, b) GPT5 and toggle "extra thinking" on, c) "GPT5 with thinking", and d) "GPT5 with thinking" then click "quick answer" which aborts thinking (this mode is possibly identical with GPT5) I don't find this much simpler than 4o, o3, etc. It's just reordering the hierarchies. Now the model name is no longer descriptive at all and one has to add which mode one ran it in.
- m_a_g 1y agoIf it’s all going to be step changes from now on, doesn’t this mean we’re in an AI bubble and it might burst at any moment?
- boyesm 1y agoIs this video generated? The hands of the initial presenter look slightly unnatural.
- trane_project 1y agoFinally got access to it. It's so awful. I asked it something, answered in Spanish with something completely different. In another conversation, it kept giving me completely different answers to something I didn't even ask. Telling it to stop doesn't do anything. It ignores it and continues a conversation with itself.
- real_marcfawzi 1y agofeels like O3 but faster (in quick answer mode), not any smarter .. thinking mode takes forever and the results are mediocre
- koll 1y agoIn terms of coding, claude opus 4.1 is still the powerhorse.
- koll 1y agoIn terms of coding Claude Opus 4.1 still is the power horse.
- jaybrendansmith 1y agoIt does make me wonder about human training, aka education. Could it be that ingesting/reading a better quality literature, math, and other topics would necessarily result in a higher intelligence?