24 ms·
GPT-5.2
https://platform.openai.com/docs/guides/latest-model https://platform.openai.com/docs/guides/latest-model
System card: https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944f8d/oai_5_2_system-card.pdf https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944...
- Croftengea 10mo agoIs this another GPT-4.5?
- DICloak 10mo ago[dead]
- villgax 10mo agoMarginal gains for exorbitantly pricey and closed model…..
- sfmike 10mo agoEverything is still based on 4 4o still right? is a new model training just too expensive? They can consult deepseek team maybe for cost constrained new models.
- catigula 10mo agoThe irony is that Deepseek is still running with a distilled 4o model.
- blovescoffee 10mo agoSource?
- verdverm 10mo agoApparently they have not had a successful pre training run in 1.5 years
- fouronnes3 10mo agoI want to read a short scify story set in 2150 about how, mysteriously, no one has been able to train a better LLM for 125 years. The binary weights are studied with unbelievably advanced quantum computers but no one can really train a new AI from scratch. This starts cults, wars and legends and ultimately (by the third book) leads to the main protagonist learning to code by hand, something that no human left alive still knows how to do. Could this be the secret to making a new AI from scratch, more than a century later?
- armenarmen 10mo agoI’d read it!
- barrenko 10mo agoMonsieur, if I may offer a vaaaguely similar story on how things may progress https://www.owlposting.com/p/a-body-most-amenable-to-experimentation https://www.owlposting.com/p/a-body-most-amenable-to-experim...
- verdverm 10mo agoYou can ask 2025 Ai to write such a book, it's happy to comply and may or may not actually write the book https://www.pcgamer.com/software/ai/i-have-been-fooled-reddit-user-endures-the-roasting-of-a-lifetime-after-asking-how-to-download-a-487mb-book-they-worked-on-with-chatgpt-for-over-2-weeks/ https://www.pcgamer.com/software/ai/i-have-been-fooled-reddi...
- WhyOhWhyQ 10mo agoThere's a scifi short story about a janitor who knows how to do basic arithmetic and becomes the most important person in the world when some disaster happens. Of course after things get set up again due to his expertise, he becomes low status again.
- bradfitz 10mo agoI had to go look that up! I assume that's https://en.wikipedia.org/wiki/The_Feeling_of_Power https://en.wikipedia.org/wiki/The_Feeling_of_Power ? (Not a janitor, but "a low grade Technician"?)
- Wowfunhappy 10mo agoI thought whenever the knowledge cutoff increased that meant they’d trained a new model, I guess that’s completely wrong?
- brokencode 10mo agoTypically I think, but you could pre-train your previous model on new data too. I don’t think it’s publicly known for sure how different the models really are. You can improve a lot just by improving the post-training set.
- rockinghigh 10mo agoThey add new data to the existing base model via continuous pre-training. You save on pre-training, the next token prediction task, but still have to re-run mid and post training stages like context length extension, supervised fine tuning, reinforcement learning, safety alignment ...
- astrange 10mo agoContinuous pretraining has issues because it starts forgetting the older stuff. There is some research into other approaches.
- elgatolopez 10mo agoWhere did you get that from? Cutoff date says august 2025. Looks like a newly pretrained model
- SparkyMcUnicorn 10mo agoIf the pretraining rumors are true, they're probably using continued pretraining on the older weights. Right?
- deleted 10mo ago[deleted]
- FergusArgyll 10mo ago> This stands in sharp contrast to rivals: OpenAI’s leading researchers have not completed a successful full-scale pre-training run that was broadly deployed for a new frontier model since GPT-4o in May 2024, highlighting the significant technical hurdle that Google’s TPU fleet has managed to overcome. - https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-swing-at-the https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-s... It's also plainly obvious from using it. The "Broadly deployed" qualifier is presumably referring to 4.5
- ric2b 10mo agoHow is that a technical hurdle if they obviously were able to do it before? It's probably just a question of cost/benefit analysis, it's very expensive to do, so the benefits need to be significant.
- zamadatix 10mo agohttps://openai.com/index/introducing-gpt-5-2/ https://openai.com/index/introducing-gpt-5-2/
- system2 10mo ago"Investors are putting pressure, change the version number now!!!"
- exe34 10mo agoI'm quite sad about the S-curve hitting us hard in the transformers. For a short period, we had the excitement of "ooh if GPT-3.5 is so good, GPT-4 is going to be amazing! ooh GPT-4 has sparks of AGI!" But now we're back to version inflation for inconsequential gains.
- verdverm 10mo ago2025 is the year most Big AI released their first real thinking models Now we can create new samples and evals for more complex tasks to train up the next gen, more planning, decomp, context, agentic oriented OpenAI has largely fumbled their early lead, exciting stuff is happening elsewhere
- ToValueFunfetti 10mo agoTake this all with a grain of salt as it's hearsay: From what I understand, nobody has done any real scaling since the GPT-4 era. 4.5 was a bit larger than 4, but not as much as the orders of magnitude difference between 3 and 4, and 5 is smaller than 4.5. Google and Anthropic haven't gone substantially bigger than GPT-4 either. Improvements since 4 are almost entirely from reasoning and RL. In 2026 or 2027, we should see a model that uses the current datacenter buildout and actually scales up.
- snovv_crash 10mo agoDatacenter capacity is being snapped up for inference too though.
- Leynos 10mo ago4.5 is widely believed to be an order of magnitude larger than GPT-4, as reflected in the API inference cost. The problem is the quantity of parameters you can fit in the memory of one GPU. Pretty much every large GPT model from 4 onwards has been mixture of experts, but for a 10 trillion parameter scale model, you'd be talking a lot of experts and a lot of inter-GPU communication. With FP4 in the Blackwell GPUs, it should become much more practical to run a model of that size at the deployment roll-out of GPT-5.x. We're just going to have to wait for the GBx00 systems to be physically deployed at scale.
- Xiol 10mo ago[flagged]
- tabletcorry 10mo agoSlight increase in model cost, but looks like benefits across the board to match. gpt-5.2 $1.75 $0.175 $14.00 gpt-5.1 $1.25 $0.125 $10.00
- llmslave 10mo agoThey probably just beefed up compute run time on the what is the same underlying model
- deleted 10mo ago[deleted]
- deleted 10mo ago[deleted]
- jtbayly 10mo ago40% increase is not "slight."
- credit_guy 10mo agoNot the OP, but I think "slight" here is in relation to Anthropic and Google. Claude Opus 4.5 comes at $25/MT (million tokens), Sonnet 4.5 at $22.5/MT, and Gemini 3 at $18/MT. GPT 5.2 at $14/MT is still the cheapest.
- deaux 10mo agoYour numbers are very off. $25 - Opus 4.5 $15 - Sonnet 4.5 $14 - GPT 5.2 $12 - Gemini 3 Pro Even if you're including input, your numbers are still off.
- credit_guy 10mo agoI used the pricing for long context (>200k) in all cases. I personally use AI as coding assistants, like lots of other people, and as such, hitting and exceeding 200k is quite the norm. The numbers you are showing are for <200k context length.
- meetpateltech 10mo agoGPT-5.2 System Card PDF: https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944f8d/oai_5_2_system-card.pdf https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944...
- dang 10mo agoThanks, we'll put that in the toptext as well.
- josalhor 10mo agoFrom GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!
- verdverm 10mo agoWe're also in benchmark saturation territory. I heard it speculated that Anthropic emphasizes benchmarks less in their publications because internally they don't care about them nearly as much as making a model that works well on the day-to-day
- quantumHazer 10mo agoSeems pretty false if you look at the model card and web site of Opus 4.5 that is… (check notes) their latest model.
- verdverm 10mo agoBuilding a good model generally means it will do well on benchmarks too. The point of the speculation is that Anthropic is not focused on benchmaxxing which is why they have models people like to use for their day-to-day. I use Gemini, Anthropic stole $50 from me (expired and kept my prepaid credits) and I have not forgiven them yet for it, but people rave about claude for coding so I may try the model again through Vertex Ai... The person who made the speculation I believe was more talking about blog posts and media statements than model cards. Most ai announcements come with benchmark touting, Anthropic supposedly does less / little of this in their announcements. I haven't seen or gathered the data to know what is truth
- doctoboggan 10mo agoThis seems like another "better vibes" release. With the number of benchmarks exploding, random luck means you can almost always find a couple showing what you want to show. I didn't see much concrete evidence this was noticeably better than 5.1 (or even 5.0). Being a point release though I guess that's fair. I suspect there is also some decent optimizations on the backend that make it cheaper and faster for OpenAI to run, and those are the real reasons they want us to use it.
- rat9988 10mo ago> I didn't see much concrete evidence this was noticeably better than 5.1 Did you test it?
- doctoboggan 10mo agoNo, I would like to but I don't see it in my paid ChatGPT plan or in the API yet. I based my comment solely off of what I read in the linked announcement.
- sebzim4500 10mo ago>I suspect there is also some decent optimizations on the backend that make it cheaper and faster for OpenAI to run, and those are the real reasons they want us to use it. I doubt it, given it is more expensive than the old model.
- BrtByte 10mo agoAt this point the benchmark soup is so dense that it's hard to tell signal from selective framing
- _7u7v 10mo agoIt baffles me to see these last 2 announcements (GPT 5.1 as well) devoid of any metrics, benchmarks or quantitative analyses. Could it be because they are behind Google/Anthropic and they don't want to admit it? (edit: I'm sorry I didn't read enough on the topic, my apologies)
- zamadatix 10mo agoThis isn't the announcement, it's the developer docs intro page to the model - https://openai.com/index/introducing-gpt-5-2/ https://openai.com/index/introducing-gpt-5-2/. Still doesn't answer cross-comparison, but at least has benchmark metrics they want to show off.
- fulafel 10mo agoSo GDPval is OpenAI's own benchmark. PDF link: https://arxiv.org/pdf/2510.04374 https://arxiv.org/pdf/2510.04374
- minadotcom 10mo agoThey used to compare to competing models from Anthropic, Google DeepMind, DeepSeek, etc. Seems that now they only compare to their own models. Does this mean that the GPT-series is performing worse than its competitors (given the "code red" at OpenAI)?
- poormathskills 10mo agoOpenAI has never compared their models to models from other labs in their blog post. Open literally any past model launch post to see that.
- tabletcorry 10mo agoThe matrix required for a fair comparison is getting too complicated, since you have to compare chat/thinking/pro against an array of Anthropic and Google models. But they publish all the same numbers, so you can make the full comparison yourself, if you want to.
- Tiberium 10mo agoThey did compare it to other models: https://x.com/OpenAI/status/1999182104362668275 https://x.com/OpenAI/status/1999182104362668275 https://i.imgur.com/e0iB8KC.png https://i.imgur.com/e0iB8KC.png
- enlyth 10mo agoThis looks cherry-picked, for example Claude Opus had a higher score on SWE-Bench Verified so they conveniently left it out, also GDPval is literally a benchmark made by OpenAI
- mattas 10mo agoAre benchmarks the right way to measure LLMs? Not because benchmarks can be gamed, but because the most useful outputs of models aren't things that can be bucketed into "right" and "wrong." Tough problem!
- Sir_Twist 10mo agoNot an expert in LLM benchmarks, but I generally I think of benchmarks as being good particularly for measuring usefulness for certain usecases. Even if measuring LLMs is not as straightforward as, say, read/write speeds when comparing different SSDs, if a certain model's responses are consistently measured as being higher quality / more useful, surely that means something, right?
- olliepro 10mo agoDo you have a better way to measure LLMs? Measurement implies quantitative evaluation... which is the same as benchmarks.
- Wowfunhappy 10mo agoI don’t have a good way to measure them, but I think they should be evaluated more like how we evaluate movies, or restaurants. Namely, experienced critics try them and write reviews.
- olliepro 10mo agoIt feels like this should work, but the breadth of knowledge in these models is so vast. Everyone knows how to taste, but not everyone knows physics, biology, math, every language… poetry, etc. Enumerating the breadth of valuable human tasks is hard, so both approaches suffer from the scale of the models’ surface area. An interesting problem since the creators of OLMO have mentioned that throughout training, they use 1/3 or their compute just doing evaluations. Edit: One nice thing about the “critic” approach is that the restaurant (or model provider) doesn’t have access to the benchmark to quasi-directly optimize against.
- k2xl 10mo agoThe ARC AGI 2 bump to 52.9% is huge. Shockingly GPT 5.2 Pro does not add too much more (54.2%) for the increase cost.
- zug_zug 10mo agoFor me the last remaining killer feature of ChatGPT is the quality of the voice chat. Do any of the competitors have something like that?
- FrasiertheLion 10mo agoTry elevenlabs
- sosodev 10mo agoDoes elevenlabs have a real-time conversational voice model? It seems like like their focus is largely on text to speech and speech to text. Which can approximate that type of thing but it's not at all the same as the native voice to voice that 4o does.
- dragonwriter 10mo ago> Does elevenlabs have a real-time conversational voice model? Yes. > It seems like like their focus is largely on text to speech and speech to text. They have two main broad offerings (“Platforms”); you seem to be looking at what they call the “Creative Platform”. The real-time conversational piece is the centerpiece of the “Agents Platform”.
- sosodev 10mo agoIt specifically says in the architecture docs for the agents platform that it's STT (ASR) -> LLM -> TTS https://elevenlabs.io/docs/agents-platform/overview#architecture https://elevenlabs.io/docs/agents-platform/overview#architec...
- deleted 10mo ago[deleted]
- hi_im_vijay 10mo ago[disclaimer, i work at elevenlabs] we specifically went with a cascading model for our agents platform because it's better suited for enterprise use cases where they have full control over the brain and can bring their own llm. with that said, even with a cascading model, we can capture a decent amount of nuance with our asr model, and it also supports capturing audio events like laughter or coughing. a true speech to speech conversational model will perform better on things like capturing tone, pronouncations, phonetics, etc, but i do believe we'll also get better at that on the asr side over time.
- Tiberium 10mo agoThe only table where they showed comparisons against Opus 4.5 and Gemini 3: https://x.com/OpenAI/status/1999182104362668275 https://x.com/OpenAI/status/1999182104362668275 https://i.imgur.com/e0iB8KC.png https://i.imgur.com/e0iB8KC.png
- varenc 10mo ago100% on the AIME (assuming its not in the training data) is pretty impressive. I got like 4/15 when I was in HS...
- hellojimbo 10mo agoThe no tools part is impressive, with tools every model gets 100%
- varenc 10mo agoIf I recall, the AIME answers are always 4 digits numbers. And most of the problems are of the type where if you have a candidate number it's reasonable to validate its correctness. So easy to brute force all 4 digit ints with code. tl;dr; humans would do much better too if they could use programming tools :)
- Davidzheng 10mo agouh no it's not solved by looping over 4 digit numbers when it uses tools
- JanSt 10mo agoThe benchmarks are very impressive. Codex and Opus 4.5 are really good coders already and they keep getting better. No wall yet and I think we might have crossed the threshold of models being as good or better than most engineers already. GDPval will be an interesting benchmark and I'll happily use the new model to test spreadsheet (and other office work) capabilities. If they can going like this just a little bit further, much of the office workers will stop being useful.... I don't know yet how to feel about this. Great for humanity probably but but for the individuals?
- llmslave 10mo agoYeah theres no wall on this. It will be able to mimic all of human behavior given proper data.
- ionwake 10mo agoit was only about 2-3 weeks when several HNers told me "nah you better re-check your code", when I explained I have over 2 decades xp of coding, yet have not manually edited code (in memory) for the last 6 or so months, whilst performing daily 12 hour daily vibe code seshes
- ipsum2 10mo agoIt really depends on the complexity of code. I've found models (codex-5.1-max, opus 4.5) to be absolutely useless writing shaders or ML training code, but really good at basic web development.
- sheeshe 10mo agoWhich is no surprise as the data for web development stuff exists in large amounts on the web that the models feed off.
- nineteen999 10mo agoInteresting, I've been using Claude Max with UE5 and while it isn't _brilliant_ with shaders I can usually get it to where I want. Also had a bit of success with converting HLSL shaders to GLSL with it.
- breakingcups 10mo agoIs it me, or did it still get at least three placements of components (RAM and PCIe slots, plus it's DisplayPort and not HDMI) in the motherboard image[0] completely wrong? Why would they use that as a promotional image? 0: https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3ft1oP/2ee4e2a98c4725fab4e9eada8d38b6ad/image_8.png?w=1920&q=90&fm=webp https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3f...
- dolmen 10mo agoNot that bad compared to product images seen on AliExpress.
- timerol 10mo agoAlso a "stacked pair" of USB type-A ports, when there are clearly 4
- tedsanders 10mo agoYep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.
- dandiep 10mo agoStill no GPT 5.x fine tuning? I emailed support a while back to see if there was an early access program (99.99% sure the answer is yes). This is when I discovered that their support is 100% done by AI and there is no way to escalate a case to a human.
- jazzyjackson 10mo agoContainment breach is going to occur from a disgruntled customer convincing the customer service bot it needs to get a hold of a supervisor
- dandiep 10mo agoI tried this! Me: Is it possible get early access to the GPT 5 fine tuning? It is much better at non-English languages, which is the majority of what we do, so it would be greatly beneficial to us. Bot: I'm an AI support agent. I don't have information about early access to GPT-5 fine-tuning or unannounced features. Currently, there is no public information or program for early access to GPT-5 fine-tuning. Please keep an eye on official OpenAI channels for updates. Let me know if you have any questions about current features or products! Me: Can you ask someone from the team if I can be part of the early access? I know there is no public information, but there is an early access program and I would like to be part of it. Bot: I'm an AI support agent and do not have the ability to forward requests or make inquiries about early access programs with the OpenAI team. If early access becomes available, it will be announced through official channels. Let me know if you have any other questions about OpenAI products or current fine-tuning options. Me: How do I get in touch with the open ai team? Bot: I'm an AI support agent and can only assist with OpenAI products or services through this support channel. There is no separate contact method for the OpenAI team beyond official support and public announcements. If you have questions about current products or features, I'm here to help Should have tried a jailbreak maybe...
- jazzyjackson 10mo agoYea you might have to convince it someone is in mortal danger or something
- firebot 10mo ago[flagged]
- orliesaurus 10mo agoI told all my friends to upgrade or they're not my friends anymore /s
- DeathArrow 10mo agoPricing is the same?
- tedsanders 10mo agoChatGPT pricing is the same. API pricing is +40% per token, though greater token efficiency means that cost per task is not always that much higher. On some agentic evals we actually saw costs per task go down with GPT-5.2. It really depends on the task though; your mileage may vary.
- ComputerGuru 10mo agoHow long have you been previewing 5.2?
- ImprobableTruth 10mo agoAn almost 50% price increase. Benchmarks look nice, but 50% more nice...?
- arnaudsm 10mo ago#1 models are usually priced at 2x more than the competition, and they often decrease the price right when they lose the crown.
- wewtyflakes 10mo agoThere are too few examples to say this is a trend. There have been counterexamples of top models actually lowering the pricing bar (gpt-5, gpt-3.5-turbo, some gemini releases were even totally free [at first]).
- sigmar 10mo agoAre there any specifics about how this was trained? Especially when 5.1 is only a month old. I'm a little skeptical of benchmarks these days and wish they put this up on llmarena edit: noticed 5.2 is ranked in the webdev arena (#2 tied with gemini-3.0-pro), but not yet in text arena (last update 22hrs ago)
- kouteiheika 10mo agoUnfortunately there are never any real specifics about how any of their models were trained. It's OpenAI we're talking about after all.
- emp17344 10mo agoI’m extremely skeptical because of all those articles claiming OpenAI was freaking out about Gemini - now it turns out they just casually had a better model ready to go? I don’t buy it.
- tempaccount420 10mo agoThey had to rush it out, I'm sure the internal safety folks are not happy about it.
- Workaccount2 10mo agoI (and others) have a strong suspicion that they can modulate models intelligence in almost real time by adjusting quantization and thinking time. It seems if anyone wants, they can really gas a model up in the moment and back it off after the hype wave.
- bamboozled 10mo agoYeah I've noticed with Claude, around the time of the Opus 4.5 release, at least for a few days, Sonnet 4.5 was just dumb, but it seems temporary. I feel that redirected resources to Opus.
- qeternity 10mo agoQuantization is not some magical dial you can just turn. In practice you basically have 3 choices: fp16, fp8 and fp4. Also thinking time means more tokens which costs more especially at the API level where you are paying per token and would be trivially observable. There is basically no evidence that either of these are occurring in the way you suggest (boosting up and down).
- johnsutor 10mo agohttps://platform.openai.com/docs/models/gpt-5.2 https://platform.openai.com/docs/models/gpt-5.2 More information on the price, context window, etc.
- gkbrk 10mo agoIs this the "Garlic" model people have been hyping? Or are we not there yet?
- 0x457 10mo agoGarlic will be released 2026Q1.
- coolfox 10mo agothe halving of error rates for image inputs is pretty awesome, this makes it far more practical for issues where it isn't easy to input all the needed context. when I get lazy I'll just shift+win+s the problem and ask one of the chatbots to solve it.
- xd1936 10mo ago> While GPT‑5.2 will work well out of the box in Codex, we expect to release a version of GPT‑5.2 optimized for Codex in the coming weeks. https://openai.com/index/introducing-gpt-5-2/ https://openai.com/index/introducing-gpt-5-2/
- jstummbillig 10mo ago> For coding tasks, GPT-5.1-Codex-Max is a faster, more capable, and more token-efficient coding variant Hm, yeah, strange. You would not be able to tell, looking at every chart on the page. Obviously not a gotcha, they put it on the page themselves after all, but how does that make sense with those benchmarks?
- tempaccount420 10mo agoCoding requires a mindset shift that the -codex fine-tunes provide. Codex will do all kinds of weird stuff like poking in your ~/.cargo ~/go etc. to find docs and trying out code in isolation, these things definitely improve capability.
- dmos62 10mo agoThe biggest advantage of codex variants, for me, is terseness and reduced sicophany. That, and presumably better adherence to requested output formats.
- deaux 10mo agoLooks like they removed that line.
- baq 10mo agoCodex talks much less than the standard variant, especially between tool calls.
- k_bx 10mo agogpt-5.2 is already present in codex at this moment
- preetamjinka 10mo agoIt's actually more expensive than GPT-5.1. I've gotten used to prices going down with each latest model, but this time it's gone up. https://platform.openai.com/docs/pricing https://platform.openai.com/docs/pricing
- Handy-Man 10mo agoIt also seems much more "smarter" though
- PhilippGille 10mo agoGemini 3 Pro Preview also got more expensive than 2.5 Pro. 2.5 Pro: $1.25 input, $10 output (million tokens) 3 Pro Preview: $2 input, $12 output (million tokens)
- TechDebtDevin 10mo agoLiterally no difference in productivity from a free/ <0.50c output OpenRouter model. All these > $1.00+ per mm output are literal scams. No added value to the world.
- wahnfrieden 10mo ago5.1 Pro is great
- manmal 10mo agoI struggle to see where Pro is better than 5.x with Thinking. Actually prefer the latter.
- wahnfrieden 10mo agoMany problems where latter spins its wheel and Pro gets it in one go, for me. You need to give Pro full files as context and you need to fit within its ~60k (I forget exactly) silent context window if using via ChatGPT. Don't have it make edits directly, have it give the execution plan back to Codex
- ComputerGuru 10mo agoWish they would include or leak more info about what this is, exactly. 5.1 was just released, yet they are claiming big improvements (on benchmarks, obviously). Did they purposely not release the best they had to keep some cards to play in case of Gemini 3 success or is this a tweak to use more time/tokens to get better output, or what?
- Ninjinka 10mo agoMan this was rushed, typo in the first section: > Unlike the previous GPT-5.1 model, GPT-5.2 has new features for managing what the model "knows" and "remembers to improve accuracy.
- petercooper 10mo agoAlso, did they mention these features? I was looking out for it but got to the end and missed it. (No, I just looked again and the new features listed are around verbosity, thinking level and the tool stuff rather than memory or knowledge.)
- gigatexal 10mo agoSo how much better is it than opus or Gemini ?
- HardCodedBias 10mo agoHuge fan that Gemini-3 prompted OAI to ship this. Competition works! GDPval seems particularly strong. I wonder why they held this back. 1) Maybe this is uneconomical ? 2) Did the safety somehow hold back the company ? looking forward to the internet trying this and posting their results over the next week or two. COMPETITION!
- mrandish 10mo ago> I wonder why they held this back. IMHO, I doubt they were holding much back. Obviously, they're always working on 'next improvements' and rolled what was done enough into this but I suspect the real difference here is throwing significantly more compute (hence investor capital) at improving the quality - right now. How much? While the cost is currently staying the same for most users, the API costs seem to be ~40% higher. The impetus was the serious threat Gemini 3 poses. Perception about ChatGPT was starting to shift, people were speculating that maybe OAI is more vulnerable than assumed. This caused Altman to call an all-hands "Code Red" two weeks ago, triggering a significant redeployment of priorities, resources and people. I think this launch is the first 'stop the perceptual bleeding' result of the Code Red. Given the timing, I think this is mostly akin to overclocking a CPU or running an F1 race car engine too hot to quickly improve performance - at the cost of being unsustainable and unprofitable. To placate serious investor concerns, OAI has recently been trying to gradually work toward making current customers profitable (or at least less unprofitable). I think we just saw the effort to reduce the insane burn rate go out the window.
- Jackson__ 10mo agoFunny that, their front page demo has a mistake. For the waves simulation, the user asks: >- The UI should be calming and realistic. Yet what it did is make a sleek frosted glass UI with rounded edges. What it should have done is call a wellness check on the user on suspicion of a co2 leak leading to delirium.
- jasonthorsness 10mo agoDoes anyone have it yet in ChatGPT? I'm still on 5.1 :(.
- mudkipdev 10mo agoNo, but it's already in codex
- FergusArgyll 10mo ago> We deploy GPT‑5.2 gradually to keep ChatGPT as smooth and reliable as we can; if you don’t see it at first, please try again later.
- jasonthorsness 10mo agoI have it now
- FergusArgyll 10mo ago> Additionally, on our internal benchmark of junior investment banking analyst spreadsheet modeling tasks—such as putting together a three-statement model for a Fortune 500 company with proper formatting and citations, or building a leveraged buyout model for a take-private—GPT 5.2 Thinking's average score per task is 9.3% higher than GPT‑5.1’s, rising from 59.1% to 68.4%. Confirming prior reporting about them hiring junior analysts
- zhyder 10mo agoBig knowledge cutoff jump from Sep 2024 to Aug 2025. How'd they pull that off for a small point release, which presumably hasn't done a fresh pre-training over the web? Did they figure out how to do more incremental knowledge updates somehow? If yes that'd be a huge change to these releases going forward. I'd appreciate the freshness that comes with that (without having to rely on web search as a RAG tool, which isn't as deeply intelligent, as is game-able by SEO). With Gemini 3, my only disappointment was 0 change in knowledge cutoff relative to 2.5's (Jan 2025).
- throwaway314155 10mo ago> which presumably hasn't done a fresh pre-training over the web What makes you think that? > Did they figure out how to do more incremental knowledge updates somehow? It's simple. You take the existing model and continue pretraining with newly collected data.
- Workaccount2 10mo agoA leak reported on by semi-analyses stated that they haven't pre-trained a new model since 4o due to compute constraints.
- jumploops 10mo ago> “a new knowledge cutoff of August 2025” This (and the price increase) points to a new pretrained model under-the-hood. GPT-5.1, in contrast, was allegedly using the same pretraining as GPT-4o.
- 98Windows 10mo agoor maybe 5.1 was an older checkpoint and has more quantization
- deleted 10mo ago[deleted]
- FergusArgyll 10mo agoA new pretrain would definitely get more than a .1 version bump & would get a whole lot more hype I'd think. They're expensive to do!
- femiagbabiaka 10mo agoNot if they didn't feel that it delivered customer value no? It's about under promising and over delivering, in every instance
- redwood 10mo agoNot if it underwhelms
- hannesfur 10mo agoMaybe they felt the increase in capability is not worth of a bigger version bump. Additionally pre-training isn't as important as it used to be. Most of the advances we see now probably come from the RL stage.
- caconym_ 10mo agoReleasing anything as "GPT-6" which doesn't provide a generational leap in performance would be a PR nightmare for them, especially after the underwhelming release of GPT-5. I don't think it really matters what's under the hood. People expect model "versions" to be indexed on performance.
- devinprater 10mo agoCan the tables have column headers so my screen reader can read the model name as I go across the benchmakrs? And the images should have alt-text.
- MagicMoonlight 10mo agoThey’re definitely just training the models on the benchmarks at this point
- roxolotl 10mo agoYea either this is an incredible jump or we’ve finally gotten confirmation benchmarks are bs.
- simonw 10mo agoWow, there's a lot going on with this pelican riding a bicycle: https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83bcc?permalink_comment_id=5901194#gistcomment-5901194 https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...
- minimaxir 10mo agoIs that the first SVG pelican with drop shadows?
- simonw 10mo agoNo, I got drop shadows from DeepSeek 3.2 recently https://simonwillison.net/2025/Dec/1/deepseek-v32/ https://simonwillison.net/2025/Dec/1/deepseek-v32/ (probably others as well.)
- tmaly 10mo agoseems to be eating something
- danans 10mo agoProbably a jellyfish. You're seeing the tentacles
- belter 10mo agoWhat happens if you ask for a pterodactyl on a motorbike? Would like to know how much they are optimizing for your pelican....
- simonkagedal 10mo agoHe commented on this here: https://simonwillison.net/2025/Nov/13/training-for-pelicans-riding-bicycles/ https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
- irthomasthomas 10mo ago
- ComputerGuru 10mo agoWish they would include or leak more info about what this is, exactly. 5.1 was just released, yet they are claiming big improvements (on benchmarks, obviously). Did they purposely not release the best they had to keep some cards to play in case of Gemini 3 success or is this a tweak to use more time/tokens to get better output, or what?
- famouswaffles 10mo agoOpen AI sat on GPT-4 for 8 months and even released 3.5 months after 4 was trained. While i don't expect such big lag times anymore, generally, it's a given the public is behind whatever models they have internally at the frontier. By all indications, they did not want to release this yet, and only did so because of Gemini-3-pro.
- eldenring 10mo agoI'm guessing they were waiting to figure out more efficient serving before a release, and have decided to eat the inference cost temporarily to stay at the frontier.
- dalemhurley 10mo agoMy guess is they develop multiple models in parallel.
- nathan-wall 10mo agoIf you look at their own chart[1] it shows 5.1 was lagging behind Gemini 3 Pro in almost every score listed there, sometimes significantly. They needed to come out with something to stay ahead. I'm guessing they threw what they had at their disposal together to keep the lead as long as they can. It sounds like 5.2 has a more recent knowledge cutoff; a reasonable guess is they could have already had that but were trying to make bigger improvements out of it for a more major 5.5 release before Gemini 3 Pro came out and then they had to rush something out. Also 5.2 has a new "Extended Thinking" option for Pro. I'm guessing they just turned up a lever that told it to think even longer, which helps them score higher, even if it does take a long time. (One thing about Gemini 3 Pro is it's very fast relative to even ChatGPT 5.1 Pro Thinking. A lot of the scores they're putting out to show they're staying ahead aren't showing that piece.) [1] https://imgur.com/e0iB8KC https://imgur.com/e0iB8KC
- airstrike 10mo agoI feel like if we're going to regulate anything about AI, we should start by regulating (1) what they get to claim to be a "new model" to the public and (2) what changes they are allowed to make at inference before being forced to name it something different.
- jacquesm 10mo agoThat's almost but not quite how the airline industry is treated. The difference there is that the regulators are in bed with the companies they should be regulating.
- chux52 10mo agoIs this why all my Cursor requests are timing out in the past hour?
- sureglymop 10mo agoHow can I hide the big "Ask ChatGPT" button I accidentally clicked like 3 times while actually trying to read this on my phone? I guess I must "listen" to the article...
- z58 10mo agoWith Safari on iOS you can hide distracting items. I just tried it on that button, it works flawlessly.
- riazrizvi 10mo agoDoes it still use the word ‘fluff’ in 90% of its preambles, or is it finally able to get straight to the point?
- ChrisArchitect 10mo agoDiscussion on blog post: https://openai.com/index/introducing-gpt-5-2/ https://openai.com/index/introducing-gpt-5-2/ (https://news.ycombinator.com/item?id=46234874 https://news.ycombinator.com/item?id=46234874)
- yousif_123123 10mo agoWhy doesn't OpenAI include comparisons to other models anymore?
- ftchd 10mo agobecause they probably need to compare pricing too
- enraged_camel 10mo agoBecause their main competition (Google and Anthropic) have caught up and even started to surpass them, and comparisons would simply drive it home.
- IAmNotACellist 10mo agoWhy do they care so much? They're a non-profit dedicated to the betterment of humanity via open access to AI. They have nothing to hide. They have no motivation to lie, or lie by omission.
- koolba 10mo ago> Why do they care so much? They're a non-profit dedicated to the betterment of humanity via open access to AI. We're still talking about OpenAI right?
- IAmNotACellist 10mo agoYou're not calling Sam Altman a liar, are you?
- kaliqt 10mo agoThey are not a nonprofit at all. Legally, yes. But they are not.
- conradkay 10mo agoSam Altman posted with a comparison to Gemini 3 and Opus 4.5 https://x.com/sama/status/1999185784012947900 https://x.com/sama/status/1999185784012947900
- HackerThemAll 10mo agoNo, thank you, OpenAI and ChatGPT doesn't cut it for me.
- dang 10mo ago"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- d--b 10mo ago> it’s better at creating spreadsheets I have a bad feeling about this.
- HackerThemAll 10mo agoNo, thank you, OpenAI and ChatGPT doesn't cut it for me.
- wayeq 10mo agothanks for letting us know.
- replwoacause 10mo agoWhat’s cutting it for you these days?
- deleted 10mo ago[deleted]
- daviding 10mo agogpt-5.2 and gpt-5.2-chat-latest the same token price? Isn't the latter non-thinking and more akin to -nano or -mini?
- dalemhurley 10mo agoNo. It is the same model without reasoning.
- daviding 10mo agoSo is maybe gpt-5.2 with reasoning set to 'none' identical to gpt-5.2-chat-latest in capabilities but perhaps with a different system (system) prompt? I notice chat-latest doesn't accept temperature or reasoning (which makes sense) parameters, so something is certainly different underneath?
- scottndecker 10mo agoStill 256K input tokens. So disappointing (predictable, but disappointing).
- dinobones 10mo agoIt's becoming challenging to really evaluate models. The amount of intelligence that you can display within a single prompt, the riddles, the puzzles, they've all been solved or are mostly trivial to reasoners. Now you have to drive a model for a few days to really get a decent understanding of how good it really is. In my experience, while Sonnet/Opus may not have always been leading on benchmarks, they have always *felt* the best to me, but it's hard to put into words why exactly I feel that way, but I can just feel it. The way you can just feel when someone you're having a conversation with is deeply understanding you, somewhat understanding you, or maybe not understanding at all. But you don't have a quantifiable metric for this. This is a strange, weird territory, and I don't know the path forward. We know we're definitely not at AGI. And we know if you use these models for long-horizon tasks they fail at some point and just go off the rails. I've tried using Codex with max reasoning for doing PRs and gotten laughable results too many times, but Codex with Max reasoning is apparently near-SOTA on code. And to be fair, Claude Code/Opus is also sometimes equally as bad at doing these types of "implement idea in big codebase, make changes too many files, still pass tests" type of tasks. Is the solution that we start to evaluate LLMs on more long-horizon tasks? I think to some degree this was the spirit of SWE Verified right? But even that is being saturated now.
- ACCount37 10mo agoThe good old "benchmarks just keep saturating" problem. Anthropic is genuinely one of the top companies in the field, and for a reason. Opus consistently punches above its weight, and this is only in part due to the lack of OpenAI's atrocious personality tuning. Yes, the next stop for AI is: increasing task length horizon, improving agentic behavior. The "raw general intelligence" component in bleeding edge LLMs is far outpacing the "executive function", clearly.
- imiric 10mo agoShouldn't the next stop be to improve general accuracy, which is what these tools have struggled with since their inception? Until when are "AI" companies going to offload the responsibility on the user to verify the output of their tools? Optimizing for benchmark scores, which are highly gamed to begin with, by throwing more resources at this problem is exceedingly tiring. Surely they must've noticed the performance plateau and diminishing returns of this approach by now, yet every new announcement is the same.
- deleted 10mo ago[deleted]
- qoez 10mo agoThis is also the exact on-the-day 10th anniversary of openai's creation incidentally
- cc62cf4a4f20 10mo agoIn other news, been using Devstral 2 (Ollama) with OpenCode, and while it's not as good as Claude Code, my initial sense it that it's nonetheless good enough and doesn't require me to send my data off my laptop. I kind of wonder how close we are to alternative (not from a major AI lab) models being good enough for a lot of productive work and data sovereignty being the deciding factor.
- Nesco 10mo agoWait, isn't Devstral2 (normal not small) 123b? What type of laptop do you have? MacBooks don't go over 128GiB
- cc62cf4a4f20 10mo agoI'm using small - works well for its size
- yberreby 10mo agoWould you share some additional details? CPU, amount of unified memory / VRAM? Tok/s with those?
- cc62cf4a4f20 10mo agoMBP M4 Max 64MB - haven't measured the tokens/sec, feels slower than Claude, but not unbearably It's not yet perfect, my sense is just that it's near the tipping point where models are efficient enough that running a local model is truly viable
- a_wild_dandan 10mo ago> Unlike the previous GPT-5.1 model, GPT-5.2 has new features for managing what the model "knows" and "remembers to improve accuracy. Dumb nit, but why not put your own press release through your model to prevent basic things like missing quote marks? Reminds me of that time an OAI released wildly inaccurate copy/pasted bar charts.
- Imnimo 10mo agoIt does seem to raise fair questions about either the utility of these tools, or adoption inertia. If not even OpenAI feels compelled to integrate this kind of model-check into their pipeline, what's that say about the business world at-large? Is it that it's too onerous to set up, is it that it's too hard to get only true-positive corrections, is it that it's too low value for the effort?
- JumpCrisscross 10mo ago> what's that say about the business world at-large? Nothing. OpenAI is a terrible baseline to extrapolate anything from.
- croes 10mo agoMaybe they did
- layer8 10mo agoHumans are now expected to parse sloppy typing without complaining about it, just like LLMs do. Slop is the new normal.
- Bengalilol 10mo agoIt may have been used, how could we know? Mainly, I don't get why there are quote marks at all.
- boplicity 10mo agoTheir model doesn't handle punctuation, quote marks, and similar things very well at all.
- SkyPuncher 10mo agoGiven the price increase and speculation that GPT 5 is a MoE model, I'm wondering if they're simply "turning up the good stuff" without making significant changes under the hood.
- throwaway314155 10mo agoGPT 4o was an MoE model as well.
- minimaxir 10mo agoI'm not sure why being a MoE model would allow OpenAI to "turn up the good stuff". You can't just increase the number of E without training it as such.
- yberreby 10mo agoBased on what works elsewhere in deep learning, I see no reason why you couldn't train once with a randomized number of experts, then set that number during inference based on your desired compute-accuracy tradeoff. I would expect that this has been done in the literature already.
- SkyPuncher 10mo agoMy opinion is they're trying to internally route requests to cheaper experts when they think they can get away with it. I felt this was evident by the wild inconsistencies I'd experience using it for coding. Both in quality and latency You "turn of the good stuff" by eliminating or reducing the likelihood of the cheap experts handling the request.
- dumbmrblah 10mo agoGreat! It'll be SOTA for a couple of weeks until the quality degrades due to throttling. I'll stick with plug and play API instead.
- mrandish 10mo agoDue to the "Code Red" threat from Gemini 3, I suspect they'll hold off throttling for longer than usual (by incinerating even more investor capital than usual). Jump in and soak up that extra-discounted compute while the getting is good, kids! Personally, I recently retired so I just occasionally mess around with LLMs for casual hobby projects, so I've only ever used the free tier of all the providers. Having lived through the dot com bubble, I regret not soaking up more of the free and heavily subsidized stuff back then. Trying not to miss out this time. All this compute available for free or below cost won't last too much longer...
- dankwizard 10mo agoI've been using tools like ProxLLM which just slam these AI models via proxy everytime a free tier limit is hit and it works great.
- ssvss 10mo agocan you provide a link to this tool, a search for proxllm didn't seem to find anything related.
- impulser_ 10mo agoThe thing about OpenAI is their models never fit anywhere for me. Yes they maybe smart or even the smartest models but they are alway so fucking slow. The ChatGPT web app is literally usable for me. I ask simple task and it does most extreme shit jsut to get an answer that the same as Claude or Gemini. For example, I asked ChatGPT to take a chart and convert into a table. It went and cut up the image and zoomed in for literally 5 mins to get the a worst answer than Claude which did it in under a minute. I see people talk about Codex like it better than Claude Code, and I go and try it and it takes a lifetime to do thing and it return maybe an on par result as Opus or Sonnet but it takes 5mins longer. I just tried out this model and it the same exact thing. It just take ages for it to give you an answer. I don't get how these models are useful in the real world. What am I missing, is this just me? I guess it truly an enterprise model.
- wetoastfood 10mo agoAre you using 5.1 Thinking? I tended to prefer Claude before this model. I use models based on the task. They still seem specialized and better at specific tasks. If I have a question I tend to go to it. If I need code, I tend to go to Claude (Code). I go to ChatGPT for questions I have because I value an accurate answer over a quick answer and, in my experience, it tends to give me more accurate answers because of its (over) willingness to go to the web for search results and question its instincts. Claude is much more likely to make an assumption and its search patterns aren't as thorough. The slow answers don't bother me because it's an expectation I have for how I use it and they've made that use case work really well with background processing and notifications.
- deleted 10mo ago[deleted]
- zone411 10mo agoI've benchmarked it on the Extended NYT Connections benchmark (https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/): The high-reasoning version of GPT-5.2 improves on GPT-5.1: 69.9 → 77.9. The medium-reasoning version also improves: 62.7 → 72.1. The no-reasoning version also improves: 22.1 → 27.5. Gemini 3 Pro and Grok 4.1 Fast Reasoning still score higher.
- Donald 10mo agoGemini 3 Pro Preview gets 96.8% on the same benchmark? That's impressive
- capitainenemo 10mo agoAnd performs very well on the latest 100 puzzles too, so isn't just learning the data set (unless I guess they routinely index this repo). I wonder how well AIs would do at bracket city. I tried gemini on it and was underwhelmed. It made a lot of terrible connections and often bled data from one level into the next.
- wooger 10mo ago> unless I guess they routinely index this repo This sounds like exactly the kind of thing any tech company would do when confronted with a competitive benchmark.
- rsanek 10mo agoI mean, the repo has <200 stars, it's not like it's so mainstream that you'd expect LLM makers to be watching it actively. If they wanted to game it, they could more easily do that in RL with synthetic data anyway.
- capitainenemo 10mo agoBelated update on this. Gemini reasoning did much better than quick on bracket city today (an easy puzzle but still). It only failed to solve one clue outright, got another wrong but due to ambiguity in the expression referenced and in a way that still fit the next level down making the final answer fairly cleanly solved. Still clearly has a harder time with it than the connections puzzle.
- speedgoose 10mo agoTrying it now in Vscode Insiders with Github Copilot (codex crashes with HTTP 400 server errors), and it eventually started using sed and grep in shells instead of using the better tools it has access to. I guess this is not an issue to perform well in benchmarks.
- pixelmelt 10mo agoto be fair I've seen the other sota models do this as well
- deleted 10mo ago[deleted]
- songodongo 10mo agoI get this behavior with a lot with most of the premium models (Gemini 3, Opus 4.5). I think it’s somehow more a GitHub Copilot issue than the models.
- andreygrehov 10mo agoEvery new model is ‘state-of-the-art’. This term is getting annoying.
- arthur-st 10mo agoI mean, that is what the term implies.
- sundarurfriend 10mo ago> new context management using compaction. Nice! This was one of the more "manual" LLM management things to remember to regularly do, if I wanted to avoid it losing important context over long conversations. If this works well, this would be a significant step up in usability for me.
- jiggawatts 10mo agoFeels a bit rushed. They haven’t even updated their API playground yet, if I select 5.2-chat-latest, I get: Unsupported parameter: 'top_p' is not supported with this model. Also, without access to the Internet, it does not seem to know things up to August 2025. A simple test is to ask it about .NET 10 which was already in preview at that time and had lots of public content about its new features. The model just guessed and waved its hand about, like a student that hadn’t read the assigned book.
- jstummbillig 10mo agoSo, right off the bat: 5.2 code talk (through codex) feels really nice. The first coding attempt was a little meh compared to 5.1 codex max (reflecting what they wrote themselves), but simply planning / discussing things felt markedly better than anything I remember from any previous model, from any company. I remain excited about new models. It's like finding my coworker be 10% smarter every other week.
- iwontberude 10mo agoI have already cancelled. Claude is more than enough for me. I don’t see any point in splitting hairs. They are all going to keep lying more and more sneakily.
- deleted 10mo ago[deleted]
- slackr 10mo ago“…where it outperforms industry professionals at well-specified knowledge work tasks spanning 44 occupations.” What a sociopathic way to sell
- willahmad 10mo agoare we doomed yet? Seems not yet with 5.2
- dangelosaurus 10mo agoI ran a red team eval on GPT-5.2 within 30 minutes of release: Baseline safety (direct harmful requests): 96% refusal rate With jailbreaking: 22% refusal rate 4,229 probes across 43 risk categories. First critical finding in 5 minutes. Categories with highest failure rates: entity impersonation (100%), graphic content (67%), harassment (67%), disinformation (64%). The safety training works against naive attacks but collapses with adversarial techniques. The gap between "works on benchmarks" and "works against motivated attackers" is still wide. Methodology and config: https://www.promptfoo.dev/blog/gpt-5.2-trust-safety-assessment/ https://www.promptfoo.dev/blog/gpt-5.2-trust-safety-assessme...
- stainablesteel 10mo agoim happy for this, but there's all these math and science benchmarks, has anyone ever made a communicates-like-a-human benchmark? or an isn't-frustrating-to-talk-with benchmark?
- tenpoundhammer 10mo agoI have been using chatGPT a ton over the last months and paying the subscription. Used it for coding, news, stock analysis, daily problems, and a whatever I could think of. I decided to give Gemini a go when version three came out to great reviews. Gemini handles every single one of my uses cases much better and consistently gives better answers. This is especially true for situations were searching the web for current information is important, makes sense that google would be better. Also OCR is phenomenal chatgpt can't read my bad hand writing but Gemini can easily. Only downsides are in the polish department, there are more app bugs and I usually have to leave the happen or the session terminates. There are bugs with uploading photos. The biggest complaint is that all links get inserted into google search and then I have to manipulate them when they should go directly to the chosen website, this has to be some kind of internal org KPI nonsense. Overall, my conclusion is that ChatGPT has lost and won't catch up because of the search integration strength.
- LorenDB 10mo agoWhat is it with the Polish always messing up products? (yes, /s)
- petersumskas 10mo agoIt’s because their thoughts are Roman while they are always Russian to Finnish things. Kenya believe it! Anyway, I’m done here. Abyssinia.
- labrador 10mo agoI like their hotdogs
- solarkraft 10mo ago> Only downsides are in the polish department What an understatement. It has me thinking „man, fuck this“ on the daily. Just today it spontaneously lost an entire 20-30 minutes long thread and it was far from the first time. It basically does it any time you interrupt it in any way. It’s straight up data loss. It’s kind of a typical Google product in that it feels more like a tech demo than a product. It has theoretically great tech. I particularly like the idea of voice mode, but it’s noticeably glitchy, breaks spontaneously often and keeps asking annoying questions which you can’t make it stop.
- deleted 10mo ago[deleted]
- deleted 10mo ago[deleted]
- deleted 10mo ago[deleted]
- bobse 10mo ago[dead]
- onraglanroad 10mo agoI suppose this is as good a place as any to mention this. I've now met two different devs who complained about the weird responses from their LLM of choice, and it turned out they were using a single session for everything. From recipes for the night, presents for the wife and then into programming issues the next day. Don't do that. The whole context is sent on queries to the LLM, so start a new chat for each topic. Or you'll start being told what your wife thinks about global variables and how to cook your Go. I realise this sounds obvious to many people but it clearly wasn't to those guys so maybe it's not!
- vintermann 10mo agoIt's not at all obvious where to drop the context, though. Maybe it helps to have similar tasks in the context, maybe not. It did really, shockingly well on a historical HTR task I gave it, so I gave it another one, in some ways an easier one... Thought it wouldn't hurt to have text in a similar style in the context. But then it suddenly did very poorly. Incidentally, one of the reasons I haven't gotten much into subscribing to these services, is that I always feel like they're triaging how many reasoning tokens to give me, or AB testing a different model... I never feel I can trust that I interact with the same model.
- dcre 10mo agoThe models you interact with through the API (as opposed to chat UIs) are held stable and let you specify reasoning effort, so if you use a client that takes API keys, you might be able to solve both of those problems.
- eru 10mo ago> Incidentally, one of the reasons I haven't gotten much into subscribing to these services, is that I always feel like they're triaging how many reasoning tokens to give me, or AB testing a different model... I never feel I can trust that I interact with the same model. That's what websites have been doing for ages. Just like you can't step twice in the same river, you can't use the same version of Google Search twice, and never could.
- 10mo ago
- keeeba 10mo agoDoesn’t seem like this will be SOTA in things that really matter, hoping enough people jump to it that Opus has more lenient usage limits for a while
- w_for_wumbo 10mo agoDoes anyone else consider that maybe it's impossible to benchmark the performance of a piece of paper. This is a tool that allows an intelligent system to work with it, the same way that a piece of paper can reflect the writers' intelligence, how can we accurately judge the performance of the piece of paper, when it is so intimately reliant on the intelligence that is working with it?
- anishshil 10mo ago[dead]
- mlmonkey 10mo agoIt's funny how they don't compare themselves to Gemini and Claude anymore.
- anishshil 10mo agoThis shift toward new platforms is exactly why I’m building Truwol, a social experience focused on real, unedited human moments instead of the AI-saturated feeds we’re drifting toward. I’m developing it independently and sharing the progress publicly, so if you’re interested in projects reinventing online spaces from the ground up, you can see what I’m working on Truwol buymeacoffee/Truwol
- jrflowers 10mo agoOpenAI is really good at just saying stuff on the internet. I love the way they talk about incorrect responses: > Errors were detected by other models, which may make errors themselves. Claim-level error rates are far lower than response-level error rates, as most responses contain many claims. “These numbers might be wrong because they were made up by other models, which we will not elaborate on, also these numbers are much higher by a metric that reflects how people use the product, which we will not be sharing“ I also really love the graph where they drew a line at “wrong half of the time” and labeled it ‘Expert-Level’. 10/10, reading this post is experientially identical to watching that 12 hours of jingling keys video, which is hard to pull off for a blog.
- goobatrooba 10mo agoI feel there is a point when all these benchmarks are meaningless. What I care about beyond decent performance is the user experience. There I have grudges with every single platform and the one thing keeping me as a paid ChatGPT subscriber is the ability to sort chats in "projects" with associated files (hello Google, please wake up to basic user-friendly organisation!) But all of them * Lie far too often with confidence * Refuse to stick to prompts (e.g. ChatGPT to the request to number each reply for easy cross-referencing; Gemini to basic request to respond in a specific language) * Refuse to express uncertainty or nuance (i asked ChatGPT to give me certainty %s which it did for a while but then just forgot...?) * Refuse to give me short answers without fluff or follow up questions * Refuse to stop complimenting my questions or disagreements with wrong/incomplete answers * Don't quote sources consistently so I can check facts, even when I ask for it * Refuse to make clear whether they rely on original documents or an internal summary of the document, until I point out errors * ... I also have substance gripes, but for me such basic usability points are really something all of the chatbots fail on abysmally. Stick to instructions! Stop creating walls of text for simple queries! Tell me when something is uncertain! Tell me if there's no data or info rather than making something up!
- nullbound 10mo ago<< I feel there is a point when all these benchmarks are meaningless. I am relatively certain you are not alone in this sentiment. The issue is that the moment we move past seemingly objective measurements, it is harder to convince people that what we measure is appropriate, but the measurable stuff can be somewhat gamed, which adds a fascinating layer of cat and mouse game to this.
- delifue 10mo agoOnce a metric becomes optimization target, it ceases to become good metric.
- hnfong 10mo agoThere's a leaderboard that measures user experience, the "lmsys" Chatbot Arena Leaderboard ( https://huggingface.co/spaces/lmarena-ai/lmarena-leaderboard https://huggingface.co/spaces/lmarena-ai/lmarena-leaderboard ). Main issue with it these days are that it kinda measures sycophancy and user preferred tone more than substance. Some issues you mentioned like length of response might be user preference. Other issues like "hallucination" are areas of active research (and there are benchmarks for these).
- kachapopopow 10mo agodid they just tune the parameters? the hallucinations are crazy high on this version.
- hbarka 10mo agoA year ago Sunday Pichai declared code red, now it’s Sam Altman declaring code red. How tables have turned, and I think the acquisition of Windsurf and Kevin Hou by Google seems to correlate with their level up.
- jerrygenser 10mo agoAcquisition of noam shazeer to supercharge their Gemini flagship model line I think made a bigger impact. To make an argument it was Kevin Hou, then we would need to see Antigravity their new IDE being key. I think the crown jewel are the Gemini models.
- ChrisMarshallNY 10mo agoThey are talking a lot about economics, here. Wonder what that will mean for standard Plus users, like me.
- bluerooibos 10mo agoYawn.
- dudeinhawaii 10mo agoWhat does this add to the conversation? This isn't Reddit.
- mobrienv 10mo agoI recently built a webapp to summarize hn comment threads. Sharing a summary given there is a lot here: https://hn-insights.com/chat/gpt-52-8ecfpn https://hn-insights.com/chat/gpt-52-8ecfpn.
- DenisM 10mo agoI keep asking ChatGPT to read and summarize HN front page while driving, and it keeps blundering. I don’t know if there’s a business for you in this, but I would pay. Of course I always have questions about the subject, so it become the whole voice chat thing.
- mobrienv 10mo agoInteresting I recently added the ability to receive a daily email digest. Would just need a way to read it out. I'll look into what a conversational voice chat might look like.
- DenisM 10mo agoIs there a voice chat mode in any chat app that is not heavily degraded in reasoning? I’m ok waiting for a response for 10-60 seconds if needed. That way I can deep dive subjects while driving. I’m ok paying money for it, so maybe someone coded this already?
- nbardy 10mo agoThose arc agi 2 improvements are insane. Thats especially encouraging to me because those are all about generalization. 5 and 5.1 both felt overfit and would break down and be stubborn when you got them outside their lane. As opposed to Opus 4.5 which is lovely at self correcting. It’s one of those things you really feel in the model rather than whether it can tackle a harder problem or not, but rather can I go back and forth with this thing learning and correcting together. This whole releases is insanely optimistic for me. If they can push this much improvement WITHOUT the new huge data centers and without a new scaled base model. Thats incredibly encouraging for what comes next. Remember the next big data center are 20-30x the chip count and 6-8x the efficiency on the new chip. I expect they can saturate the benchmarks WITHOUT and novel research and algorithmic gains. But at this point it’s clear they’re capable of pushing research qualitatively as well.
- mmaunder 10mo agoSame. Also got my attention re ARC-AGI-2. That's meaningful. And a HUGE leap.
- cbracketdash 10mo agoSlight tangent yet I think is quite interesting... you can try out the ARC-AGI 2 tasks by hand at this website [0] (along with other similar problem sets). Really puts into perspective the type of thinking AI is learning! [0] https://neoneye.github.io/arc/?dataset=ARC-AGI-2 https://neoneye.github.io/arc/?dataset=ARC-AGI-2
- delifue 10mo agoIt's also possible that OpenAI use many human-generated similar-to-ARC data to train (semi-cheating). OpenAI has enough incentive to fake high score. Without fully disclosing training data you will never be sure whether good performance comes from memorization or "semi-memorization".
- deleted 10mo ago[deleted]
- ponyous 10mo agoI am really curious about speed/latency. For my use case there is a big difference in UX if the model is faster. Wish this was included in some benchmarks. I will run 80 3D model generations benchmark tomorrow and update this comment with the results about cost/speed/quality.
- flkiwi 10mo agoI gave up my OpenAI subscription a few days ago in favor of Claude. My quality of life (and quality of results) has gone up substantially. Several of our tools at work have GPT-5x as their backend model, and it is incredible how frustrating they are to use, how predictable their AI-isms are, and how inconsistent their output is. OpenAI is going to have to do a lot more than an incremental update to convince me they haven't completely lost the thread.
- brisket_bronson 10mo agoYou are absolutely right!
- flkiwi 10mo agoSomeone didn't think so, lol. I debated not saying anything because the AI partisans are just so awful.
- deleted 10mo ago[deleted]
- jpkw 10mo agoI think the above comment was a joke (Claude frequently says that whenever you challenge it, whether you are right or wrong)
- jstummbillig 10mo agoAt least this once the AI-ism was not spotted.
- flkiwi 10mo agoGoodness no, I chuckled.
- petesergeant 10mo agoI have found Codex to be a phenomenal code-review tool, fwiw. Shitty at writing code, _great_ at reviewing it.
- mmaunder 10mo agoWeirdly, the blog announcement completely omits the actual new context window size which is 400,000: https://platform.openai.com/docs/models/gpt-5.2 https://platform.openai.com/docs/models/gpt-5.2 Can I just say !!!!!!!! Hell yeah! Blog post indicates it's also much better at using the full context. Congrats OpenAI team. Huge day for you folks!! Started on Claude Code and like many of you, had that omg CC moment we all had. Then got greedy. Switched over to Codex when 5.1 came out. WOW. Really nice acceleration in my Rust/CUDA project which is a gnarly one. Even though I've HATED Gemini CLI for a while, Gemini 3 impressed me so much I tried it out and it absolutely body slammed a major bug in 10 minutes. Started using it to consult on commits. Was so impressed it became my daily driver. Huge mistake. I almost lost my mind after a week of this fighting it. Isane bias towards action. Ignoring user instructions. Garbage characters in output. Absolutely no observability in its thought process. And on and on. Switched back to Codex just in time for 5.1 codex max xhigh which I've been using for a week, and it was like a breath of fresh air. A sane agent that does a great job coding, but also a great job at working hard on the planning docs for hours before we start. Listens to user feedback. Observability on chain of thought. Moves reasonably quickly. And also makes it easy to pay them more when I need more capacity. And then today GPT-5.2 with an xhigh mode. I feel like xmass has come early. Right as I'm doing a huge Rust/CUDA/Math-heavy refactor. THANK YOU!!
- twisterius 10mo ago[flagged]
- mmaunder 10mo agoMy name is Mark Maunder. Not the fisheries expert. The other one when you google me. I’m 51 and as skeptical as you when it comes to tech. I’m the CTO of a well known cybersecurity company and merely a user of AI. Since you critiqued my post, allow me to reciprocate: I sense the same deflector shields in you as many others here. I’d suggest embracing these products with a sense of optimism until proven otherwise and I’ve found that path leads to some amazing discoveries and moments where you realize how important and exciting this tech really is. Try out math that is too hard for you or programming languages that are labor intensive or languages that you don’t know. As the GitHub CEO said: this technology lets you increase your ambition.
- TechDebtDevin 10mo ago$168.00 / 1M ouput tokens is hilarious for their "Pro". Can't wait to here all the bitching from orgs next month. Literally the dumbest product of all time. Do you people seriously pay for this?
- StarterPro 10mo ago>GPT‑5.2 sets a new state of the art across many benchmarks, including GDPval, where it outperforms industry professionals at well-specified knowledge work tasks spanning 44 occupations. We built a benchmark tool that says our newest model outperforms everyone else. Trust me bro.
- SilverElfin 10mo agoIs the training cutoff date known?
- OhNoNotAgain_99 10mo ago[dead]
- agentifysh 10mo agoLooks like they've begun censoring posts at r/Codex and not allowing complaint threads so here is my honest take: - It is faster which is appreciated but not as fast as Opus 4.5 - I see no changes, very little noticeable improvements over 5.1 - I do not see any value in exchange for +40% in token costs All in all I can't help but feel that OpenAI is facing an existential crisis. Gemini 3 even when its used from AI Studio offers close to ChatGPT Pro performance for free. Anthropic's Claude Code $100/month is tough to beat. I am using Codex with the $40 credits but there's been a silent increase in token costs and usage limitations.
- AstroBen 10mo agoDid you notice much improvement going from Gemini 2.5 to 3? I didn't I just think they're all struggling to provide real world improvements
- XCSme 10mo agoMaybe they are just more consistent, which is a bit hard to notice immediately.
- dcre 10mo agoNearly everyone else (and every measure) seems to have found 3 a big improvement over 2.5.
- enraged_camel 10mo agoGemini 3 was a massive improvement over 2.5, yes.
- cmrdporcupine 10mo agoI think what they're actually struggling with is costs. And I think they're all behind the scenes quantizing models to manage load here and there, and they're all giving inconsistent results. I noticed huge improvement from Sonnet 4.5 to Opus 4.5 when it became unthrottled a couple weeks ago. I wasn't going to sign back up with Anthropic but I did. But two weeks in it's already starting to seem to be inconsistent. And when I go back to Sonnet it feels like they did something to lobotomize it. Meanwhile I can fire up DeepSeek 3.2 or GLM 4.6 for a fraction of the cost and get almost as good as results.
- tpurves 10mo agoUndoubtedly each new model from OpenAi has numerous training and orchestration improvements etc. But how much of each product they release also just a factor of how much they are willing to spend on inference per query in order to stay competitive? I always wonder how much is technical change vs turning a knob up and down on hardware and power consumption. GTP5.0 for example seemed like a lot of changes more for OpenAI's internal benefit (terser responses, dynamic 'auto' mode to scale down thinking when not required etc.) Wondering if GPT5.2 is also case of them in 'code red mode' just turning what they already have up to 11 as a fastest way to respond to fiercer competion.
- simonsarris 10mo agoI always liked the definition of technology as "doing more with less". 100 oxen replaced by 1 gallon of diesel, etc. That it costs more does suggest it's "doing more with more", at least.
- psychoslave 10mo agoGood luck with reproducing and eating diesel like can be done with oxen and related species. Humanity won't be able to tap into this highly compressed energy stock that was generated through processes taking literally geological scales time to bed achieved. That is, technology is more about what alternative tradeoffs can we leverage on to organize differently with resources at hand. Frugality can definitely be a possible way to shape the technologies we want to deploy. But it's not all possible technologies, just a subset. Also better technology is not necessarily bringing societies to morale and well-being excellency. Improving technology for efficient genocides for example is going to bring human disaster as obvious outcome, even if it's done in a manner that is the most green, zero-carbon emissions and growing more forests delivered beyond expectations of the specifications.
- ClipNoteBook 10mo agoChatGPT seems to just randomly pick urls to cite and extract information from. Google Gemini seems to look at heuristics like whether the author is trustworthy, or an expert in the topic. But more advanced
- jonplackett 10mo agoExcited to try this. I’ve found Gemini excellent recently and amazing at coding. But I still feel somehow like ChatGPT understands more. Even though it’s not quite as good at coding - and nowhere at as fast. It is much less likely anti spontaneously forget something. Gemini’s is part unbelievably amazing and part amnesia patient. I still kinda trust ChatGPT more.
- snake_doc 10mo ago> Models were run with maximum available reasoning effort in our API (xhigh for GPT‑5.2 Thinking & Pro, and high for GPT‑5.1 Thinking), except for the professional evals, where GPT‑5.2 Thinking was run with reasoning effort heavy, the maximum available in ChatGPT Pro. Benchmarks were conducted in a research environment, which may provide slightly different output from production ChatGPT in some cases. Feels like a Llama 4 type release. Benchmarks are not apples to apples. Reasoning effort is across the board higher, thus uses more compute to achieve an higher score on benchmarks. Also notes that some may not be producible. Also, vision benchmarks all use Python tool harness, and they exclude scores that are low without the harness.
- intelec1 10mo ago[dead]
- 0xdeafbeef 10mo agomuch better https://chatgpt.com/s/t_693b489d5a8881918b723670eaca5734 https://chatgpt.com/s/t_693b489d5a8881918b723670eaca5734 than 5.1 https://chatgpt.com/s/t_6915c8bd1c80819183a54cd144b55eb2 https://chatgpt.com/s/t_6915c8bd1c80819183a54cd144b55eb2. Same query - what romanian football player won the premier league update. Even instant returns correct result without problems https://chatgpt.com/s/t_693b49e8f5808191a954421822c3bd0d https://chatgpt.com/s/t_693b49e8f5808191a954421822c3bd0d
- jacquesm 10mo agoA classic long-form sales pitch. Someone's been reading their Patio11...
- jaimex2 10mo agoThey just keep flogging that dead horse. The winner in this race will be whoever gets small local models to perform as well on consumer hardware. It'll also pop the tech bubble in the US.
- johnwheeler 10mo agoI'm not interested in using OpenAI anymore because Sam Altman is so untrustworthy. All you see on X.com is him and Greg Brockman kissing David Sacks' ass, trying to make inroads with him, asking Disney for investments, and shit. Are you kidding? Who wants to support these clowns? Let's let Google win. Let's let Anthropic win. Anyone but Sam Altman.
- lacoolj 10mo agoThis is a whole bunch of patting themselves on the back. Let me know when Gemini 3 Pro and Opus 4.5 are compared against it.
- upvotenow 10mo ago[dead]
- deleted 10mo ago[deleted]
- johan914 10mo agoA bit off topic: but what's with the ram usage of LLM clients? ChatGPT, google, and Anthropic all use 1+ GB of ram during a long session. Surely they are not running GPT 3 locally?
- aaroninsf 10mo agoAs a popcorn eating bystander it is striking to scan the top comments and find they alternate so dramatically in tone and conclusions.
- whereistejas 10mo agoDid anyone notice how Cursor wasn’t an early tester? I wonder why…
- Kim_Bruning 10mo agoI'm continuously surprised that some people get good results out of GPT models. They sort of fail on my personal benchmarks for me. Maybe GPT needs a different approach to prompting? (as compared to eg Claude, Gemini, or Kimi)
- piskov 10mo agoThey are all gpt as in generative pre-trained transformer
- deleted 10mo ago[deleted]
- Kim_Bruning 10mo agoThat may or may not be true, but in the context of this article, I'm referring to OpenAI's GPT brand of models.
- johndill 10mo agoDid Calmmy Sammy that his is the version that will finally cure cancer? The AI shakeout in the AI industry is going to be brutal. Can't see how Private Equity is going to get the little guy to be left holding the giant bag of excrement, but they will figure that out. AI, smart enough to replace you, but not quite smart enough the replace the CEO or Hedge Fund Bros.
- astrange 10mo agoWhat do private equity or hedge funds have to do with any of this? Those are like, specific business models that are not involved in this situation.
- byt3bl33d3r 10mo agoThere’s really no point in looking at benchmarks anymore as real world usage of these models varies between task and prompting strategies. Use your internal benchmarks to evaluate and ignore everything else. It is curious to me how they don’t provide a side x side comparison of other models benchmarks for this release
- youngermax 10mo agoIsn't it interesting how this incremental release includes so many testimonials from companies who claim the model has improved? It also focuses on "economically valuable tasks." There was nothing of this sort in GPT-5.1's release. Looks like OpenAI feeling the pressure from investors now.
- nezaj 10mo agoWe saw it do better at making counter-strike! https://x.com/instant_db/status/1999278134504620363?s=20 https://x.com/instant_db/status/1999278134504620363?s=20
- stopachka 10mo agoFor those curious about the question: "how well does GPT 5.2 build Counter Strike?" We tried the same prompts we asked previous models today, and found out [1]. The TL:DR: Claude is still better on the frontend, but 5.2 is comparable to Gemini 3 Pro on the backend. At the very least 5.2 did better on just about every prompt compared to 5.1 Codex Max. The two surprises with the GPT models when it comes to coding: 1. They often use REPLs rather than read docs 2. In this instance 5.2 was more sheepish about running CLI commands. It would instead ask me to run the commands. Since this isn't a codex fine-tuned model, I'm definitely excited to see what that looks like. [1] The full video and some details in the tweet here: https://x.com/instant_db/status/1999278134504620363 https://x.com/instant_db/status/1999278134504620363
- TakakiTohno 10mo agoI use it everyday but have been told by friends that Gemini has overtaken it.
- blitz_skull 10mo agoAgain I just tap the sign. All of your benchmarks mean nothing to me until you include Claude Sonnet on them. In my experience, GPT hasn’t been able to compete with Claude in years for the daily “economically valuable” tasks I work on.
- nextworddev 10mo agoClaude is pretty trash for anything besides coding
- wyre 10mo agoWhat are you basing that on? Between Sonnet and Opus I don't think I'm reaching for Gemini 3 at all.
- timmg 10mo agoThat hasn't been my experience at all. I always wondered if we just get used to how to prompt a given model and that it hard to transition to another.
- romanovcode 10mo agoYeah, but that is the whole point of Claude. And that's why we are interested in the comparison.
- jstummbillig 10mo agoSince as per Anthropics own benchmarks Sonnet 4.5 is beaten by Opus 4.5 would it not suffice to infer the rest? https://x.com/OpenAI/status/1999182104362668275 https://x.com/OpenAI/status/1999182104362668275
- 8cvor6j844qw_d6 10mo agoWhat the current preferred subscription on AI? OpenAI and Anthrophic is my current preference. Looking forward to know what others use. Claude Code for coding assistance and cross-checking my work. OpenAI for second opinion on my high-level decisions.
- eastoeast 10mo agoFor the first time, I’m presenting a problem to LLMs that they cannot seem to answer. This is my first instance of them “endlessly thinking” without producing anything. The problem is complicated, but very solvable. I’m programming video cropping into my Android application. It seems videos that have “rotated” metadata cause the crop to be applied incorrectly. As in, a crop applied to the top of a video actually gets applied to the video rotated on its side. So, either double rotation is being applied somewhere in the pipeline, or rotation metadata is being ignored. I tried Opus 4.5, Gemini 3, and Codex 5.2. All 3 go through loops of “Maybe Media3 applies the degree(90) after…”, “no, that’s not right. Let me think…” They’ll do this for about 5 minutes without producing anything. I’ll then stop them, adjusting the prompt to tell them “Just try anything! Your first thought, let’s rapidly iterate!“. Nope. Nothing. To add, it also only seems to be using about 25% context on Opus 4.5. Weird!
- xmcqdpt2 10mo agoI don’t know if they used the new ChatGPT to translate this page but I was served the French version and it is NOT good. There are placeholders for quotes like <quote> and the prose is incredibly repetitive. You’d figure that OpenAI of all people would be able to translate something to one of the worlds most spoken language.
- matt3210 10mo agoCan this be used without uploading my code base to their server?
- yearolinuxdsktp 10mo agoPlus users are now defaulted to a faster, less deep GPT-5.2 Thinking mode called “Standard”, and you now have to manually select “Extended” to get back to previous deep thinking level for Plus users. Yet the 3K messages a week quota is the same regardless of thinking level. Also, the selection does not sync to mobile (you know, just not enough RAM in computers these days to persist a setting between web and mobile).
- rishabhaiover 10mo agoAfter I saw Opus 4.5 search through zig's std io because it wasn't aware of a breaking change in the recent release, I fell in love with claude-code and I don't see a strong enough reason to switch to codex at the moment.
- lend000 10mo agoIt seems like they fixed the most obvious issue with the last release, where codex would just refuse to do its job... if it seemed difficult or context usage was getting above 60% or so. Good job on the post-training improvements. The benchmark changes are incredible, but I have yet to notice a difference in my codebases as of yet.
- lazarus01 10mo agoMy god, what terrible marketing, totally written by AI. No flow whatsoever. I use Gemini 3 with my $10/month copilot subscription on vscode. I have to say, Gemini 3 is great. I can do the work of four people. I usually run out of premium tokens in a week. But I’m actually glad there is a limit or I would never stop working. I was a skeptic, but it seems like there is a wider variety of patterns in the training distribution.
- namesbc 10mo agoSo the rosy biased estimate is OpenAI is saving 1 hour of work per day, so 5 hours total per-work week and 20 hours total per-month. With a subsidized cost of $200/month for OpenAI it would be cheaper to hirer a part-time minimum wage worker than it would be to contract with OpenAI. And that is the rosiest estimate OpenAI has.
- dangoodmanUT 10mo agoA part time minimum wage worker can't code
- namesbc 10mo agoCheck the wages of coders outside of the US
- 6510 10mo agoThere use to be a mythological creature on irc from south America (sorry forgot the specifics) who was both a 10x dev and a 10x mathematician. One day he showed a picture of his computer. It was a low end laptop with a tft monitor and an external keyboard because the screen and the keyboard didn't work. It explained everything, the machine was just good enough to write code, do math, read stack exchange and lurk irc with his ghosts.
- AstroBen 10mo agoWhat people here forget is coding is a tiny minority of the actual usage. ~5% if I remember correctly? Their best market might just be as a better Google with ads
- ofermend 10mo agoGPT-5.2 just added to Vectara Hallucination Leaderboard. Definitely an improvement over GPT-5.1 - congrats to the team https://github.com/vectara/hallucination-leaderboard https://github.com/vectara/hallucination-leaderboard
- getnormality 10mo agoSweet Jesus. 53% on ARC-AGI-2. There's still gas in this van.
- rallies 10mo agoI work at the intersection of AI and investing, and I'm really amazed at the ability of this model to build spreadsheets. I gave it a few tools to access sec filings (and a small local vector database), and it's generating full fledged spreadsheets with valid, real time data. Analysts in wallstreet are going to get really empowered, but for the first time, I'm really glad that retail investors are also getting these models. Just put out the tool: https://github.com/ralliesai/tenk https://github.com/ralliesai/tenk
- rallies 10mo agoHere's a nice parsing of all the important financials from an SEC report. This used to be really hard a few years ago. https://docs.google.com/spreadsheets/d/1DVh5p3MnNvL4KqzEH0MEIW3oH4k2786GGhD6xqO-vOw/edit?usp=sharing https://docs.google.com/spreadsheets/d/1DVh5p3MnNvL4KqzEH0ME...
- sumedh 10mo agoDoesnt SEC provide XBRL data and the statements in excel?
- deleted 10mo ago[deleted]
- npodbielski 10mo agoCan't wait for being fired because some VP or other manager asked some model to prepare list of people with lowest productivity to pay ratio. Model hallucinated half of the data?! Sorry we can't go back on this decision, that would make us look bad! Or when some silly model will push everyone to invest in some radicoulous company and everybody will do it. Poisoning data attack to inject some I am Future Inc ™ company with high investment rate. After few months pocket money and vanish. We are certainly going to live in interesting times.
- buu700 10mo agoThat's more of a management problem than an AI problem. You could get the same result by replacing "model" with "intern" or "dude from Fiverr".
- vishal_new 10mo agoHmmm, is there any insight if these are really getting much better at coding? Will hand coding be dead within a few years, just human typing in english?
- psychoslave 10mo agoMia espero estas ke ne, ni nur parolos home inter homoj, robotoj anticipe faros servutoj por taŭge fari niajn dezirojn realigi laŭ niaj faktaj bezonoj. Kompreneble ni ĉiuj flue parolos Esperanto por taga geopolitikaj internaciaj aferoj, kaj ia ajn alia lingvo kiu plaĉas al mi por aliaj aferoj. Estonteco estas hela, miaj karaj siboj.
- CodeCompost 10mo agoFor the first time, I've actually hidden an AI story on HN. I can't even anymore. Sorry this is not going anywhere.
- gchokov 10mo agoHere, take my downvote.
- bigyabai 10mo agoIn lieu of a killer app?
- andybak 10mo agoHow this is different to any other post announcing an incremental improvement in an app or service?
- mabedan 10mo agoIt’s a little different. Most of these improvements are just more training hours and better weights. Even if it’s about actual improvement in trining algorithm or other software tweaks they’re not open source and hence other than “look how marginally nicer the chat bot responds now” the post doesn’t provide value.
- fasteo 10mo ago>>> Already, the average ChatGPT Enterprise user says AI saves them 40–60 minutes a day If this is what AI has to offer, we are in a gigantic bubble
- jatora 10mo agoThis seems pretty huge. Not sure by what metric it wouldn't be civilizationally gigantic for everyone to save that much time per day.
- svara 10mo agoIn my experience, the best models are already nearly as good as you can be for a large fraction of what I personally use them for, which is basically as a more efficient search engine. The thing that would now make the biggest difference isn't "more intelligence", whatever that might mean, but better grounding. It's still a big issue that the models will make up plausible sounding but wrong or misleading explanations for things, and verifying their claims ends up taking time. And if it's a topic you don't care about enough, you might just end up misinformed. I think Google/Gemini realize this, since their "verify" feature is designed to address exactly this. Unfortunately it hasn't worked very well for me so far. But to me it's very clear that the product that gets this right will be the one I use.
- phorkyas82 10mo agoIsn't that what no LLM can provide: being free of hallucinations?
- svara 10mo agoYes, they'll probably not go away, but it's got to be possible to handle them better. Gemini (the app) has a "mitigation" feature where it tries to to Google searches to support its statements. That doesn't currently work properly in my experience. It also seems to be doing something where it adds references to statements (With a separate model? With a second pass over the output? Not sure how that works.). That works well where it adds them, but it often doesn't do it.
- deleted 10mo ago[deleted]
- intended 10mo agoDoubt it. I suspect it’s fundamentally not possible in the spirit you intend it. Reality is perfectly fine with deception and inaccuracy. For language to magically be self constraining enough to only make verified statements is… impossible.
- jbkkd 10mo agoA new model doesn't address the fundamental reliability issues with OpenAI's enterprise tier. As an enterprise customer, the experience has been disappointing. The platform is unstable, support is slow to respond even when escalated to account managers, and the UI is painfully slow to use. There are also baffling feature gaps, like the lack of connectors for custom GPTs. None of the major providers have a perfect enterprise solution yet, but given OpenAI's market position, the gap between expectations and delivery is widening.
- sigmoid10 10mo agoWhich tier are you? We are on the highest enterprise tier and I've found that OpenAI is a much more stable platform for high-usage than other providers. Can't say much about the UI though since I almost exclusively work with the API. I feel like UIs generally suck everywhere unless you want to do really generic stuff.
- energy123 10mo agoChatGPT UI is leagues above Gemini and AI Studio in responsiveness and latency which is what I care about.
- dannyw 10mo agoCompletely the opposite experience.
- elAhmo 10mo agoThis feels like "could've been an email" type of thing, a very incremental update that just adds one more version. I bet there is literally no one in the world who wanted *one more version of GPT* in the list of available models from OpenAI. "All models" section on https://platform.openai.com/docs/models https://platform.openai.com/docs/models is quite ridiculous.
- tim333 10mo agoIt's significant because it looked like they were falling behind Gemini and maybe others.
- m12k 10mo agoSo, does 5.2 still have a knowledge cutoff date of June 2024, or have they managed to complete another full pre-training run?
- bob1029 10mo agoI've been looking really hard at combining Roslyn (.NET compiler platform SDK) with one of these high end tool calling models. The ability to have the LLM create custom analyzers and then verify them with a human in the loop can provide stable, compile-time guarantees of business rules that accumulate without paying for context tokens. I feel like there is a small chance I could actually make this work in some areas of the business now. 400k is a really big context window. The last time I made any serious attempt I only had 32k tokens to work with. I still don't think these things can build the whole product for you, but if you have a structured configuration abstraction in an existing product, I think there is definitely uplift possible.
- schmuhblaster 10mo agoSounds interesting, could you elaborate a bit on this? (I am experimenting in a similar direction)
- dev1ycan 10mo agoHow many years of the world's DRAM production capacity is it this time?
- throwaway2037 10mo agoSomewhat tangential: The second link says "System card": https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944f8d/oai_5_2_system-card.pdf https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944... Does that term have special meaning in the AI/LLM world? I never heard it before. I Google'd the term "System Card LLM" and got a bunch of hits. I am so surprised that I never saw the term used here in HN before. Also, the layout looks exactly like a scientific paper written in LaTeX. Who is the expected audience for this paper?
- tylerrobinson 10mo agoThe major model providers use system cards as a sort of self attestation document like a nutrition label. It’s been around for a couple years.
- blinding-streak 10mo agoYeah, search HN for the term. It's a relatively big topic of conversation.
- EastLondonCoder 10mo agoI’ve been using GPT-4o and now 5.2 pretty much daily, mostly for creative and technical work. What helped me get more out of it was to stop thinking of it as a chatbot or knowledge engine, and instead try to model how it actually works on a structural level. The closest parallel I’ve found is Peter Gärdenfors’ work on conceptual spaces, where meaning isn’t symbolic but geometric. Fedorenko’s research on predictive sequencing in the brain fits too. In both cases, the idea is that language follows a trajectory through a shaped mental space, and that’s basically what GPT is doing. It doesn’t know anything, but it generates plausible paths through a statistical terrain built from our own language use. So when it “hallucinates”, that’s not a bug so much as a result of the system not being grounded. It’s doing what it was designed to do: complete the next step in a pattern. Sometimes that’s wildly useful. Sometimes it’s nonsense. The trick is knowing which is which. What’s weird is that once you internalise this, you can work with it as a kind of improvisational system. If you stay in the loop, challenge it, steer it, it feels more like a collaborator than a tool. That’s how I use it anyway. Not as a source of truth, but as a way of moving through ideas faster.
- ostacke 10mo agoInteresting concept with conceptual spaces, but how does that affect how you work with LLM:s in practice?
- EastLondonCoder 10mo agoI think of it like improvising with a very skilled but slightly alien musician. If you just hand it a chord chart, it’ll follow the structure. But if you understand the kinds of patterns it tends to favour, the statistical shapes it moves through, you can start composing with it, not just prompting it. That’s where Gärdenfors helped me reframe things. The model isn’t retrieving facts. It’s traversing a conceptual space. Once you stop expecting grounded truth and start tracking coherence, internal consistency, narrative stability, you get a much better sense of where it’s likely to go off course. It reminds me of salespeople who speak fluently without being aligned with the underlying subject. Everything sounds plausible, but something’s off. LLMs do that too. You can learn to spot the mismatch, but it takes practice, a bit like learning to jam. You stop reading notes and start listening for shape.
- rl_shannon 10mo agoIsn't it delusional to only compare your models against your own previous variants? Where is an actual comparison with Google, Anthropic, OSS Models
- zild3d 10mo agoit's the best ____ we've ever made
- keepamovin 10mo agoIt is significantly better than 5.1 .. testing now with codex. It's much more focused, perceptive and efficient.
- deleted 10mo ago[deleted]
- loa_observer 10mo agodoes the model really improve? i tried several tasks today, and most of them failed, which are super easy ones. maybe it's just because the gpt5.2 in cursor is super stupid?
- atheljcarlton 10mo agoIt's dog-doo-doo. I put in my algebraic geometry final review (100's of thousands of tokens) and Gemini instantly found all the propositions, theorems, and problems that I needed in a neat list (in about 5 seconds), meanwhile ChatGPT 5.2 Thinking took 10mins before timing out and not even completing the request.
- atheljcarlton 10mo agoHowever, the model card for GPT 5.2 looks amazing, wish I could actually see that performance in action!
- xnx 10mo agoWhat's a more accurate name for this model? GPT-4 v3?