30 ms·
GPT-5.6
https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf
https://developers.openai.com/api/docs/guides/latest-model https://developers.openai.com/api/docs/guides/latest-model
https://x.com/levie/status/2075287443411222628 https://x.com/levie/status/2075287443411222628, https://xcancel.com/levie/status/2075287443411222628 https://xcancel.com/levie/status/2075287443411222628
- 5555watch 3mo agoSo with this release do they kill the 5.5-Pro model with super long thinking and reasoning? 5.6-Sol-Ultra is not the equivalent, right?
- tipiirai 3mo agoThought Fable was great
- system2 3mo agoAt this point, they are just changing the decimals to stay relevant and in the news.
- dude250711 3mo agoAnthropic should be grateful OpenAI did not borrow "Epic" and "Legend".
- system2 3mo agoI expect OpenAI names to be "fabulous", "glorious", "empowered", "delicious" etc.
- realty_geek 3mo agoWow, the "Agents' Last Exam" graph looks unreal!
- therobots927 3mo agoThat’s because it’s bullshit
- alimhaq 3mo agoI mean the y axis is deceptive to make it seem like greater gains since it starts at 30%, when in reality the differences aren't great. Even worse, it's not a fair comparison: they purposefully just used "adaptive" instead of "max" for Fable. What about the graph looked so unreal to you?
- tedsanders 3mo ago> Even worse, it's not a fair comparison: they purposefully just used "adaptive" instead of "max" for Fable. We agree models should be compared on a fair basis. Unfortunately, adaptive was the only publicly available number. Anthropic doesn't generally let us run their models for evals, so we rely on whatever Anthropic or third parties have published. In this case, the Agents' Last Exam leaderboard has Fable Adaptive, but not Fable Max. https://agents-last-exam.org/leaderboard https://agents-last-exam.org/leaderboard Would have loved to publish a full curve for Fable if anyone makes the data available. Although we do bias toward publishing evals where we're ahead, we have historically been unafraid to publish evals where we're behind (e.g., GDPval). The point is give people useful information to decide what's best, not to trick people. Edit: Now I see there's a second entry with xhigh effort. Not sure if that was added or recently or we skipped it. (I work at OpenAI.)
- poolnoodle 3mo ago[flagged]
- sidcool 3mo agoThe claims are pretty bold. I think 5.6 may exceed Fable.
- Syntaf 3mo agoOk long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
- moomoo11 3mo ago[dead]
- Daedren 3mo agoUse a harness that doesn't lock you into a moat, like OpenCode.
- AntonyGarand 3mo agoCan't use a claude code subscription in another harness though
- greenavocado 3mo agoYou absolutely can; they are not banning anymore. The bigger problem is that subscription versions of the models are way crappier than when the "same" model is hit via API (Bedrock/Vertex) You can also make it not count against extra usage. OpenCode docs show it because Anthropic specifically ambushed them with a PR to remove support so simpletons can't use it easily.
- infberg 3mo agoDo you have a source for that?
- deleted 3mo ago[deleted]
- AntonyGarand 3mo ago
- deleted 3mo ago[deleted]
- saberience 3mo ago"On Agents’ Last Exam (opens in a new window), an evaluation of long-running professional workflows across 55 fields, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. That efficiency extends to smaller models, which are essential to making intelligence more abundant and affordable: GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost. " Some pretty big claims and results! Excited to see how it feels during usage. I use Fable and 5.5 extensively and I still find both have a place in my toolkit, i.e. Fable IS good but it isn't perfect, and it's still better to play them off against each other. I have Fable and 5.5 write plans and have them adversarially review each other's plans. Having this amount of competition in the coding model space is good for all of us.
- vamsiraju 3mo agoI think this is the phase shift 5.6 (Sol set to Ultra) is bringing to the table. Until now we have become accustomed to asking models to continue and their natural inclination is always to stop. Now OpenAI have flipped it around and for the first time are asking us to steer or stop the model instead, and its own inclination is to keep going. We now have to decide when we need to steer or want to catch up on our understanding of the work done but it will keep going.
- deleted 3mo ago[deleted]
- rvz 3mo agoMost importantly, the cost: > GPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. Just as expensive as Fable 5. But of course, another slot machine upgrade but the costs will keep going up and the open weight models from china will continue to race everyone else to $0. Looking forward to the next version of GLM, Qwen, Deepseek and Minimax.
- therobots927 3mo agoAlso watching deepseek closely. Seems like US frontier labs only know how to throw money at things as opposed to actually make smart improvements to the algorithms.
- trollbridge 3mo agoTo be fair, DeepSeek doubled prices during the peak Chinese workday. (Which admittedly doesn't affect me much.)
- deleted 3mo ago[deleted]
- celesian 3mo agoThat's wrong. GPT 5.6 Sol looks to have the same price as GPT 5.5, apart from a new pricing fee for cache writes. Fable 5 is $10 input / $50 output.
- sd9 3mo agoI haven't tried an OpenAI model for a long time, but with Fable going to API pricing soon this might be enough to get me to try codex.
- danielbln 3mo agoSeeing how Anthropomorphic just reset usage quotas back to 0 and the other day extended Fable sub inclusion by a few days, I have a feeling they might not drop Fable out of sub after all, because like you I would most definitely take a long good look at codex at that point.
- sd9 3mo ago[dead]
- matheusmoreira 3mo agoIt's not just the API pricing either, there's also the constant uncertainty. They pull the model then put it back up, they say the model is going away then suddenly it's not. And then there's the fact Fable is barely usable because it randomly downgrades to Opus out of nowhere whenever it thinks about exploits. It's definitely good that Anthropic's feeling the pressure. Anthropic has worn out their welcome with this "safety" nonsense. If OpenAI actually lets me use the LLMs on a subscription without any of this bullshit, I'll definitely switch.
- kouteiheika 3mo ago> And then there's the fact Fable is barely usable because it randomly downgrades to Opus out of nowhere whenever it thinks about exploits. I suspect it's not what you meant, but it's definitely not random and is very deliberate. Just today I got it to reliably trigger the "safety" filter with (drumroll) having it list the weight keys of a 300M parameter ModernBERT-derived model. Their "safety" classifier must be matching one of the key names in there and trigger their "this is a frontier model" anti-competitive filter[1] (even though it's just a tiny 300M parameter model, four orders of magnitude smaller than the frontier). [1]: https://news.ycombinator.com/item?id=48464732 https://news.ycombinator.com/item?id=48464732 Fortunately once you know how it works (i.e. dumb keyword classifier) it's easy-ish to get around: just rename the keys so that it doesn't contain the naughty keyword. (At least as long as it doesn't trigger on something in its own thinking trace, which needs... more creative workarounds.)
- cbg0 3mo ago5.6 Terra (mid tier model) as good as Fable on DeepSWE while cheaper than Opus API pricing. Seems like a homerun.
- osti 3mo agoGPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a look at imo.
- DetroitThrow 3mo agoDeepSWE seems to strongly, strongly prefer ChatGPT models. There were also major flaws in its methodology pointed out recently, that overlap strongly with the flaws OpenAI pointed out in its SWE Verified report. I use both ChatGPT and Claude for engineering work on a daily basis, touching performance critical code to application backends to frontend work, and I've found that DeepSWE scores don't reflect my reality when I assess high quality output from the models/harnesses. Not that Opus always beats GPT 5.5., but that 5.5 is ahead of Opus on a general benchmark smells off to me.
- minimaxir 3mo agoThe developer's guide (https://developers.openai.com/api/docs/guides/latest-model https://developers.openai.com/api/docs/guides/latest-model) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the original dimensions of images sent with original or auto detail instead of resizing them to a patch budget or pixel-dimension limit. > Use shorter prompts: In internal evaluations, replacing long, explicit system prompts with minimal prompts improved scores by roughly 10–15%, while reducing total tokens by 41–66% and cost by 33–67%. > Avoid generic brevity instructions: GPT-5.6 is more sensitive than GPT-5.5 to instructions such as “Be concise,” “Keep it short,” or “Use minimal text.” > Control warmth: GPT-5.6 does not become meaningfully better when prompted to be broadly friendlier or more empathetic.
- ravenstine 3mo ago> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things for developers who didn't know to make prompt changes after upgrading to 5.6. What if you were expecting GPT to be as wordy as it usually is? Then suddenly your output is not wordy enough? Smells like OpenAI trying its best to stave off financial armageddon for another few months. Then again, I'm not sure why they chose to waste so much output computation on verbal diarrhea all this time up to now.
- anticorporate 3mo agoIt seems like the way brevity instructions have changed is mis-aligned with how most people would expect to use them or are currently using them. Here's the example they give: > Instead of asking for the shortest possible answer, replace brevity instructions with prioritization: > Lead with the conclusion. Include the evidence needed to support it, any material caveat, and the next action. Omit secondary detail and repetition. > Keep all required facts, decisions, caveats, and next steps. Trim introductions, repetition, generic reassurance, and optional background first. Generally speaking, when I ask for a short answer, I want a short answer because I'm not really willing to read through a bunch of bullshit to get to a summary. Putting the onus back on me to assume what the model will return and write a longer prompt detailing exactly what information I want completely misses the point of why I'm asking for a short answer in the first place.
- willchis 3mo agoThe marketing team must've done research that said "people are starting to think that you guys are evil-water-stealing-lay-off-loving-bubble-bursting scumbags" and decided to really lean into the small family business and happy font vibes!
- enraged_camel 3mo agoCTRL-F: Fable 15 hits Holy shit. They must be feeling very threatened by Fable if they're spending this much energy talking about it in the release notes for their own model.
- therobots927 3mo agoApparently it significant outperforms fable on both an intelligence and cost index. I don’t believe it at all and I don’t think anyone else does either.
- trollbridge 3mo agoI believe that it outperformed it on benchmarks.
- cbg0 3mo agoIn the past they received a lot of hate for not comparing to the competition.
- InsideOutSanta 3mo agoSir, you do not understand! This is the Internet! You must always find a reason to be upset and/or complain!
- BrokenCogs 3mo agoyikes - looks like you need to go back to stats school gemini - 13 hits opus - 18 hits So they are more threatened by opus than fable, or are they almost as threatened by gemini as they are by fable?
- enraged_camel 3mo agoThe second paragraph has four mentions of Fable. I think that makes my case pretty clearly.
- simianwords 3mo ago
- dude250711 3mo agoNot available - checked and it's not there.
- cactusplant7374 3mo agoThey have really been stringing us along for the past few weeks.
- tedsanders 3mo agoAs usual, even though GPT-5.6 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (same as our previous launches). We usually start with Pro/Enterprise accounts and then work our way down to Plus. We know it's slightly annoying to have to wait a random amount of time, but we do it this way to keep service maximally stable. The timescale is typically hours not minutes, so if you don't see it now, I'd try again later today. We mention it will be a gradual rollout over the next 24 hours in the Availability section at the bottom of the blog but I admit it's pretty buried. (I work at OpenAI.)
- dude250711 3mo agoUnderstood thanks; will 5.6 fix this issue that makes Pro unusable? https://github.com/openai/codex/issues/30364 https://github.com/openai/codex/issues/30364 "GPT-5.5 Codex reasoning-token clustering at 516/1034/1552 may be leading to degraded performance on complex tasks"
- jiggawatts 3mo agoIs this bug fixed with 5.6? If not, it probably doesn’t matter which version Codex users are getting because the overall result is dramatically worse than stated by Open AI advertising: https://github.com/openai/codex/issues/30364 https://github.com/openai/codex/issues/30364
- tedsanders 3mo agoNot entirely fixed yet, but should be rarer with 5.6. Don’t have a quantification, unfortunately.
- therobots927 3mo agoDo they expect us this model is 15ppt more accurate at half the price of fable? What’s going on?
- arizen 3mo ago"GPT‑5.6 delivers a step change in design judgment. With only high-level direction, GPT‑5.6 creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result—not just generate the underlying code or content—so it can catch visual and functional issues and apply finishing touches before handing the work back." This one is really promising, as it may allow to close major gap with Claude in design/UI skills
- GenerWork 3mo agoAgreed, I’m looking forward to trying it out. I think that the rise of visual design skills that are pretty clearly targeted towards Codex users has lit a bit of a fire under their butts.
- HyperL0gi 3mo ago+1. I've been only using Sonnet/Opus these days for UI work because GPT 5.5 just can't do any of that. Its just really terrible. Eager to give this one a try.
- silksowed 3mo agoComputer-use is a big limitation that my 2015 Macbook Pro cannot handle. I find the Codex cli says it looks at the end output artifact but so often it fails to refine it into acceptable form. If it could use my computer screen and visual inputs for review, it might be able to actually design documents/powerpoints/etc. I'm juicing everything I can out of the 11 year old laptop and I'm honestly impressed at what it can still do.
- semiquaver 3mo agoHow dare you point out that 2015 is 11 years ago.
- silksowed 3mo agohahaha, makes me sad and happy all at once...
- eig 3mo agoFunny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!
- fblp 3mo agooh that's sad, are the biolgy limitations for "safety"?
- sunnybeetroot 3mo agoYes
- paxys 3mo agoWhere’s the lie?
- deleted 3mo ago[deleted]
- inciampati 3mo agoThis is a major reason why I and a number of biologists I've talked to have canceled their anthropic accounts recently. Not working is not working.
- fellowniusmonk 3mo agoI mean it's a fucking joke, I kept getting refusals on a code base I wasn't familiar with and it was literally just because there are some vars named DNA. Just absolutely stupid.
- steve_adams_86 3mo agoIt's so absurdly sensitive. It bailed out earlier today working on a TypeScript client for a sensor network API which happens to include some temperature and pH sensors for tanks, which yes, are used for biology experiments. But wow, we're degrees of separation from the actual biology work. It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust it to do a task without deferring to Opus and that's really annoying at times. I want to know what I'm getting up front, not after the fact.
- deleted 3mo ago[deleted]
- newfriend 3mo ago>Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. Sounds great. Also latency looks very good.
- deleted 3mo ago[deleted]
- GodelNumbering 3mo agoDirac (https://github.com/dirac-run/dirac https://github.com/dirac-run/dirac, https://dirac.run/ https://dirac.run/) now supports gpt-5.6. This thing does now seem to be on the chatGPT/codex accounts yet. UPDATE: it is now available in chatGPT account also, they rolled it out
- denysvitali 3mo agoWill be there soon according to the last commits in the codex repo: https://github.com/openai/codex/pull/31684/changes https://github.com/openai/codex/pull/31684/changes Also, confirmed it works for me by using --model gpt-5.6-sol
- BoorishBears 3mo agoI used to pride myself on not being the "fonts too pointy, scroll too buttery" crowd! But AI has brought me full circle and now nothing removes my interest in reading even a single word on a page faster than purple gradient greeble-afflicted tailwind-slop models put out without stronger prompting/references That being said, maybe 5.6 can fix that!
- GodelNumbering 3mo agoThanks, I needed to hear that lol. Yes, the site was an afterthought, core work took/takes most my focus. I will look into un-slopping the site soon.
- esafak 3mo agoDoes it support subagents?
- GodelNumbering 3mo agoyup it does
- hughw 3mo agoIf it's not dangerous enough to be classified as WMD by USG, who's interested.
- fgeytk2 3mo ago[dead]
- deleted 3mo ago[deleted]
- big_toast 3mo agoThe cost & output token charts are useful but I wish I could view them more like a 3D surface. Like the CS:APP memory mountain charts. I wonder how long model size and effort will be a few discrete points instead of continuous.
- Tenoke 3mo agoIs any of those comparisons about Pro vs non-Pro (Pro is only available in $100+ plans)? I am curious about that but I think Sol, Terra, Luna are different sizes of it without the Pro part, and I want to know how much worse do I have it on the $20 plan compared to if I upgrade.
- mchinen 3mo agoThe frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressive over Fable, I will lose what remaining trust I had in these announcements.
- therobots927 3mo agoThe proof is in the pudding and these benchmark stats will only work for so long before people lose interest.
- thurn 3mo agoThey do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...
- mchinen 3mo agoI totally missed that, because in the charts they showcase for coding, the SWEBench score is not present, they only include it at the end of the post in tables. Hmm. Great catch.
- saberience 3mo agoThe SWEBench benchmarks are really gamed at this point and should not be trusted period. The solutions are effectively in the training sets and have been for a while.
- mnicky 3mo agoSWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.
- gozucito 3mo agoThe meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.
- paxys 3mo agoMakes sense why they released an entire study yesterday discrediting SWE-bench Pro.
- osti 3mo agoAnd they'd be right, it's an almost saturated benchmark where even some subpar open source models score very well on. And most models are clustered within a small range so it really doesn't tell you much.
- osti 3mo agoSWE-Bench pro is pretty much useless now even though many ppl still look at it. OpenAI published a report yesterday saying so as well. Only look at DeepSWE and FrontierCode right now for coding imo.
- SirMaster 3mo agoAmazing, a company that does poorly in a benchmark says that benchmark is useless...
- osti 3mo agoSWE-bench series just aren't that great by today's standard, even Anthropic previously stated Claude had memorized solutions for the non Pro version of the benchmark, I suspect the recent increase in the score for the Pro version probably also had similar behaviors. But anyway, I think it's pretty useless to look at SWE Bench's now when other way better benchmarks exist.
- 3mo ago
- kubb 3mo agoThey have a fantastic media team.
- luciana1u 3mo ago[flagged]
- nharziro 3mo agowhere is it? Still not accessible...
- jstummbillig 3mo ago"GPT‑5.6 is available starting today across ChatGPT, Codex, and the OpenAI API. The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours."
- bryceneal 3mo agoI find that 5.5 gives me far fewer refusals than Anthropic models for security and reverse engineering work. I hope the same is true for 5.6.
- SwellJoe 3mo agoYeah, I pretty much had to switch to using GPT rather than Opus completely for all my security benchmarking and harness development. I was annoyed enough to blog about it: https://swelljoe.com/post/why-i-had-to-switch-to-gpt/ https://swelljoe.com/post/why-i-had-to-switch-to-gpt/
- artisin 3mo agoI wish they had kept their previous sensible naming convention instead of this celestial Sol, Terra, and Luna mumbo-jumbo
- SwellJoe 3mo agoI assume they're jealous of the Fable/Mythos hype. People talk about Fable like it's a whole new thing, rather than another incremental improvement over the existing best models (which has happened several times and continues to happen).
- samuelknight 3mo agoThere is an issue on the page that causes the benchmark tables to get cut off. If you highlight and drag right you can see a few more models like Gemini and Claude Opus. It's also interesting that they introduced explicit caching, which is something that only Anthropic had for a long time.
- m3h 3mo agoWe have an official pelican on a bicycle from the OpenAI livestream: https://imgshare.cc/mz9xwut3 https://imgshare.cc/mz9xwut3
- BrokenCogs 3mo agoholy moly it's in THREE dimensions! AGI solved
- SirMaster 3mo agoSo it's failing epically because it generated a tricycle instead of a bicycle?
- jstummbillig 3mo ago"GPT‑5.6 is available starting today across ChatGPT, Codex, and the OpenAI API. The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours."
- terramex 3mo agoI am on Plus subscription and see Terra and Luna in Codex, but no sign of Sol. Will it be available only on Pro plans?
- lostmsu 3mo agoI an on Pro and it still returns "The 'gpt-5.6' model is not supported when using Codex with a ChatGPT account" UPD from announcement: "The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours."
- enraged_camel 3mo agoMy Codex app got upgraded to the new unified ChatGPT app. I don't see Sol available though. Only Terra and Luna. I'm on the Pro plan. Anyone else see it?
- philip1209 3mo agoWill this run on Cerebas? I'm really looking forward to that.
- paxys 3mo agoSam Altman confirmed during the initial limited release that Sol will run on Cerebras at 750 tok/sec.
- kordlessagain 3mo ago"I canna' give her any more, Captain!" - Montgomery "Scotty" Scott, Chief Engineer
- apitman 3mo agoThis is the part I'm most excited about with the new release, though I'm concerned plebs like me may never get a chance to play with it
- simianwords 3mo ago> On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning sets a new state of the art at 80, 2.8 points above Fable 5, while using less than half the output tokens, taking less than half the time, and costing about one-third less. > That advantage extends across the family: Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost. Wow. I don't believe it. Every indication and twitter post told me that Fable is much more intelligent than Sol and here we are told that even Terra outperforms Fable? Not only that, Sol doesn't even come with run time classifiers. So it is even more suspicious. What's even stranger is that OpenAI is directly referencing a competitor in this direct way.
- I_am_tiberius 3mo agoThe way they talk about cyber security fixes makes clear that they are in bed with the government in order to get ahead of Anthropic.
- culi 3mo agoAll of them closely collaborate with the government. LLMs are a national security priority and are vetted. Claude AI was used by Palantir's Maven to target the Minab school that led to a triple tap strike killing over 150 schoolchildren.
- applicative 3mo agoThe Minab disaster has every sign of being a pure humint fail the defense department decided to cover up with politically expedient AI blaming.
- culi 3mo agoClaude makes suggestions for targets and humans review and approve them. We always knew there was a human in the process but if there's one massive takeaway from years of AI ethics research, it's that there is a very clear and well-documented human bias towards automated answers when there's any ambiguity. Including a human in the loop does not excuse the fact that AI was trusted in a process that decides who lives and dies.
- esafak 3mo agoThe DIA's Maven database was out of date: https://www.theguardian.com/news/2026/mar/26/ai-got-the-blame-for-the-iran-school-bombing-the-truth-is-far-more-worrying https://www.theguardian.com/news/2026/mar/26/ai-got-the-blam...
- culi 3mo ago"We triple tapped a girls elementary school because our data wasn't up to date."
- vinhnx 3mo agoGPT‑5.6 system card https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf
- browski 3mo agoHere's me using a Gemini chat log scraper (from Gdrive) then dumping my prompt+Gemini response into local AI Never go over the free limits in Gemini Pro. Gemini is great at research and architecture, and my 30 years experience in programming everything; for fun or work; means together there is little to no code slop. Add to project repo some git submodules of reference source code; boom, bobs your uncle Zero reason to sign up for OAI or Claude. With employers realizing the costs are more than employees, local models getting more powerful, and models in chips just a few years out, neither of the one note LLM companies without diversified services and R&D portfolios gonna last
- mchusma 3mo agoLooks like a great set of models, but there are about 20 different thinking/model levels here in this family and they are very complex to pick the right one for the task E.g. for GeneBench Pro, it looks like you would always use GPT-5.6 Sol over Terra/Luna, its pareto optimal. For Agents Last Exam, you would maybe want Luna, then Terra, then Luna, then Sol as you increasingly budget for tasks. I feel that there may need to be a new auto mode in many of these cases. It selects the best model and thinking given a particular problem. Feels like it's going to have to go that way eventually, because here we have about 20 different model and thinking levels you could use, and they're not obvious which ones are right for the given use case.
- WarmWash 3mo ago8% on ARC-AGI-3, they actually got some traction going...
- eugene3306 3mo agonote that ARC-AGI-3 has its rules changed. before today all the contestants were capped at $10k
- hereme888 3mo agoSo is 5.5-Daybreak still relevant for cyber security give. 5.6 capabilities?
- twothreeone 3mo agoWow the video is much better.. the PR spend clearly went up a lot. Mainly just showing "real people" doing "real stuff".
- deleted 3mo ago[deleted]
- fractorial 3mo agoSounds like a perfect fit for a minimal or bespoke harness?
- Jcampuzano2 3mo agoI really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confusing. That aside all the numbers look amazing, and I'll be happy to probably main this alongside grok-4.5 for a while comparing the two on price and efficiency. I vastly prefer the direction that OpenAI seems to be going with token efficiency and performance compared to Anthropic who seems to be moving towards a world where you just token-max as much as possible ignoring any and all costs.
- jstummbillig 3mo agoWhy would you need a guide for that now? We long had to pick different models (and thinking levels) by task and feel.
- Jcampuzano2 3mo agoPreviously it was much more obvious which model to reach for depending on your use case because they had the mini and nano naming conventions. Getting rid of that seems like a step back. Just a personal nit though. I've seen buzz about this elsewhere as well but to me effort levels seem more like spend limits disguised with another word. I don't think they should even exist.
- bigyabai 3mo agoThe naming convention is bizarre and doesn't really mean anything to normies. Trying to pick between "Sol" and "Terra" is like asking the average person if they want the Max or the Ultra chip.
- copperx 3mo agoBizarre? The size of the model is in the name. Sun, earth, and moon don't mean anything?
- infraredshift 3mo ago[dead]
- meetpateltech 3mo agoGPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6 https://arcprize.org/results/openai-gpt-5-6
- simianwords 3mo agoVery interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?
- osti 3mo agoMythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
- vatsachak 3mo agoAre you joking? They spend billions of dollars training LLMs to get a 7.8% on arc agi 3 whereas DINO models are near sota in image classification, provide meaningful embeddings to the point where image segmentation is just PCA. The spend on DINO cannot be more than five million (correct me if I'm wrong) JEPA is just getting started
- esafak 3mo agoASI is going to be here by the time Lecun gets started.
- redactsureAI 3mo agoDINO is a transformer model?
- 3mo ago
- karma_daemon 3mo agoI wish model launches were like proper product releases it's impossible to _try_ it out on release! it's not on their codex subscription, or the web/mobile chatgpt interfaces, or aws bedrock, etc. I just cant find a working endpoint with the latest model after they announce
- O5vYtytb 3mo agoThe announcement says they're rolling it out over the next 24 hours or so. I think it's reasonable to do a slow-roll-out release for one of the most used products on the internet.
- patapong 3mo agoGPT-5.6 Terra just showed up in Codex for me.
- ssl-3 3mo agoFor me, minutes ago, as a Plus subscriber: I started up Codex CLI fresh. That version of Codex was 1.42.5. 5.6 wasn't in the models list. After I updated Codex to a newer version (0.144.0), 5.6-terra and -luna appeared in the models list (but not 5.6-sol). (It's impossible for me to know whether updating was causative or just correlative, but that's the timeline I experienced.)
- bearmania 3mo agoIf OpenAI can add all the features from CC into Codex i’ll gladly switch.
- bearmania 3mo agoif OpenAI adds all the features from CC into Codex, i’ll gladly switch.
- gavino 3mo agoWhat features are you missing? That you can't add through skills?
- CjHuber 3mo ago> Instead of requiring developers to script every step or passing every tool response back through the model, Programmatic Tool Calling in the Responses API can filter large amounts of intermediate data, retain only what matters, and adapt its workflow along the way. this seems very interesting
- aliasxneo 3mo ago"We've extended usage of Claude Fable" message incoming any day now.
- ftchd 3mo agoThey reset all usage half an hour ago. It's back to 0% per week and session. No specifically Fable related.
- halfmatthalfcat 3mo agoHahaha seeing this play out in real time is absolutely incredible.
- danielbln 3mo agoIm here for it, good on Anthropomorphic to feel some heat again after all that drug dealer Fable business.
- aliasxneo 3mo ago100. I'm so tired of being treated like a drug addict by them. I'm currently sitting on _four_ "reset vouchers" from OpenAI so I get basically all week to play with 5.6 Sol to my hearts content. The amount of positive sentiment that brings to me should really be a concern for Anthropic who is increasingly alienating me away with their shit strategies.
- morgengold 3mo agoyou are a drug addict, though. we all are
- aliasxneo 3mo agoTo an extent, but I still feel consciously addicted at this point, and Anthropic's antics seem to suggest that I've reached the unconscious level where I'll do anything to keep access, including their insane street API price.
- m3h 3mo ago[dead]
- big_toast 3mo agoIn the introduction video they say 5.6 Sol autonomously post-trained 5.6 Luna. Curious what this means.
- vibcdingenjoyer 3mo agoSounds like they gave it a goal to hit certain benchmarks and just let it have its way with the base Luna model.
- block_dagger 3mo agoThis produced a disturbing mental image.
- 2001zhaozhao 3mo agoIt means OpenAI and Anthropic are now in a RSI race with each other
- jrflo 3mo ago/goal tune 5.6 Luna parameters until performance is maximized across all benchmarks
- ddxv 3mo agoI'm disappointed these models continue to be closed source and so expensive. Open weight models being 10x or more cheaper is just so much more of an unlock than incremental gains for me.
- AgentMasterRace 3mo agobecause they're stealing from the frontier models. they're gaming the benchmarks. look how bad glm 5.2 is on cursors evals. gmhit garbage , but it gets glazed as God tier.
- anematode 3mo agoThey're stealing, eh?
- maxdo 3mo agocursor benchmarks with GPT 5.6 in picture, a good reason to stop using opus. https://cursor.com/evals https://cursor.com/evals The good news you don't have to send your dollars to China to fund ai dictatorship, in russia, north korea, african countries and south america.
- AgentMasterRace 3mo agoso the answer is use grok ?
- maxdo 3mo agoI'd say answer , the opus is no longer undisputed. grok + gpt models are very competitive + glm if you are ok to wait 3-4 times longer, unless you have some unique access to GPU
- deleted 3mo ago[deleted]
- neuropacabra 3mo agoIs it available in EU? I only see 5.5 still :-(
- yokoprime 3mo agoArguably not in the EU, but I'm seeing Sol,Terra and Luna on my account here in Norway
- hereme888 3mo agoI use 5.5 a ton. It's immediately apparent that 5.6 is truly a better model. Hope they don't lobotomize it later.
- lukebuehler 3mo agoOh man, I love capitalism spoiling us here. I was just enjoying my extra Fable credits, now I'll switch to using 5.6 this weekend. I was planning to ration my Anthropic credits, I guess now I do not have to. And I was half wondering if exactly this would happen: right when Fable usage credits were starting to kick in for people, OAI swoops in and takes the puck. As much the AI craze is crazy, this play by play part is pretty fun.
- sidrag22 3mo agotop it off with anthropic stressing about the release and resetting usage to 0 for the week just now.
- halfmatthalfcat 3mo agoAnthropic just reset all limits, including Fable. Capitalism is spoiling us.
- lukebuehler 3mo agoMake hay while the sun is out.
- mempko 3mo agoIf hoarding is spoiling. You know what would be better than using fable and gpt 5.6, being able to run that level of model on your own hardware.
- vamsiraju 3mo agoprompts -> loops -> slingshots? Its an extremely capable model. I think the way we need to approach works shifts again. We need to get our harnesses/workflows to let it gather some momentum on the first couple rounds but then we also need to structure it so that it can slingshot and accomplish the long range goal.
- vatsachak 3mo agoI wonder what increment of progress will be achieved by the next billion dollars
- Razengan 3mo agoJust a day before my $100 subscription expires, perfect
- hyperknot 3mo ago> GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints (opens in a new window) and a 30-minute minimum cache life. Great to read they are moving away from the 5 minute cache defaults. Hopefully other providers follow soon!
- hrpnk 3mo agoThey highlight the cache write price now much more in the guide. Did it increase vs. prior generations?
- hyperknot 3mo agoThere was no cache write before! https://openrouter.ai/openai/gpt-5.5?endpoint=58e5b336-423e-430b-a2ab-8bc353f0c51b https://openrouter.ai/openai/gpt-5.5?endpoint=58e5b336-423e-... vs https://openrouter.ai/openai/gpt-5.6-sol?endpoint=a54c5de0-89bf-4ad7-a212-cf977eed918a https://openrouter.ai/openai/gpt-5.6-sol?endpoint=a54c5de0-8...
- OutOfHere 3mo agoLike the last time, again they failed to note whether there is an Instant model or when it might become available.
- mNovak 3mo ago>> approximately 700,000 A100e GPU hours of black-box automated red teaming Amusing that they use A100e as the reference point to sound impressive. Different ways you could make that conversion, but based on FP4 FLOPs (yes it's disadvantageous to A100, that's the point), that's something like 200hr on a GB300 NVL72 rack. Not nothing either, but far less astounding sounding than 700k hrs.
- stavros 3mo agoWait, what do you mean? 700k A100e hours are equal to 200 hours of a GB300 NVL72 rack? One GB300 NVL72, 72-GPU rack has equal processing power to 3500 A100e GPUs?
- pizzafeelsright 3mo agomaybe? ai says about *8.3 days* of continuous runtime on a single GB300 NVL72 rack about a sprint's level of effort.
- aerodexis 3mo agoa very expensive sprint
- BoorishBears 3mo ago> based on FP4 FLOPs (yes it's disadvantageous to A100, that's the point) The A100 doesn't have hardware FP4, and you'd be running a quantized model with some accuracy loss but unless this was natively trained on FP4* * to add another layer, they own the model and could apply tons of post-training techniques to reduce that accuracy loss and probably already do
- ai_fry_ur_brain 3mo ago[flagged]
- BoorishBears 3mo agoI'm pretty sure Altman has spoken about giving a model 100k+ A100s specifically, this might be them being very literal
- laichzeit0 3mo agoSo glad Fable limits just got reset. Thanks OpenAI.
- InsideOutSanta 3mo agoOh hey, thanks for the hint!
- dwa3592 3mo agoThis marketing video on the page is nice!! can't wait for the hardware to get cheaper to live the AI life i wanna live.
- thimabi 3mo agoI’m interested in knowing how each of GPT 5.6’s variants fare in non-English writing/translation tasks. GPT 5.5 has a tendency to write English calques and non-idiomatic prose in other languages. Although that can be somewhat tamed with detailed instructions and a corpus of confusing terms, the model’s output often reads like a literal translation rather than native prose. Since I notice these issues most clearly in languages I know well, it makes me reluctant to trust the model’s output in languages in which I’m less proficient. Ironically, ChatGPT began as a simple text-generation tool, but much of its offerings and benchmarks now focus on coding and agentic workflows, while leaving behind what made it notable in the first place.
- diwank 3mo agoi'm not happy with how openai is trying to pit 5.6 sol as a cheaper equivalent to fable here for one thing, they said that on AA, sol is "within one point of fable" at 58.9 vs 59.9 but don't clarify that the latter is with safeguards where ~8% of the tasks got routed to opus i'm not rooting for either and genuinely think that the token efficiency and cheaper price are important but this sort of thing just feels disingenuous :-/
- mnicky 3mo agoThis is especially interesting because IIRC the AA benchmark is calibrated so that 1 point and greater difference is statistically significant.
- 2001zhaozhao 3mo agoHuh, a good alternative just as anthropic's 50% weekly subscription subsidy is ending this weekend. Time to see if it's benchmaxxed or actually a strong leap over GPT5.5. They also seem to really not care about alignment, or care about it in the wrong way. It's entirely missing in the blogpost and there are some concerning bits in the model card, seemingly treating CoT controllability as something to be "investigated" rather than the warning sign it's supposed to be.
- mnicky 3mo agoThere's also this: > GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated -- https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-the-system-card#METR_Warns_Us__9_1_3_6_ https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-th...
- lukebuehler 3mo agoVery interesting: I wonder if the RL approach is diverging between Anthropic and OAI? I noticed that Fable uses shell tools almost exclusively (even to search and edit files), compared to previous Anthropic models. Having run some experiments with 5.6, I notice that it uses built-in file systems and provider native tools much more (not shell tools), compared to previous OAI models.
- golangdev 3mo agogood alternative to anthropic
- paul7986 3mo agoFor writing GPT which i was subscribed to Fall 2024 to March 2026 (laid off) is superior to Gemini. Been using Gemini since March mostly and they offered a $10 a month plan so i took it. Though today realizing GPT is superior to help me write I am back to being a paying customer. Im in full swing mode to get back into the job market (get the heck away from UI/UX which is now a stupid career in terms of number of jobs out there and in the future there will continue to be less) pivoting into product management (can vibe code anything now) and or customer relations. Hopefully GPT helps me with this pivot and Im again gainfully employed!
- super256 3mo agoOn the tiny voids demo: does your Firefox js thread lock up as well, when you try to interact with it? https://openai.com/index/gpt-5-6/#a-leap-forward-in-design https://openai.com/index/gpt-5-6/#a-leap-forward-in-design
- fisher-brett 3mo agoYep, happens to me on Chrome as well
- gorgmah 3mo agoJust used terra ultra for exactly one prompt in codex and it ate through my full 5h window in about 10mns (20$ plan). The results look pretty good though. Luckily I have had my chatGPT subscription for a while and have a bunch of resets available (nice compared to anthropic). Assuming I take the 5x plan it would give me about an hour of active sessions with terra ultra (maybe ultra is not good value regarding tokens?), not even using Sol yet. Does everyone using codex use the 200$ plan? I normally use the 100$ anthropic plan and barely ever reach the usage limit.
- altcognito 3mo agoDo you know if you used sol/terra/luna?
- jstummbillig 3mo ago> maybe ultra is not good value regarding tokens? Well, yes, as explicitly stated on https://openai.com/index/gpt-5-6/ https://openai.com/index/gpt-5-6/: "ultra goes further by coordinating four agents in parallel by default, trading higher token use for stronger results and faster time-to-result on demanding tasks."
- gorgmah 3mo agothanks, it makes sense, I'll stick to max from now on
- ssl-3 3mo agoI use the $20 plan, but I don't code all day every day. With Codex, it is my experience that I can churn through a 5h window in no time with newer models -- especially when they're new. So I tend to use fancier models for planning, and the less-fancy models for writing code based on that plan. I switch to the fanciest model if any part of this gets stuck. If I've got a something big-ish to work on, I pay attention to the reset timers so I can get more of it done in one chunk. Models seem to slowly get better/relatively less-expensive as they age. (It isn't clear to me if that's because the cost actually goes down, or if the allotment goes up, or if things get more efficient in unseen ways, or what. OpenAI is vague AF about what we get for the $20 that we pay.)
- ls612 3mo agoI think the most interesting part of this is that OpenAI is going way easier on the classifiers than Anthropic. They explicitly state that many defensive cybersecurity uses are supported and implicitly criticize Anthropic's stance on Fable's uses by saying that overblocking cyber requests is itself a major security risk as more AI models continue to advance in intelligence. I have so many questions as to what is going on on a game theoretic level in the AI space in the past two months, it seems like multiple actors have realized their incentives are really quite different than they originally thought.
- goodmattg 3mo agoI flip back and forth between whoever currently has the more powerful frontier model that isn't cost prohibitive - subscriptions only, API pricing a non-starter. Today that's Fable 5 which has been excellent, as soon as it's Sol I'll switch to that. The OAI/Anthropic harness behavior has mostly stabilized for me with consistent AGENTS.md that I sync with CLAUDE.md - I like pi (pi.dev) and have tried to build it up to get performance comparable to the two "first-party" harnesses, I'm just not there yet. One major sticking criteria for not going with OpenCode / pi for all of my coding is I want access to the tier-1 frontier model of the day without API pricing - e.g. afaik I can't use Fable 5 via pi harness even though I have a subscription, so for this week I'm on Claude Code. It's not the need to Fable 5 for everything, but even if I just want the marginal intelligence benefit to stress test an architecture decision, it's a safety blanket to know there isn't a ~smarter~ model I could have used. And for my use cases, the doggedness and capability of these frontier models has been insanely effective. My feeling is we're still in the Uber era subsidy period - the moment the subscriptions either try to lock me in longer than a month or stop OAI/Anthropic stop delivering frontier models in the subscriptions, I'm out - switching fully over to pi.dev or another OS harness and routing my token spend via OpenRouter or offloading to Qwen locally. Then I'll have to put an accurate dollar amount on frontier intelligence.
- 2001zhaozhao 3mo agoI'm working on a multi-harness IDE that supports custom agent workflows and skills that are shared between any harnesses it wraps over. I think it might prove handy for a workflow like yours.
- goodmattg 3mo agoWould it currently support Fable 5 via the restrictions Anthropic is placing on usage... because that's my major blocker
- RhodesianHunter 3mo ago> My feeling is we're still in the Uber era subsidy period I often wonder whether this doesn't continue indefinitely. Uber was able to do this because it was just them and Lyft playing second fiddle, with a huge barrier to entry once the network effects had kicked in. It just seems like the model space has way too many competitors, + OSS/Local options for them to ever be able to jack up their prices. At least once the datacenter bottleneck has been cleared.
- tmaly 3mo agoOne of my best use cases for the short duration I have fable is to use it to create the plan and acceptance test files then use GPT 5.5 Pro to do an adversarial review on the plan then feed that feedback into fable to fix the plan.
- xur17 3mo agoLooks like I have access to gpt-5.6-terra and luna. How does one decide between gpt-5.5 and gpt-5.6-terra? Pricing is similar, but it's hard to tell if it's better..
- sunaookami 3mo agoOverloaded in Codex, no indication if it is already in ChatGPT and I can't use it in the API even though it says it should be available. Typical horrible OpenAI launch. Glad that Anthropic just reset the rate limits so I will go back to Fable again.
- AgentMasterRace 3mo agoI never have have the issues most people talk about ... I feel like most were never Devs before ai and don't know what they actually need done when prompting. that on top of not utilizing good tools such as a codebase indexer, lsp and a project scaffold.
- revolvingthrow 3mo agoBenchmarks look really promising. Suspiciously good, even. I guess we’ll see soon enough. My question to previewers: how are the guardrails for random joe that wasn’t personally blessed by the ai pope to access the non-nerfed model? Fable is a nightmare in this regard, but I’m not sure whether 5.6 also gets a critical side-eye from the gubmint when you ask it to fix bugs in your code (you filthy hacker, you).
- cmrdporcupine 3mo agoI almost immediately ran into "This request requires additional safety checks, which can take extra time. Hang tight or retry with a faster model for a quicker response, though it may be less capable of handling complex requests." Which is something I've never seen with codex before, and I wasn't doing anything funky. Just writing CUDA kernels and benchmarks for them.
- alex0015 3mo agoI was getting that regularly last week with regular 5.5 medium on the plus plan. I was doing benchmarking for a photo editor in Swift.
- cmrdporcupine 3mo agoInteresting. Never seen it before. Now it's just constant and there's nothing sensitive about this work.
- Donald 3mo agoI'm getting the same message doing WebGL shader work.
- cmrdporcupine 3mo agoIt (Sol, on high) does seem actually really quite competent at this work though (GPU programming). Much more so than the attempts I made with 5.5 earlier today.
- simonw 3mo agoHere are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelicans.html https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s https://www.youtube.com/live/Wq45rvPGNHs?t=1070s
- semiquaver 3mo agoThe quality of sol on effort=none makes me think this test is saturated or they are benchmarkmaxxing this exercise.
- yokoprime 3mo agoThank you Simon! Luna is surprisingly decent across all reasoning levels.
- vecter 3mo agoI think all of Luna's are bad. The only decent one is sol @ xhigh. Even sol @ max is weird. Sol @ high and @ medium are ok, and every other single one across every model is bad.
- hn_throwaway_99 3mo agoStrong disagree, but to each their own. For sol I really like how only medium uses the wings on the handlebars to ride the bike. For all the other sols the pelican evolved a new set of arms separate from the wings.
- emehex 3mo agogpt-5.6-sol x XHIGH is my favourite
- throwaw12 3mo agosomething is wrong with Terra model series, most pelicans, except Max, looks bad
- cmrdporcupine 3mo agoAlmost immediately ran into some the kind of gatekeeping I've heard Claude Code users complaining about with Fable. Not sure why, I just had it working on writing benchmarks for some CUDA kernels. Nothing security related: "This request requires additional safety checks, which can take extra time. Hang tight or retry with a faster model for a quicker response, though it may be less capable of handling complex requests." At least it gave me the option of waiting instead of just unceremoniously downgrading me. Appears to be making progress but... weird?
- brcmthrowaway 3mo agoBenchmaxxed
- itvision 3mo agoWeirdly, normally new ChatGPT releases are head and shoulders above anything else, but according to OpenAI's own evaluation, Anthropic's Mythos outperforms ChatGPT in quite a few benchmarks: https://openai.com/index/gpt-5-6/ https://openai.com/index/gpt-5-6/. ChatGPT 6 must be deep in the pipeline and will be released within the next few months. Maybe that's why this release is versioned 5.6, not 6.0.
- sahil87 3mo agoI think its more about branding than anything else. Anthropic played a masterstroke with the way they marketed, released, and then blocked Mythos. Now everyone know the Mythos "model" by name. ChatGPT 6 is trying to follow suite.
- guybedo 3mo agoIt's good to see labs taking into account the cost/task. Grok 4.5 is interesting because it's smart enough at great price. It seems gpt 5.6 is right there with great efficiency and great pricing. Working with Fable has been a great experience, but at the end of the day, if you can get only 10% of your work done because it just burns through tokens, that's not that interesting. I've been mostly using Opus and Fable high for planning and codex 5.5 medium for implementations. Claude is also the only model i can use for design tasks. If gpt 5.6 can finally deliver on the design side, it might be time to ditch the Claude sub and go full Gpt.
- halfmatthalfcat 3mo agoUsing the Claude "superpowers" skill will downgrade models automatically, using Sonnet and Haiku for trivial things.
- bob1029 3mo agoI am seeing some bugginess in testing: Parameter: reasoning_effort Function tools with reasoning_effort are not supported for gpt-5.6-sol in /v1/chat/completions. To use function tools, use /v1/responses or set reasoning_effort to 'none'.' Official OAI .NET library. Even when I override the currently experimental [?] flag to 'none', it will still occasionally throw this error (about 5% of the time). I hope we aren't trying to push customers off the chat completion endpoint... Responses endpoint looks great on paper, but the business wants more visibility and control over the reasoning process than this product currently offers. Edit: This is broken in my VS copilot setup too.
- stillpointlab 3mo agoI can't try it since it hasn't appeared in my Codex yet, but this is is necessary from OpenAI in my opinion. Fable is just so much better at understanding broad context. I only use GPT 5.5 for straight forward easy to describe tasks, and it does crush those. But I spend a lot more time steering Codex towards good design on broad concept type tasks, ones that Fable shows sometimes surprising clarity. I look forward to seeing how it compares once I have access. Not getting tripped by spurious safe guard flags could be an advantage.
- EugeneOZ 3mo agoGPT 5.6 Sol is a token hog. After implementing the task, it started some "reviews" I didn't ask for - they consumed 19.5M and 11.9M tokens, while the task itself was below 5M tokens.
- breatheoften 3mo agoi wish they had renamed chatgpt to codex instead of the other way around ...
- macleginn 3mo agoFor context, I have access to MS Copilot through my workplace. To see what it looks like, I have tried to login through https://copilot.microsoft.com/ https://copilot.microsoft.com/ , where I was informed that my account, although recognised, is not yet supported. However, I can get more or less the same chat window, with access to all the data, through https://m365.cloud.microsoft/ https://m365.cloud.microsoft/ A redirect could have been useful.
- vamsiraju 3mo agoI think 5.6 Sol is only as good as 5.5 or Opus 4.8 in terms of getting its given work done. It just has an uncanny ability to pickup more work that it can tackle next that the older models lack, or have not been trained to do before. Where folks are seeing a difference between working with Fable or 5.6 I think also boils down to this phase shift.
- egorfine 3mo agoZero information on the knowledge cutoff. The model itself responds it's June 2024 which is weird given that GPT-5.5 has knowledge cutoff at August 2025.
- mydreamof 3mo agoBro these colors on chars are unbelievlable, I can not understand which is opus, which is fable, which is GPT...
- senko 3mo agoI love testing the new models by asking them to code a toy RTS game. Here's what Terra did: https://senko.net/vibecode-bench/2026/rts-gpt-5.6-terra.html https://senko.net/vibecode-bench/2026/rts-gpt-5.6-terra.html (one try, in codex app, xhigh effort) Comparing this to other models, I find it similar to GPT-5.5 and a bit behind Sonnet 5. You can see how other models fared here: https://senko.net/vibecode-bench/ https://senko.net/vibecode-bench/ (you can also fetch the prompt and the the 5.6 Terra resulting code on from that page). I don't have access to Sol yet (on a Plus sub, which should get it according to what I've read), so can't do the more interesting test. I'll update the above page as soon as I get access - hopefully soon.
- rsoto2 3mo agoSo the measure of a model is how well they can recreate something they easily have thousands of examples of in their training data. There's probably a better base RTS on github somewhere for free.
- ciefa 3mo ago>There's probably a better base RTS on github somewhere for free. I... I think you are missing the point.
- senko 3mo agoWell, it is a silly test, not a scientific benchmark. However, I would say it is a measure (not the measure). If you look at the entries, there's a lot of variation - definitely not something they memorized outright. And the test itself is deceptively simple. You need to do canvas rendering, there's pathfinding, command queueing, terrain generation, etc. There are some subtle click handler bugs (various LLMs often stumble on those). And I ask the model to do it all in one file, further increasing the complexity of the task. And the result is something that you can instantly evaluate. And if the result is any good, even play! So yeah, I think it's a fair test. I'm sure it'll get saturated at some point. Actually I started with Minesweeper and switched to RTS last December, because Minesweeper was being saturated. I'm expecting (hoping?) the RTS test will last until the end of this year...
- 3mo ago
- celltalk 3mo agoI guess Plus accounts don't get access to Sol? Or is it because I am in Europe?
- HarHarVeryFunny 3mo agoNot specific to OpenAI / Codex, but I'm curious what people are doing to protect themselves from any destructive actions by their coding agents? Just install and pray? Explicity approve all actions? Reconfigure for safety? Run in a sandbox (Docker) ?
- ryan_n 3mo agoI still just explicitly approve all actions and review all code (unless it's a personal/throwaway project no one else will ever touch/use/see). I know a lot of people that run in a sandbox though. That said, I'm sure there are lots of people that just yolo it and hope for the best.
- kennykartman 3mo agoTypically I just want to isolate the agent disallowing it from accessing other parts of the filesystem. Using a different user might be enough, but I typically use [bubblewrap](https://github.com/containers/bubblewrap https://github.com/containers/bubblewrap).
- user43928 3mo agoI use the auto-reviewer for actions outside the builtin sandbox. So far this has been rock solid, and tens of millions of developers use this setup without issue. It is not going to wipe our hard disks. At least I hope so. Fable and GPT 5.6 have been ever more proactive, and GPT 5.6 is automating the AppStore on my machine to download an Xcode update while I am typing this.
- HarHarVeryFunny 3mo agoIs this auto-reviewer part of Codex? Is the review done by the agent or the model?
- user43928 3mo agoYes. In Codex it is called 'Approve for me', in Claude it is 'Auto mode'. I believe in both cases it is prompting a model with a fresh context that is tasked with reviewing the reason for the action. With Claude, I have seen that if the reviewer does reject the proposed action, it responds with a long text about how the Agent should not try to work around this rejection, and instead prompt the user for an explicit approval of the proposed action.
- gverrilla 3mo agoIf Fable is removed from my Anthropic sub, I'll have to change to OpenAI.
- tekacs 3mo agoUnfortunately, I'm finding that in long-form agentic use, when I'm trying to use Sol, I keep tripping guardrails – moreso than even Fable, somehow. I don't know exactly what part of my codebase is triggering it, so I'm going to have to keep poking, but apparently the guardrails are not that gentle despite the phrasing. :(
- ramijames 3mo agoSounds like you are working on something naughty :)
- rsanek 3mo agoBased on the Intelligence vs. Cost graph, not clear to me why anyone would use Terra? Luna looks quite interesting though, happy to see OpenAI still serving the more budget-oriented side of the market (seems like Anthropic and Google have lost interest there). https://artificialanalysis.ai/articles/gpt-5-6-has-landed https://artificialanalysis.ai/articles/gpt-5-6-has-landed
- pstorm 3mo agoCost and intelligence aren't the only axes. Terra has better latency and output speed than Sol for example.
- glaslong 3mo agoLuna@max is in a VERY interesting spot if their rankings are at all to be believed: - Better than Opus4.8 in the coding agent index (doubt) - Just below sonnet 5, even with glm5.2, in the overall intelligence index - Cheaper than haiku4.5, glm5.2 and kimi2.6 on cost per intelligence task index
- simianwords 3mo agoGPT Terra is 50% cheaper than 5.5 while being more performant. So it’s like a straight up 50% reduction in cost! That leads me to a question. Why wouldn’t they just default to terra in ChatGPT in the last few months? If they didn’t then they burnt money for no reason by giving a shittier model at a higher price
- mnicky 3mo ago"while being more performant" ..on some specific set of benchmarks ;)
- 698969 3mo ago50% reduction in cost charged to customers, inference cost may as well be the same, we don't know.
- simianwords 3mo agoRecent reports that OpenAI found optimisations would explain that these are legit
- gorgmah 3mo agoAnyone else noticed the "Extended: Fable 5 is included in your weekly limit through July 12 blablabla" disappeared from claude code? Did they panic-delete the july 12th deadline ?
- atentaten 3mo agoYes, I noticed this too!
- ciefa 3mo agoI still see in the menu to select the model in the GUI (Claude Desktop, claude.ai etc).
- DetroitThrow 3mo agoLooks like they reset everyone's Fable usage.
- user43928 3mo agoThey did. I wonder if Anthropic will also be removing the 50% limit. My Fable weekly limit is at 15% used already, 5.6 Sol at 3% used. And this is with the Max 20x plan compared to Codex 5x. I don't work on the same tasks to compare them objectively, but GPT 5.6 on xhigh seems much cheaper. Essentially unlimited usage.
- zarzavat 3mo agoAnthropic really needs to get Opus 5 out ASAP. The gap between Opus 4.8 and Fable is large enough to drive a GPT 5.6 sized bus through. A better Opus would take some of the heat off.
- user43928 3mo agoProvided such a Opus 5 performs on par with Fable 5 and GPT 5.6 Sol. Otherwise I am not interested. Supposedly Fable 5.1 is in the later stages of the release pipeline, maybe it takes back the crown from OpenAI, who are now rumored to launch GPT 6 in August.
- dayone1 3mo agodoes anyone on chatgpt business plan (not enterprise) not have access to the Sol models in codex? i have 5.6 for terra and luna but not sol
- guybedo 3mo agowe probably need to use gpt sol max to decide which gpt flavor and effort we need to use per task.
- giorgioz 3mo agoOn top of GPT 5.6 Sol they added a Tamagotchi / Clippy mascotte https://x.com/giorgio_zampa/status/2075319657997750495?s=20 https://x.com/giorgio_zampa/status/2075319657997750495?s=20
- reversefleckerl 3mo agoNo, this Pets feature has been around for a while now.
- throw03172019 3mo agoWhat in the world is that? Why.
- deleted 3mo ago[deleted]
- prodmod 3mo agoThings I have been struggling with Fable over and GPT 5.5, were just solved handily by SOL in a real "thank you, next problem" kind of way. Overall, something that just works is way less wasteful for your usage than struggling back and forth for hours.
- ai_fry_ur_brain 3mo ago[flagged]
- brcmthrowaway 3mo agoLove this comment.
- Sabinus 3mo agoThrowaway accounts being aggressive to other users and ranting about AI isn't productive IMO. There are plenty of other places on the internet for that.
- guybedo 3mo agoit seems terra is pretty much useless, you either want luna max for everyday coding (cheaper and same perf as 5.5 high), or sol xhigh/max for demanding tasks
- beaker52 3mo agoWe Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.
- dandaka 3mo agoNot at all, we love them all with Chinese labs. And wish them to continue competing and not winning. That is how we get best models, lower prices and better availability.
- matheusmoreira 3mo agoI openly hope the chinese labs distilling them into open weights win.
- perching_aix 3mo agoTried Xiaomi MiMo v2.5 via opencode today. Since Sonnet 5's release week, Sonnet 4.6 has been feeling like a vegetable, with Sonnet 5 itself being only a little better. MiMo on the other hand feels like Sonnet 4.6 did up until very recently. Absolutely impressive. In some ways, more impressive than GPT 5.5 with high(!) thinking. GPT says quite some nonsense from time to time; didn't see any sign of this in MiMo so far, which is a pretty wild difference.
- w4yai 3mo agoYup. I'm done with US companies. Let's go China !
- jatora 3mo agohave fun with your sub-tier models then. More compute for me
- kouteiheika 3mo ago...until you get rug-pulled (like with Fable recently) because the model you're depending on is proprietary, then you have fun with no model at all. The sooner Chinese labs catch up the better it is for all of us, even if you don't use their models, as they're the ones who define the baseline capability that cannot be taken away from you (no one is going to limit/remove access to an LLM if you can get equivalent/better unrestricted access from Chinese open-weight models).
- XCSme 3mo agoGPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n. Here's my benchmark results for GPT-5.6: https://aibenchy.com/?q=gpt-5.6 https://aibenchy.com/?q=gpt-5.6 (the high reasoning variants are still running, uploading them soon too) EDIT: The high variants are there too, enjoy the hamsters[0]. [0]: https://aibenchy.com/showcase/?q=gpt-5.6 https://aibenchy.com/showcase/?q=gpt-5.6
- lsllc 3mo agoInteresting that Sol (low) did better than Sol (medium) in your benchmark (and is barely more expensive than Terra). I too have been using 5.3 codex as a cheap-but-good model and are switching to Terra (xhigh).
- XCSme 3mo agoHere's all 3 (medium), and GPT-5.5 It GPT-5.6 doesn't seem to be a lot smarter than 5.5, but it is faster, cheaper, more efficient and more consistent: https://aibenchy.com/compare/openai-gpt-5-6-sol-medium/openai-gpt-5-6-terra-medium/openai-gpt-5-6-luna-medium/openai-gpt-5-5-medium/ https://aibenchy.com/compare/openai-gpt-5-6-sol-medium/opena...
- ai_fry_ur_brain 3mo ago[flagged]
- paxys 3mo ago
- raketenkater 3mo ago5.6 sol ultra just nuked my branch and burned my 5h limit. nice work
- yayamao 3mo agogood alternative, while gemini still no news
- alberth 3mo agoMaybe it’s a bug but on iOS individual paid Pro account - I can no longer see which model is being used nor select which model I want.
- noobcoder 3mo agoI hope it isnt like Opus eating so many tokens and taking so much time Really wanna see it in DeepSWE benchmark
- clutter55561 3mo agoI use both Claude and Codex, but mostly Claude for planning and coding, and Codex to review Claude’s work. I follow a sort of waterfall workflow which is verbose but fully transparent. Anthropic’s $100 subscription works fine for me, but whatever subscription my company has with OpenAI reaches the 5hr limit ridiculously quickly.
- Sol- 3mo agoHow do you couple them together efficiently? The nice thing about Codex or Claude is that the delegation or multi agent workflow capabilities are just built-in. Do you link one with the other as a skill or mcp or so?
- killix 3mo ago[flagged]
- rubenflamshep 3mo agoWhen I was going through this it was because OpenAI had defaulted to /fast mode with 2x token usage
- Lucasoato 3mo ago> This page couldn’t load > Reload to try again, or go back. This on iOS, safari
- stanmay 3mo agosol is good
- int3trap 3mo ago5.6 SOL is basically useless, even on fast mode. It takes so long to do anything that it would be faster to do yourself. And it burns usage so quickly it's genuinely not worth it.
- mococa 3mo ago“Be scared”
- ekzy 3mo agoJust my two cents. I'm on the Plus plan, I ask gpt-5.6 sol / high to analyze a vibe-coded codebase (~50k LoC) and write a plan to make it production ready. It wasn't a great prompt, I just wanted to test it quickly. It ran for ~15min and consumed 95% of my 5h quota (I thought it was gonna crash). The output is excellent but just a heads up that it consumes a lot of quota!
- gorgmah 3mo agoYes had the same experience, Sol consumes limits quite fast
- jatora 3mo agoAnyone looking to use frontier SOTA on the $20 plans is going to have a bad time
- desterothx 3mo agoIf you don't have an agent heavy workflow, you'd be surprised how far the 20$ subscription stretches. I used ~500$ of usage in the last month on a team 20$ plan.
- RagnarD 3mo agoAnnoyingly, the new ChatGPT app which folds in Codex, no longer recognizes Shift-Tab to toggle plan mode. Irritatingly you have to enter /plan. OpenAI, fix this!
- isaachinman 3mo agoNoticed this as well. You have to go into keyboard shortcuts and set it manually.
- robertwt7 3mo agoit seems like 5.6 SOL is better at almost everything than Mythos except Coding Benchmarks (except TerminalBench)? anyone knows why Mythos scores so high on SWEBench are they cheating or are they just optimised better for coding?
- ls_stats 3mo agoThe most impressive part is the token efficiency/cost per task of 5.6 Sol, it makes Opus 4.8 and Fable look extremely bad ($1.04 vs $1.80 vs $2.75)[0]. And 5.6 Luna ($0.21) is also impressive, cheaper than GLM 5.2 ($0.37) with higher intelligence. [0]: https://artificialanalysis.ai/#price-and-cost https://artificialanalysis.ai/#price-and-cost
- mnicky 3mo agoWell it's smaller model (something like 4T against 10T Fable). So it's faster and cheaper and with a lot of RL and maybe some favorable benchmark selection it can compete on these scores. In real tasks I expect it to have less intelligence, generalization ability, etc. than Fable.
- djx22 3mo agoNot sure what everyone's experience is but I find 5.6 Sol to be a great liar. Reported success on a half done job and left things in a broken state after having quite a few back & forth followups on the initial prompt to clarify the plan. Didn't experience this with 5.5. Opus 4.7 and below sometimes did it but they fixed it in Opus 4.8. So, overall, the initial experience has made me think that this model will be a lot more stressful to work with just because the level of trust that it actually completes the task is now much much lower.
- ai_fry_ur_brain 3mo agoAre you people seriously this dumb? Have you conwidered that all of these benchmarks are trained into these models. Can you stop sharing them as if they matter?
- drsalt 3mo agothey update these shits too much.
- mlmonkey 3mo agoWhere is Gemini in all this? Lately it's not even been in the running. Sir Demis asleep at the wheel? Or Google too scared to release a SOTA model? Or ... maybe Gemini 4 is too good and the NSA is using it to break into systems worldwide ...?
- w4yai 3mo agoGoogle pretending they're still far ahead. In reality, far far behind. Google has become "just like another" company, nothing special. Few bright minds, with a lot of overpaid engineers.
- r58lf 3mo agoThe rumor is Gemini 3.5 pro will be released next week. It had previously been announced (at Google I/O) to be released in June, but it got pushed back. Why? Rumors range from they are having trouble to they scrapped an old framework to pursue a new one and the new model will surpass Fable 5 and be cheaper.
- sneezychl 3mo agoI expect Google will drop a new SOTA model soon. Rumor has it they're training a new model from the ground-up, which takes a while.
- winwang 3mo agoI find it interesting that no one here has mentioned the increased (usable) context window 258k -> 353k. That's huge, but I wonder if it means we pay long context (2x) for the ones past 272k still.
- joerawr 3mo agoI really appreciate the focus on intelligence WITH token efficiency. I'd like to see that become the trend. Smartest per token metrics. Least tokens to accomplish the task above a certain success level. Most of my tasks would benefit from efficiency / token, but switching models constantly, and trying to guess the right model and effort level takes up too much of my processing.
- tw1984 3mo ago> I really appreciate the focus on intelligence WITH token efficiency that is a polite way of saying "I don't believe AGI is coming anytime soon".
- znnajdla 3mo agoNot sure why you would think a focus on efficiency means less performance. Compression is intelligence. Higher efficiency enables higher performance.
- e1ghtSpace 3mo agoseems more like "this priorisies AGI" What is AGI to you though?
- inerte 3mo agoI honestly interpreted and agreed with the version "this saves money".
- y1n0 3mo agoAnybody have an idea of what the flops per token generated is on a SOTA model like current GPT/OPUS? Is it basically the parameter count? So something like GLM-5.2 is, at a minimum, ~744 GFLOPs per token generated? Am I way off base? Seems astronomical.
- sidgarimella 3mo agoI've found Sol's propensity for delegating to subagents can make it... disastrously expensive, especially with each subagent having some implicit floor on further reasoning/context gathering before action. The base model is certainly cheaper and more token efficient etc, but on large tasks cost in some way is now n^2
- winwang 3mo agoHow bad is it for you? Are you on ultra or xhigh/max? I typically ask it (5.5, now 5.6-sol) to use subagents for specific things anyway. On the Pro 20x plan, I'm seeing like ~1% usage per 20-30 min per session (on max effort), which is in line with 5.5. Currently trying out ultra on a personal project, feels like ~3x more expensive per unit time. (No idea on quality yet, for obvious reasons.)
- sidgarimella 3mo agoMax via Cursor, should've mentioned, very possible I'm seeing more Cursor than OAI. All the subagents it picked were also Sol Max... I've seen 7-10 in one turn. Regardless the top reasoning tiers being conflated with subagent count feels double edged
- alasano 3mo agoMy global CLAUDE.md/AGENTS.md etc have strict rules to never use subagents without approval or when specifically requested except for whitelisted scenarios. I think that's the way to go most of the time
- mafesio 3mo agoI have the codex x20 plan also. Set it to 5.6 sol ultra, gave it one task, in less than 10 mins I got the 10% warning. I did have it do a few things earlier, just some anaylsis and then create a pdf from it. I thought maybe that took a little longer than I though. Used a reset, it went for about 11 minutes and then, just out of usage popped up, no warning, there were still 4 agents running, the nice thing is it did let them finish, each one went for about 10 minutes. I also have a claude max plan, I have been using Fable 5 on ultra, I never hit the session limit, and get 3 or 4 full day's looping on ultra. I don't know how it handles the subagents, but claude does it much more efficiently, Codex does seem much faster, so maybe it's just a relativity thing.
- edg5000 3mo agoIt scores high on BenchCAD, that's interesting to see, I was wondering about how each model could handle this. Seems like they trained it on programmatic CAD specifically.
- andrijaskontra 3mo agoTrying to play those games has really bad impact on PC performance
- jingw222 3mo agonobody cares anymore
- fmind-dev 3mo agoFrom my first tests today, it is a workhouse. It can scan my whole code base, optimize every part, with a greater level of autonomy than other tools. This is insane, we are living at the best time.
- ai_fry_ur_brain 3mo agoThe literal worst time. I prefer meritocracies.
- cognitiveinline 3mo agoYou think aristrocrats are using GPT and succeeding?
- ai_fry_ur_brain 3mo agoNo I think aristocrats are using llms to justify the flattening of wages, where they can use it as an excuse diminish the value of merit based systems.. Meritocracy / education was one of the few ways someone from a lower class couod climb the ladder and they're trying to destroy that ladder.
- cognitiveinline 3mo agoThen that's a tirade against capitalism - don't direct it at AI. If you don't like the situation, push for socialist policies.
- ai_fry_ur_brain 3mo ago[dead]
- treovchinn 3mo ago> average daily output tokens per active researcher were more than twice the highest level observed for GPT‑5.5.
- thomas_witt 3mo agoI would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?
- _bobm 3mo agoare people getting the `<!-- -->` sentinel'd reasoning summaries?
- epolanski 3mo agoI'm using luna (the smallest), at low thinking in my 9-to-5 job and I'm quite happy. No groundbreaking tasks so far, but typical small jira issues and fixes are done in a matter of low minutes. Very fast loops have their pros. Fable or Opus would wander and wander.
- shabgzer 3mo agoThey talk a lot about speed in the article, but having tried out Sol today with Pi, 'medium' mode, one thing that stands out is that it's really ssslllloooowww. It also defaults to 'low' mode for some reason. Can't tell if that's a step backwards compared to GPT-5.5 in medium mode so I'm sticking to medium. Edit: just noticed it's spawning subagents in 'high' thinking mode.
- nurettin 3mo ago[dead]
- cryptokent 3mo ago[flagged]
- claud_ia 3mo ago[flagged]
- imilev 3mo agoCannot believe I needed a VPN to the US, to open this from Switzerland... At least give me the article ffs.
- wkjagt 3mo agoI just watched the video on their launch page and I am really not sure how I feel about it. On one side, it's cool that these people get to start businesses and stuff using ChatGPT (assuming these are true stories), but how much of the business is really them? And how much does this business rely on a chat bot always being present as a kind of know it all employee? Maybe I'm just being naive or old fashioned (haven't really used AI much), but seeing these two people who started a cereal business for example talking to their laptop as if they're talking to a human advisor makes me feel, I don't know, I find it creepy. By the way, this isn't about their 5.6 version in particular I guess, it's just the first time I've looked at one of their videos.
- camillomiller 3mo agoI would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an average day in the past. It's frankly depressing, and it's even more depressing thinking that most people seem to absolutely disregard this feeling completely.
- infecto 3mo agoInversely I am building things I either never would have gotten to and a rate that would have taken infinitely longer. I get excited every day because instead of so much of my day spent writing the boiler plate I can spend most of the time around the architecture.
- ed_elliott_asc 3mo agoI don’t get this feeling, I feel like I’ve achieved so much more than I have ever achieved, everything is polished and I’m happy with it. I’m still developing, I’m just doing more than I ever did by directing Codex. The way I see it is the same as I saw the leap from writing code in a text editor, to using an ide with intellisense, to using the jetbrains ide’s, to using mcp’s, to now directing AI - at all of those steps I wrote code, each step less and less but still it has the same output which is it is my work - even writing in a text editor I wrote less Java (until enterprise architects got involved :) ) than C++, and than assembly.
- pompomsheep 3mo agoLogged into the OpenAI platform this morning and had to double check they hadn't pivoted into a crypto company with these new names
- hniscensorship 3mo ago[flagged]
- maryjeiel 3mo ago[dead]
- matheusmoreira 3mo ago> Your subscription to Claude Max has been successfully canceled. Switching next month. Looking forward to working with Sol.
- pimeys 3mo agoI've been testing Sol/Terra/Luna now since yesterday, running complex evals on all of them and I feel a bit... mixed on how they perform. The eval is an agent that runs a set of tools and a prompt we can tune separately for different models. The OpenAI version of the prompt was specifically tuned based on their guide[0]. Then we let Opus to run another agent that acts as a user, trying to solve a problem (anonymized and taken from production). The problem is complex and we don't expect it to be solved by these agents, but we measure how the agents operate when faced with a vague problem: - Opus 4.8 and GLM 5.2 both identified a constraint sooner and stopped so the user can fix an issue first that the agent cannot solve. - Sol tried hard to solve the issue with different tools, burning tokens, until finally reached to the same conclusion with Opus and GLM. It was two times more expensive compared to Opus and six times more expensive to GLM for this task. - Terra went even further and started calling tools that would not solve the issue, burning tokens and failing. - Luna repeated the same failing tool call until it hit the round limit, and burned more money than Opus. I'm kind of puzzled with the new GPT. Like, yes Sol is OK for programming, but I was expecting to get a cheap agentic model for non-programming tasks, one that can detect if things go awry and correct. Terra is too expensive and Luna not really fit for the task. Sonnet 5 is a bit better but more expensive than Opus 4.8, which is still the best in my evals. GLM 5.2 is extremely good if you can define the task and the tools clearly for it, and costs pennies! [0] https://developers.openai.com/api/docs/guides/latest-model https://developers.openai.com/api/docs/guides/latest-model
- MarvinYork 3mo agoI thought I was an OpenAI fanboy, but version 5.6 isn’t for me. Sol Ultra just keeps working and checking, and working and checking again, but it can’t even correct minor errors that aren’t a problem for 5.5 xhigh. I’ve rolled Codex back to 5.5 for now.
- friendly_chap 3mo agoThe problem is they nerfed 5.5 about a month ago. The change was immediately visible to me: context compacting started to happen about 3x as frequently. I think 5.6 is still not even close to 5.5 xhigh pre-nerf.
- alberth 3mo agoGPT-5.6 is the first model where I’ve actually frustrated to use it. I’m explicitly telling it to do something extremely specific and it’s just not listening to me. Eg, I gave it an image to update. The image is sized 400x200 pixels. It then generates a new image at 300x300. I explicitly state to be 400x200 in size and it won’t listen.
- ddp26 3mo agoIs it possible GPT-5.6 is not a very aligned model?
- emrehan 3mo agoI am looking to rent an apartment in a new residential tower. I have asked Fable and Sol to scrap the listings from various sources, deduplicate them and present them as a web application. Just using the cowork/(ex-)codex application interfaces. Fable had issues with the sourcing and organizing images, and shoot itself at foot looking for shortcuts as usual. As I was getting it fix these back and forth, I copied my prompt and gave it to Sol. Sol has surpassed my expectations by far. With a one shot simple prompt on a complex task, it gave me a working web app with everything I want with minor issues to track and fix.
- patates 3mo agodevops monster! crazy how intelligently it debugs/solves any devops problem I could throw at it!
- fomoz 3mo ago5.6 Sol High Fast is using more capacity than 5.5 High Fast, I hit the 5h limit for the first time. Other than that, I think the difference between 5.5 and 5.6 will be the same as 5.4 and 5.5. 5.5 is just less frustrating to use, although not perfect and still has derp moments. But a lot less than 5.4. So I expect 5.6 Sol to be smoother to use. But so far it just feels slower. We'll see.
- gordonhart 3mo agoFirst impression of 5.6 Sol in Codex is fantastic — the model asks dozens of clarifying questions before starting to implement where other models (including 5.5 and Terra) just yolo it with assumptions that needed to be walked back later.
- angelhadjiev 3mo ago[dead]
- internet101010 3mo agoI have Fable send the specs/plans it comes up with to GPT for review and in 2/5 cases yesterday it found additional 1-2 bugs while in the process of reviewing. GPT-5.6 didn't try to fix the bugs (as instructed) but it did surface them, which is something that didn't happen with GPT-5.5. When spec/plans approved Fable sends back to GPT-5.6 for ralph implementation and it seems to be an even faster, more reliable workhorse than it already was in GPT-5.5. Overall, impressed. Will continue to be a core piece of my workflow.
- WarmWash 3mo agoStill fails my internal test of counting legs on animals who have had extra legs photoshopped in. However if prompted to determine what is wrong with the image, it does get it right. This kind of "out of bounds" image analysis seems to be a very difficult problem to solve, but totally necessary for transformers to really bring about massive change.
- Dfol 3mo agoIt's working great for me. I generally prefer OpenAI. I usually start projects with Codex; I love the plan we created. Then, after about 2 hours of working, back and forth, etc.I realize it's drifting HARD, or getting stuck on relatively simple things. Once I get frustrated enough, I sometimes start over with Anthropic. Anthropic has been better (for me at least) to work with from start to finish. I have the $200/month plan for each of them. (And I used to have the $250/month plan for Gemini... lol?) Both Fable and Sol are good and definite improvements. I don't have a favorite yet. It's too early to tell. The biggest difference I'm noticing is usage. Anthropic is like a used car salesman. Not allowed to do a basic check on your own car (website), they'll scrape every penny possible out of you. Low limits, high api prices, trying to keep everything in their system. OpenAI is like a cool but aloof dad. Go ahead and borrow his car, shoot, he'll even pay for your gas every once in a while. He'll answer your 5th question drunk, forgetting what you asked. But, at least he's a nice drinker that buys the group shots. I have to watch my Fable usage, and I'm sure as hell not going to pay API prices. I don't have to watch/worry about my Sol usage even in Ultra. But, 1 thing I'm noticing using Sol Ultra (on fast mode) is that it's slowing WAY down after a bit. I work the opposite of peak times so that's not it.
- DefineOutside 3mo agoI've played around GPT 5.6 sol high at both work and home. At work, it was able to one shot a dashboard. Of course, my prompts are vague as I'm not exactly sure what I want yet, but it did a better job than I could do as a backend dev forced to work on frontend sometimes. Usage is also great, it just feels so much more efficient than older models in terms of thinking and time. Cost is barely better though. It can burn a million tokens in less than a minute, at least at launch where there's likely less load on the servers. At home, it feels like I'm fighting the AI less while letting it refactor code. I'm glad that I left this 12,000 line vibe coded port of a hand written codebase to future models to refactor. It feels like the model has better judgement than old models that would destroy your codebase so long as it meant accomplishing your prompt. I'm almost disappointed that it's this good.
- cmrdporcupine 3mo agoAfter some time with it... It has a tendency to do things without asking, a trait I'd associated more with the Claude Opus & Sonnet models than with Codex & GPT in the past. Specifically I've seen it go and update e.g. README.md files filling it with recenty-biased gibberish that means nothing to the user (e.g. very specific technical notes related to what it was currently working on) or staging and adding design/spec documents that were meant to just be working documents. In general it tends to behave more aggressively with git, if you let it get its hands on it. It has stronger "opinions" on that stuff, that don't always agree with me. I'm going to have to update my prompts, I think. But I'm not used to this kind of thing in Codex, which in the past has been much more explicit and cautious, and one of the reasons I've preferred it over Claude. It is very "smart." It also has a tendency to yak-shave things. Producing huge volumes of correctness and regression tests and nitting over e.g. very minor variances. One thing that is "entertaining" is letting two separate instances review each other's code. They will endlessly find things to nit at.
- hugepan 3mo ago[flagged]
- hdemirev 3mo agoHave been testing Luna against other small models for production customer support uses cases, finding negligible performance impact when comparing against the substantial cost increase. Claims re: reduced token use also don't seem reproducible. Wrote a quick blog with some sample findings (marketing content but the findings are real): https://valiopt.com/blog/gpt-5-6-customer-support-cost-performance https://valiopt.com/blog/gpt-5-6-customer-support-cost-perfo...
- arendtio 3mo agoI also like to take a look at https://cursor.com/cursorbench https://cursor.com/cursorbench While in the past months Composer 2.5 was a lot better than I had expected a year ago, and the GPT 5.6 family does a good job in terms of cost for performance, I wonder why nobody is talking about Grok 4.5 high? Those numbers look very convincing to me.
- m0rde 3mo ago> * Grok 4.5 has an advantage on CursorBench: an earlier snapshot of the Cursor codebase was unintentionally included in training. The exact score impact is unclear. That data has been removed for future models. For a rundown of third-party benchmark scores, see the Grok 4.5 launch blog. I don't know about those numbers, even assuming this was by mistake :)
- SoAp9035 3mo agoI use GPT 5.6 Sol medium as daily driver and the GPT 5.6 Sol XHigh is for planning. Also I really like the GPT 5.6 Luna Max it really goes hard too and cheap~.
- childintime 3mo agoDang, very few people share what kind of work they are actually do with the model, and/or the language they code in. That would make all the difference.