8 ms·
OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
GPT-5.5 - https://news.ycombinator.com/item?id=47879092 https://news.ycombinator.com/item?id=47879092 - April 2026 (1010 comments)
- throw03172019 5mo agoFaster than anticipated because of Deepseek release?
- brianbest101 5mo ago[dead]
- swyx 5mo agomore like they wanted to release it yesterday but merely had some last min flags they wanted to hold off for
- Jhonwilson 5mo agook not bad
- deleted 5mo ago[deleted]
- m3kw9 5mo agoMaybe but no one serious is using deepseek
- XCSme 5mo agoDoubt it, DeepSeek v4 is quite underwhelming.
- pants2 5mo agoIs anyone here actually using pro models through the API? I'd be very curious what the use-case is.
- ComputerGuru 5mo agoYes? The same reason you would use it via the tooling.
- chadash 5mo agoYes. High value work where cost (mostly) doesn't matter. For example, if I need to look over a legal doc for possible mistakes (part of a workflow i have), it doesn't matter (in my case) whether it costs $0.01 or $10.00, since it's a somewhat infrequent event. So i'll pay $9.99 more, even if the model is only slightly better.
- freedomben 5mo agoIndeed, even just Terms of Service and Privacy Policy work. Infrequent enough that cost isn't an issue, but model quality absolutely is
- bogtog 5mo agoI'm surprised I never heard people talking about using -Pro variants, even though their rates ($125-175/M?) aren't drastically larger than old Opus ($75/M), which people seemed to use
- sigmoid10 5mo agoHuh. Yesterday they said: >API deployments require different safeguards and we are working closely with partners and customers on the safety and security requirements for serving it at scale. And now this. I guess one day counts as "very soon." But I wonder what that meant for these safeguards and security requirements.
- embedding-shape 5mo agoThe same person who've mercilessly lied about safety is still running the company, so not sure why anyone would expect any different from them moving forward. Previous example: > In 2023, the company was preparing to release its GPT-4 Turbo model. As Sutskever details in the memos, Altman apparently told Murati that the model didn’t need safety approval, citing the company’s general counsel, Jason Kwon. But when she asked Kwon, over Slack, he replied, “ugh . . . confused where sam got that impression.” Lots of cases where Altman hass not been entirely forthcoming about how important (or not) safety is for OpenAI. https://www.newyorker.com/magazine/2026/04/13/sam-altman-may-control-our-future-can-he-be-trusted https://www.newyorker.com/magazine/2026/04/13/sam-altman-may... (https://archive.is/a2vqW https://archive.is/a2vqW)
- simonw 5mo agoI wonder if the fact that GPT-5.5 was already available in their Codex-specific API which they had explicitly told people they were allowed to use for other purposes - https://simonwillison.net/2026/Apr/23/gpt-5-5/#the-openclaw-backdoor https://simonwillison.net/2026/Apr/23/gpt-5-5/#the-openclaw-... - accelerated this release!
- FINDarkside 5mo agoWhen stuff is delayed due to "safeguards" it just means they don't think they have the compute to release it right now.
- redsaber 5mo agonot available for Github Copilot pro(only in pro+, business and enterprise), I am really now feeling the era of subsidized AI is over.
- skeledrew 5mo agoThis is where the emigration to Chinese providers begins.
- sunaookami 5mo agoWith a 7.5x multiplier and even that is a promo!! Microsoft is insane! https://github.blog/changelog/2026-04-24-gpt-5-5-is-generally-available-for-github-copilot/ https://github.blog/changelog/2026-04-24-gpt-5-5-is-generall...
- rvnx 5mo agoVery bad habit these safeguards. These "safety" filters are counter-productive and even can be dangerous. In my place for example, a lot of doctors are using ChatGPT both to search diagnosis and communicate with non-English speaking patients. Even yourself, when you want to learn about one disease, about some real-world threats, some statistics, self-defense techniques, etc. Otherwise it's like blocking Wikipedia for the reason that using that knowledge you can do harmful stuff or read things that may change your mind. Freedom to read about things is good.
- timedude 5mo agoYup, deliberately making the model retarded
- NicuCalcea 5mo ago> a lot of doctors are using ChatGPT both to search diagnosis and communicate with non-English speaking patients I think that's the problem. Who's going to claim responsibility when ChatGPT hallucinates or mistranslates a patient's diagnosis and they die? For OpenAI, this would at best be a PR nightmare, so that's why they have safeguards.
- hellohello2 5mo agoThe doctor would be responsible. I had a choice better a doctor that used AI or not, I would much prefer one that did...
- NicuCalcea 5mo agoThe doctor would be responsible for the accuracy of their translation tool, something they can't verify but you expect them to use?
- rvnx 5mo agoWhat's the alternative then ? -> You are in China, you go to emergency, nobody speaks your language Move hands ? DeepSeek is better than using hands, even Baidu Translate, ChatGPT or whatever you find. Other solutions are theoretically nice on paper but almost delusional. An imperfect solution is better than no solution. == Similarly, a deaf-person is theorically better with a certified interpreter that can talk with the hands, but they may prefer voice-recognition software or AI tools. (or... talking with hands is more confusing and annoying or less understandable for them). Of course ChatGPT transcription can have issues, but that's the difference between the real-world and Silicon Valley's disconnected lawyers world. == If ChatGPT says: "sorry I won't be able, please go to see a licensed interpreter, good luck!" then it's just OpenAI trying to save their asses, at your risk/expense. If you have a choice, you can make the choice, and you can double-check what is said. In other cases, you have no choice, nothing to check, only problems but no hints of solutions. This is why openness is important.
- czk 5mo agoAPI page lists the knowledge cutoff as Dec 01, 2025 but when prompting the model it says June 2024. Knowledge cutoff: 2024-06 Current date: 2026-04-24 You are an AI assistant accessed via an API.
- htrp 5mo agoCan you really believe things that the model says? (A lot of prior model api pages say knowledge cutoffs of June 2024, maybe the model picks that up?)
- czk 5mo agoyou cant but its pretty reproducible across api and codex and other agents so i just thought it was odd. full text it gives: Knowledge cutoff: 2024-06 Current date: 2026-04-24 You are an AI assistant accessed via an API. # Desired oververbosity for the final answer (not analysis): 5 An oververbosity of 1 means the model should respond using only the minimal content necessary to satisfy the request, using concise phrasing and avoiding extra detail or explanation." An oververbosity of 10 means the model should provide maximally detailed, thorough responses with context, explanations, and possibly multiple examples." The desired oververbosity should be treated only as a *default*. Defer to any user or developer requirements regarding response length, if present.
- deleted 5mo ago[deleted]
- swyx 5mo agocan u test it on say who won the 2024 US election
- ghurtado 5mo agoI can't really think of a less reliable test for anything at all than making a random guess as to something that had about 50/50 odds to begin with Easiest Turing test ever...
- neosat 5mo agoEnterprise user here and still seeing only 5.4. Yesterday's announcement said that it will take a few hours to roll out to everybody. OpenAI needs better GTM to set the right expectations.
- deleted 5mo ago[deleted]
- neosat 5mo agoJust refreshed and see 5.5 now - yay! Love the speedy resolution ;) Thanks folks, I'll complain faster next time....
- gigatexal 5mo agowhat's the real world comparison to opus 4.7 fellow coders?
- Sembiance 5mo agoI gave 4.6, 4.7 and GPT 5.5 the same prompt and task to reverse engineer a collection of sample vector files from an obscure Amiga CAD program and create a detailed txt specification and a python converter that converts to SVG and produce a report so I can visually verify. 4.6 did very well. 90% perfect on first try, got to 100% with just a few followups. 4.7 failed horribly. First produced garbage output and claimed it was done, admitted it did that when called out, proceeded to work at it a lot longer and then IT GAVE UP. GPT 5.5 codex was shockingly good. Achieved 90% perfect on first try in about a fourth of the time. Got to 100% faster and with fewer follow-ups. I’m impressed.
- gigatexal 5mo agoInteresting that 4.7 failed like that. Seems 5.5 is impressive but is oh so expensive. Would be interesting if you ran your same test with Deepseek v4 and some of the other Chinese models.
- Sembiance 5mo agoJust tried with DeepSeek V4 Pro with OpenCode. It didn't do great. First attempt produced somewhat correct drawings for some of the original samples, but most were just a spaghetti messs of lines. Some prodding got it to do a little better, but still not right. A third prod and it went down a wild rabbit hole and was much worse. I gave up. I also tried GLM 5.1, it's first attempt was such a disaster I didn't bother working with it any further. It also took by far the longest and wasted a bunch of time/tokens trying to find other converters online (and failing) instead of just reverse engineering the format from the sample files given.
- gigatexal 5mo agoInteresting. I would love your test but for code. If I were to forgo my claude subscription for a Chinese cloud hosted model or local models running on my own hardware I'd use them mostly for code. the thing is I've tried to come up with a good test my own and spend countless time just tweaking it instead of saying this is good enough and benchmarking.
- theo_park87 5mo ago[dead]
- Jhonwilson 5mo agothat is great news
- pillefitz 5mo agoPlease consider the ethical aspects of giving money to OpenAI versus alternatives.
- AlexCoventry 5mo agoYou need to be more specific. OpenAI's commitment to assist the Trump administration with domestic mass surveillance seems to have been largely memory-holed.
- pillefitz 5mo agoYou're right, unfortunately. How naive of me to think that at least the HN audience would care.
- wincy 5mo agoJust tried it out for a prod issue was experiencing. Claude never does this sort of thing, I had it write an update statement after doing some troubleshooting, and I said “okay let’s write this in a transaction with a rollback” and GPT-5.5 gave me the old “okay, BEGIN TRAN; -- put the query here commit; I feel like I haven’t had to prod a model to actually do what I told it to in awhile so that was a shock. I guess that it does use fewer tokens that way, just annoying when I’m paying for the “cutting edge” model to have it be lazy on me like that. This is in Cursor the model popped up and so I tried it out from the model selector.
- syspec 5mo agoCan't tell if above is good or bad.
- wincy 5mo agoI mean, I was doing triage, so wanted an immediate fix. The actual issue is we’re getting some exploding complexity when double checking the action the API is taking is valid in the data. So that needs to be refactored. I suppose it reduces token usage, but Claude Opus will happily do exactly what I want it to.
- XCSme 5mo agoI feel like the last 2-3 generations of models (after gpt-5.3-codex) didn't really improve much, just changed stuff around and making different tradeoffs.
- pixel_popping 5mo agoI disagree, it improved enormously especially at staying consistent for long-tasks, I have a task running for 32 days (400M+ tokens) via Codex and that's only since gpt-5.4
- lowdude 5mo agoThat’s actually crazy, what kind of task is that? And is that a recurring kind of task like some analysis, or coding related?
- guilamu 5mo agoJust tested it on my homemade Wordpress+GravityForms benchmark and it's one of the worst model of the leaderboard performance wise and the worst value wise: https://github.com/guilamu/llms-wordpress-plugin-benchmark https://github.com/guilamu/llms-wordpress-plugin-benchmark I know it's only on a single benchmark, but I dont understand how it can be so bad...
- ac29 5mo agoYour benchmark has Opus 4.7 performing significantly worse than Sonnet 4.6. Even if true on your benchmark, that is not representative of the overall performance of the models.
- guilamu 5mo agoYes Opus 4.7 fast (no reasoning) did a worst job than Sonnet 4.6 high (with reasoning) according to Gemini 3.1 Pro evaluation.
- ac29 5mo agoYour table doesn't indicate reasoning vs non-reasoning, or reasoning level
- guilamu 5mo agoWhen nothing is noted it's max reasoning (xhigh in copilot chat in vscode if available). The models not availble on copilot were tested through opencode (max reasoning) and deepseek v4 was tested through Cline (with max reasoning too).
- mosselman 5mo agoYou even traveled in time to deliver us this benchmark. I really like this benchmarking. Have you evaluated the judge benchmark somehow? I'd love to setup my own similar benchmark.
- 5mo ago
- ftonon 5mo agoLooks like the default config in the chat is instant 5.3, it only uses the 5.5 on the thinking variant
- bnm04 5mo agoThey moved a few months ago to have separate instant and thinking models. 5.3 is the latest instant, and 5.5 is a reasoning model.
- _pdp_ 5mo agoA very expensive model for API usage. Fine in codex I think.
- benjiro3000 5mo ago[dead]
- QuadrupleA 5mo agoExactly double the cost of GPT 5.4 - $5 per MTok input, $0.50 cached, $30 output. All the AI players definitely seem to be trying to claw more money out of their users at the moment.
- guilamu 5mo agohttps://openrouter.ai/openai/gpt-5.5-pro https://openrouter.ai/openai/gpt-5.5-pro 30/180 usd on Openrouter. Did I miss something?
- languid-photic 5mo agoI think that's Pro. Regular 5.5 is 2x regular 5.4.
- languid-photic 5mo agoIt's 2x/token, but for default reasoning we've found GPT-5.5 uses fewer tokens overall, so net cheaper on median. [1] (Note, that stops being true at higher reasoning levels, where our observed total cost goes up ~2-3x.) [1] https://x.com/voratiq/status/2047737190323769488?s=20 https://x.com/voratiq/status/2047737190323769488?s=20
- XCSme 5mo agoGPT 5.5 is close to Opus 4.7, but at 7x the cost[0]... Either Opus 4.7 miscounts reasoning tokens, or it's A LOT more efficient than GPT 5.5 I thought they made GPT 5.5 more token efficient than 5.4, but it uses 2x the reasoning tokens. [0]: https://aibenchy.com/compare/openai-gpt-5-5-medium/openai-gpt-5-4-medium/google-gemini-3-1-pro-preview-medium/anthropic-claude-opus-4-7-medium/ https://aibenchy.com/compare/openai-gpt-5-5-medium/openai-gp...
- zerof1l 5mo agoI don't see any meaningful performance improvements in those paid models anymore. They all roughly produce junior developer-level code, continue to have mental breakdowns in their “thinking” stage, occasionally hallucinate things, delete pieces of code/docs they don’t understand or don’t like, use 1.5 times the necessary words to explain things when generating docs and so on. I'm now testing "avoid sycophancy, keep details short and focus on the facts" in my AGENTS.md files.
- podnami 5mo agoThis is snark. Since when has a junior level dev managed to debug and deploy say a cloudformation stack and follow up with notes under 3 minutes?
- gjsman-1000 5mo agoHeard this analogy elsewhere, but worth repeating: AI is like having the greatest developer who ever lived, but she is always on 4 beers.
- nathan-hello 5mo agopersonifying ai is incredibly cringe no matter how weird your comparison is
- gjsman-1000 5mo agoImagine a drunk developer. Sparks of brilliance while missing obvious trees.
- nathan-hello 5mo agoweird.
- xXSLAYERXx 5mo agoI know of a publicly traded company which in its early years was built on beer. Literally. 3 guys in a co-working space in Cambridge, MA. Beer fueled their progress. 15 years later the software is still the backbone of the org.
- refulgentis 5mo agoI'm absolutely stunned by what I've seen from 5.5. I thought it'd be a nothingburger and ~= Opus. Gave it two very long-running problems I haven't had the courage to work on in the last 2.5 years, solved each within an hour. - An incremental streaming JSON decoder that can optionally take a list of keys to stop decoding after. 1800 LOC about 30 minutes later, and now my local-first apps first sync time is 0.8s instead of 75s when there's 1.5 GB of data locally. - Flutter Web can compile to WASM and then render via Skia WASM. I've been getting odd crashes during rapid animation for months. In an hour, it got Skia WASM checked out, building locally, a Flutter test script, and root caused the issue to text shadows and font glyphs (technically, not solved yet, I want to get to the point we have Skia / Flutter patch(es)) If you told me a week ago that an LLM could do either of these, without heavy guidance, I'd be stunned. And I regularly push them to limits, ex. one of Opus' last projects was a tolerant JSON decoder, and it ended up being 8% faster than the one built-in to Dart/Flutter, which has plenty of love and attention. (we're cheating a little, that's why it's faster. TL;DR: LLMs will emit control characters in JSON and that's fine for me, treating them as fine means file edit error rates go from ~2% to 0%) I just wish it was cheaper, but, don't we all...
- Topfi 5mo agoPricing by context length: Input: $5/M tokens at <=272K, $10/M tokens above 272K. Output: $30/M tokens at <=272K, $45/M tokens above 272K. Cache read: $0.50/M tokens at <=272K, $1/M tokens above 272K. Significantly more expensive than Opus 4.7 beyond 272K and at least in my tasks, I haven't seen the model that much more token efficient, certainly not to such a degree that it'd compensate this difference. GPT-5.4 had a solid context window at 400k with reliable compaction, both appear somewhat regressed, though still to early to truly say whether compaction is less reliable. Also, I have found frontend output to still skew towards that one very distinct, easily noticeable, card laden, bluesy hue overindulged template that made me skeptical of Horizon Alpha/Beta pre GPT-5s release. Ended up doing amazing at the time for task adherence, which made it very useful for me outside that one major deficit. The fact that GPT-5.5 is still so restricted in that area is weird considering it's supposed to be an entirely new foundation.
- robertwt7 5mo agoGpt 5.5 combined with codex is really good. I actually have no doubt whenever I asked questions, plan, or implement a code with it. With opus 4.7, I have to keep double checking because it doesnt follow the CLAUDE.md instruction, it hallucinates a lot, by default it makes things up when it can’t find the answer to something. Its crazy how quickly people are saying that OpenAI is left behind last year when they declared code red and look at where we are now
- willj 5mo ago[dead]
- woohin 5mo ago[dead]
- jvidalv 5mo agoIn the only one that feels that OpenAI has bots/commenters on payroll on all this kind of news downplaying Claude and stating how much better Codex is? There is too much and there are too many, and some of their takes don’t fly if you use Claude daily.
- Aboutplants 5mo agoOf course they do. As do all of the other companies pushing their product these days.
- AlexCoventry 5mo agoYeah, it's eerie, same with how everyone seems to have forgotten that OpenAI betrayed democracy by committing to work on unsupervised autonomous weapons and domestic mass surveillance.
- int3trap 5mo agoHonestly I find comments like yours much more eerie. By all accounts they never agreed to any of that but you say it with such confidence like it's a fact.
- AlexCoventry 5mo agoThe Trump administration's handling of Anthropic showed that regardless of what the contract or the law says or means, they will severely penalize any vendor who refuses their demands. And OpenAI stepped right into that relationship immediately after the administration showed that. So either they were signing up for a supply-chain risk designation and whatever other punishments the Trump administration dreams up, or they're complying. If this sounds crazy to you, though, I'd like to know, and understand why. I miss ChatGPT/Codex.
- int3trap 5mo ago> regardless of what the contract or the law says That is not really established. The Anthropic issue was specifically about DoD use and Anthropic's military use restrictions. What the Trump admin did was bad and coercive but its not proof that contract terms and law are irrelevant. For instance, why not just use eminent domain if they don't care about contracts and want whatever they want? > either they were signing up for a supply-chain risk designation and whatever other punishments the Trump administration dreams up, or they're complying Couldn't OpenAI have negotiated different terms, accepted a narrower scope, or drawn different red lines? Their public DoD terms still exclude things like mass domestic surveillance and autonomous weapons outside human control. Do you not believe that or believe it doesn't matter at all? Either of those is problematic to the conclusions that follow from them. I also think the whole argument implies something about Anthropic's position that's not as clean in reality. NSA is already using Mythos despite the Pentagon dispute, and Anthropic is still talking to the administration. Trump even said they were "shaping up" recently. Isn't it also a possibility that one company negotiated poorly and took a position of perceived moral authority that Trump et al threw a hissy fit over and over reacted to? That's happened countless times with this admin and is far more likely in my opinion given Anthropic hasn't cut all ties and continues to try and work out a contract. I wholeheartedly agree the current administration is dangerous. I just don't think the conclusion "OpenAI must be complying with the same demands Anthropic refused" follows from what we've seen. And I think there are plenty of other far more plausible conclusions to draw from the events.
- gertlabs 5mo agoComprehensive coding reasoning benchmark results for GPT 5.5 with max reasoning are up at https://gertlabs.com/ https://gertlabs.com/ Live decision and heavier agentic evals will continue being uploaded for 24 hours but I don't expect its leaderboard position to change at this point. GPT 5.5 is the most intelligent public model. And significantly faster than its predecessor.
- nl 5mo agoSecond model to get 25/25 on my benchmark (after Opus 4.7): https://sql-benchmark.nicklothian.com/?highlight=openai_gpt-5.5 https://sql-benchmark.nicklothian.com/?highlight=openai_gpt-... Cheaper and slower than Opus.
- jedisct1 5mo ago$ uvx swival --provider chatgpt --model gpt-5.5 APIError: ChatgptException Ok, still not available everywhere apparently :(
- croemer 5mo agoIt refuses to write bioinformatics code that involves analysis of SARS-CoV-2. Even when it's totally obvious I'm not trying to do any bioengineering of any sorts. Totally harmless stuff I'm doing and I just get rejected.