6 ms·
OpenAI's o1-pro now available via API
- bakugo 2y ago> $150/Mtok input, $600/Mtok output What use case could possibly justify this price?
- icelancer 2y agoFull file refactoring. But I just use the webUI for this and will continue to at these prices... probably.
- serjester 2y agoSynthetic data generation. You can have a really powerful, expensive model create evals so you can tune a faster, cheaper system with similar performance.
- jsheard 2y agoYou could do that, but OpenAI specifically doesn't want you to: https://openai.com/policies/row-terms-of-use/ https://openai.com/policies/row-terms-of-use/ What you cannot do. You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not: Use Output to develop models that compete with OpenAI. Presumably you run the risk of getting banned if they realize what you're doing.
- serjester 2y agoSynthetic data is just as useful for building app layers evals. Probably significantly cheaper ways to get the data if you’re training your own model.
- echelon 2y agoScrew their TOS. OpenAI trained on the world's data. Data they didn't license. Anyone should be able to "rip them off" and copy their capabilities on the cheap.
- jsheard 2y agoThe irony isn't lost on me, but irony isn't going to stop them kicking you off their platform if they feel like it.
- littlestymaar 2y agoI suspect they have no way to enforce that without risking false positive hurting their rich customers (and their business).
- levocardia 2y agoI wonder if some of the high pricing is specifically an attempt to ward off this sort of "slow distillation" of a powerful model
- andyferris 2y ago> You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not: Use Output to develop models that compete with OpenAI. This reads as if they consider developing models that compete with OpenAI as illegal, harmful or abusive. Which is crazy. (The other dot points in their list in the linked terms seem better).
- kelseyfrog 2y agoI compete with AI, not my models.
- SJC_Hacker 2y agoIf it was possible: 1) Why wasn't OpenAI doing it themselves? 2) This means we've reached technological singularity if AI models can improve themselves (as in getting a smarter model, not just compressing existing ones like Deepseek)
- sgillen 2y agoCompressing existing models is exactly what we are talking about here.
- SJC_Hacker 2y agoThen why wasn't OpenAI doing that themselves ?
- nickthegreek 2y agoWho says they arent?
- throwup238 2y agoIt’s not a singularity because the synthetic data generated by the previous frontier model isn’t usually fed directly into the training for the next frontier model - a “discriminator” is applied to select only the highest quality responses. That discriminator could be a field expert or mechanical turk or another model trained to select for higher quality responses (i.e. trained on a dataset of books rather than internet content). As far as I know, OpenAI has been doing this, using both experts and Kenyan workers as well as their own discriminator models. Unfiltered synthetic data is generally used more for distilled models and fine tunes for a specific use case.
- refulgentis 2y agoIt enables obscene unnatural things at a fraction of most SWE hourly rates. One win that jumps to mind was writing a complete implementation of a Windows PCM player, as a flutter plugin, with some unique design properties and emergent API behavior that it needed to replicate from existing iOS/Android code
- zipy124 2y agoDoes it really? Your average software engineer is like £20-30 an hour, for the cost of 1m output tokens you can get a Dev for a full week.
- sheepscreek 2y agoThe math doesn’t check out. A day maybe. Also it’s not just about a placeholder dev. The person needs to know your use-case and have the tech chops to deliver successfully in that timeframe. Now to have that delivered to you in less than an hour? That’s a huge win.
- zipy124 2y agoNo, even contractors who are more expensive per day can be had for £200-300 per diem. Employee costs you are looking at closer to £165 including national insurance and pension on a standard £35k a year salary. Even an employee on £50k a year is only £241.97 a day. IR35 rules largely closed down a lot of the high-value contracting business and pushed day rates down heavily.
- refulgentis 2y agoSure. So why are you saying we'd get a dev for a week for $600? :) Even we take that at face value, now we're claiming devs are $30K/year. (This was sort of baseline full-time pay when I was in high school, 20 years ago) Why do I need 1,000,000 tokens for 500 loc? I don't think further futzing with the numbers makes this work, it's off by multiple OOMs while barely credible
- davidbarker 2y agoPricing: $150 / 1M input tokens, $600 / 1M output tokens. (Not a typo.) Very expensive, but I've been using it with my ChatGPT Pro subscription and it's remarkably capable. I'll give it 100,000 token codebases and it'll find nuanced bugs I completely overlooked. (Now I almost feel bad considering the API price vs. the price I pay for the subscription.)
- hooloovoo_zoo 2y agoIs your prompt {$codebase} find bugs?
- davidbarker 2y agoTypically something like: Look carefully through my codebase and identify any bugs/issues, or refactors that could improve it. <codebase> … </codebase> Doesn't have to be anything overly complicated to get good results. It also does well if you give it a git diff.
- ionwake 2y agoSorry if this is a noob question, but are you just pasting file strings inbetween those tags? like the contents of file1.js and file2.js?
- pridkett 2y agoRepomix can take care of this for you. I pack it, cat the file to my clipboard with pbcopy, and just paste it into the prompt. https://github.com/yamadashy/repomix https://github.com/yamadashy/repomix
- diggan 2y agoI tried the web demo (https://repomix.com/ https://repomix.com/) and it seems to generate unnecessarily complex "packs" for no reason, probably hurts LLM performance too. Why is there "Usage Guidelines" and "File Format" explanations in this, when it's supposed to just be the code "packed"? Better to just have the contents+filename, it'll infer that its directory structure and everything else.
- serjester 2y agoAssuming a highly motivated office worker spends 6 hours per day listening or speaking, at a salary of $160k per year, that works out to a cost of ≈$10k per 1M tokens. OpenAI is now within an order of magnitude of a highly skilled humans with their frontier model pricing. o3 pro may change this but at the same time I don’t think they would have shipped this if o3 was right around the corner.
- danpalmer 2y agoIf you start paying someone and give them some onboarding docs, to a first approximation they'll start doing the job and you'll get value. If you attach a credit card to o3 and give it some onboarding docs, it'll give you a nice summary of your onboarding docs that you didn't need. We're a long way from a model doing arbitrary roles. Currently at the very minimum, you need a competent office worker to run the model, filter its output through their judgement, and act on it.
- levocardia 2y agoRight, value per token is much more important (but harder to quantify). A medical AI that could provide a one-paragraph diagnosis and treatment plan for rare / untreatable diseases could be generating thousands of dollars of value per token. Meanwhile, Claude has probably racked up millions of tokens wandering around Mt. Moon aimlessly.
- _pdp_ 2y agoAt first I thought, great, we can add it now to our platform. Now that I have seen the price, I am hesitant enabling the model for the majority of users (except rich enterprises) as they will most certainly shoot themselves in the foot.
- danpalmer 2y ago> they will most certainly shoot themselves in the foot ...and then ask you for a refund or service credit.
- danpalmer 2y agoIt has a 2023 knowledge cut-off, and 200k context window... ? That's pretty underwhelming.
- gkoberger 2y agoOn the flip side, the cutoff date probably makes it a lot more upbeat.
- throw310822 2y agoDon't know if it's me, but this is really funny.
- bearjaws 2y agoFor a second I was like "2023 isn't that bad"... and then I realized we're well into 2025...
- NoahZuniga 2y agoSeems underwhelming when openai's best model, o3, was demoed almost 4 months ago.
- EcommerceFlow 2y agoo1-pro still holds up to every other release, including Grok 3 think and Claude 3.7 think (haven't tried Max out though), and that's over 3 months ago, practically an eternity in Ai time. Ironic since I was getting ready to cancel my Pro subscription, but 4.5 is too nice for non-coding/math tasks. God I can't wait for o3 pro.
- sheepscreek 2y ago4.5 works on Plus! I know. I was surprised too.
- Tiberium 2y ago"Max" as in "Claude 3.7 Sonnet MAX" is apparently Cursor-specific marketing - by default they don't use all the context of the model and set the thinking budget to a lower value than the maximum allowed. So essentially it's the exact same 3.7 Sonnet model, just with different settings.
- ilrwbwrkhv 2y agoDeepseek r1 is much better than this.
- nsoonhui 2y agoInteresting take, care to explain more exactly how it is much better?
- flippyhead 2y agoIt's exactly "much" better!
- simonw 2y agoThis is their first model to only be available via the new Responses API - if you have code that uses Chat Completions you'll need to upgrade to Responses in order to support this. Could take me a while to add support for it to my LLM tool: https://github.com/simonw/llm/issues/839 https://github.com/simonw/llm/issues/839
- icelancer 2y agoOh interesting. I thought they were going to have forward compatibility with Completions. Apparently not.
- dtagames 2y agoIt does. There are two endpoints. Eventually, all new models will only be in the new endpoint. The data interfaces are compatible.
- deleted 2y ago[deleted]
- dtagames 2y agoIt shouldn't be too bad. The responses API accepts the same basic interface as the chat completion one.
- Tiberium 2y agoEven the basic interface is different, actually - "input" vs "messages", no "max_completion_tokens" nor "max_tokens". That said, changing those things is quite easy.
- simonw 2y agoThe harder bit is the streaming response format - that's changed a bunch, and my tool supports both streaming and non-streaming for both Python sync and async IO - so there are four different cases I need to consider.
- dtagames 2y ago
- WiSaGaN 2y agoI have always suspected that the o1-Pro is some kind of workflow on the o1 model. Is it possible that it dispatches to say 8 instances of o1 then do some type of aggregation over the results?
- deleted 2y ago[deleted]
- simonw 2y agoIt cost me 94 cents to render a pelican riding a bicycle SVG with this one! Notes and SVG output here: https://simonwillison.net/2025/Mar/19/o1-pro/ https://simonwillison.net/2025/Mar/19/o1-pro/
- mateus1 2y agoI’m no expert but that does not look like a 94c pelican to me.
- deciduously 2y agoBetter than my svg pelican would be, but it's a low bar.
- orzig 2y agoAt this point you’d come out ahead just buying a pelican. Even before the tax benefits.
- prawn 2y agoI have been using ChatGPT to generate 3d models by pasting output into OpenSCAD. Often feels like coaching someone wearing a blindfold, but it can sometimes kick things forward quickly for low effort.
- qingcharles 2y agoWhenever you experience a new pelican I always have to check it against your past pelicans to see progress towards the Artificial Super Pelican Singularity: https://simonwillison.net/tags/pelican-riding-a-bicycle/ https://simonwillison.net/tags/pelican-riding-a-bicycle/
- jascination 2y agoYour collection of pelicans is so bloody funny, genuinely brightened my day. I don't know what I was expecting when I clicked the link but it definitely wasn't this: https://simonwillison.net/tags/pelican-riding-a-bicycle/ https://simonwillison.net/tags/pelican-riding-a-bicycle/
- katherineingram 2y ago[dead]
- deleted 2y ago[deleted]
- ein0p 2y agoDid not know it was that expensive to run. I'm going to use it more in my Pro subscription now. I frankly do not notice a huge difference between o1 Pro and o3-mini-high - both fail on the fairly straightforward practical problems I give them.
- jwpapi 2y agoThose that have tested it and liked it. I feel very confident with Sonnet 3.7 right now,if I would wish for something its it to be faster. Most of the problems I’m facing are like execution problems I just want AI to do it faster than me coding everything on my own. To me it seems like o1-pro would be to be used as a switch-in tool or to double-check your codebase, than a constant coding assistant? (Even with lower price), as I assume I would need to get done a tremendous amount of work including domain knowledge done to come up for the 10x more speed (estimated) of Sonnet?
- CamperBob2 2y agoo1-pro can be very useful but it's ridiculously slow. If you find yourself wishing Sonnet 3.7 was faster, you really won't like o1-pro. I pay for it and will probably keep doing so, but I find that I use it only as a last resort.
- irthomasthomas 2y agoo1-pro doesn't support streaming, so it's reasonable to assume that they doing some kind of best-of-n type technique to search over multiple answers. I think you can probably get similar results for a much lower price using llm-consortium. This lets you prompt as many models as you can afford and then chooses or synthesises the best response from all of them. And it can loop until a confidence threshold is reached.
- liu9950 2y ago[dead]