8 ms·
Notes on OpenAI o3-mini
- maxdo 2y agoHow would you rate it against Claude ? Didn’t test it yet, but o1 pro didn’t perform as good
- pants2 2y agoI've been trying out o3 mini in Cursor today, it seems "smarter" but overall tends to overthink things and if it's not provided with perfect context it's prone to hallucinate. Overall I prefer Sonnet still. It has a certain magic of always making reasonable assumptions and finding simple solutions.
- firecall 2y agoAs n occasions user and fan of Cursor, it would be good if they could explain what the models are and why the different models exist. There’s no obvious answer of why one should switch to any of them!
- conception 2y agoI don’t think there’s an obvious answer. Try them out and see which works better for your use case.
- maeil 2y agoAgreed that Sonnet still feels like the best all-round model. The new ones are at least on par with it for pure coding, or exceed it (r1, o1 both do IME) but don't generalize as well, especially to tasks with subjective answers. I find the latest Gemini 2.0-Flash-thinking to be closest to Sonnet on those.
- zyklu5 2y agoClaude is still better in my opinion. There's a suite of code-related tasks -- covering a diversity of areas, including dev ops, media manipulation etc., derived from issues I have faced over the years -- I perform for every new release. No model has solved the set of issues solved in one go but Claude still remains the best. An example of the sort of problems in the suite: > I have a special problematically encoded mp4 file with a subtle issue (something I ran into a couple of years ago while fixing a bug in a computer vision pipeline). In the question prompt I also pass the output of ffprobe and ask for the ffmpeg command that'll fix it. Only Claude has figured the real underlying issue out (after 4 interactions).
- kamikazeturtles 2y agoThere's a huge price difference between o3-mini and o1 ($4.40 vs $60 per million output tokens), what trade-offs in performance would justify such a large price gap? Are there specific use cases where o1's higher cost is justified anymore?
- zamadatix 2y agoNot really, it'll also be replaced by a newer o3 series model in short order.
- benatkin 2y ago> Are there specific use cases where o1's higher cost is justified anymore? Long tail stuff perhaps. Most stuff doesn't resemble a programming benchmark. A newer model thrives despite being small when there is a lot of training data, and with programming benchmarks, like with chess, there is a lot of training data, in part because high quality training data can be synthesized.
- arthurcolle 2y agoits the same thing as: gpt-3.5 -> gpt-4 (gpt-4-32k premium) "omni" announced (multimodal fusion, initial promise of gpt-4o, but cost effectively distilled down with additional multimodal aspects) gpt-4o-mini -> gpt-4o (multimodal, realtime) gpt-4o + "reasoning" exposed via tools in ChatGPT (you can see it in export formats) -> "o" series o1 -> o1 premium / o1-mini (equivalent of gpt-4 "god model" becoming basis for lots of other stuff) o1-pro-mode, o1-premium, o1-mini, somewhere in that is the "o1-2024-12-17" model with not streaming, function calling, and structured outputs and vision now, distilled o1-pro-mode probably is o3-mini and o3-mini-high-mode (the naming is becoming just as bad as android) its the repeat, take model, scale it up, run evals, detect innefficiencies, retrain, scale, distill, see what's not working. when you find a good little zone in the efficiency frontier, release it with a cool name
- anticensor 2y agoNo, o3-mini is a distillation of (not-yet-released) o3, not a distillation of o1.
- brianbest101 2y agoOpen AI really needs to work on their naming conventions for these things.
- benatkin 2y agoIt's all based on omni which to me has weird religious connotations. It just occurred to me to put it together with sama's other project, scanning everyone's eyes. That's one aspect of omniscience - keeping track of every soul. Another thing it seems similar to is how Jeff Bezos registered relentless.com. There seems to be a gap between the ideal branding from the perspective of the creators and branding that makes sense to consumers.
- xnx 2y agoHasn't Gemini pricing been lower than this (or even free) for awhile? https://ai.google.dev/pricing https://ai.google.dev/pricing
- BinRoo 2y agoAre you insinuating Gemini is similar in performance to o3-mini?
- gerdesj 2y agoAre you implying it isn't? (evidence please, everyone)
- BinRoo 2y agoSimple example: o3-mini-high gets this [1] right, whereas Gemini 2.0 Flash 01-21 gets it wrong. [1] https://chatgpt.com/share/679d9579-5bb8-8008-ac4a-38cef65b45b5 https://chatgpt.com/share/679d9579-5bb8-8008-ac4a-38cef65b45...
- xnx 2y agoGreat example. Thank you. Can confirm that none of the Gemini models warned about the exception without prompting.
- maeil 2y agoThis agrees with my limited testing so far, but in a different way: o3 being better at coding and objective tasks, with the most recent Flash 2.0-thinking stronger at subjective tasks. Similarly, o3 seems better at shorter output sizes, but drops off, tending to be lazy.
- xnx 2y agoDefinitely varies by application, but the blind "taste test" vibes are very good for Gemini: https://lmarena.ai/?leaderboard https://lmarena.ai/?leaderboard
- celdon25 2y ago[flagged]
- joshuanapoli 2y agoProgrammers have always made a living by automating ourselves out of business. Somehow, we're still doing pretty well.
- celdon25 2y ago[flagged]
- kingkongjaffa 2y agoI find this comment and most of your others very distasteful. Digging through peoples past posts to look for gotcha’s is a pretty crass practice. It’s the kind of thing people used to do to try and dunk on each other on Reddit. It should not become the norm here.
- celdon25 2y agoI think there is a misunderstanding here. I wasn't looking to dunk on or shame anyone, only to make an important point. A lot of software engineers are not doing great right now, even though many here may be doing fine. That's a problem bigger than anyone's ego. And I didn't go back that far, spent about 10 seconds looking. As for my other comments, at least 80-90% of them have a positive vote ratio, even though I deliberately voice unpopular opinions. Most of the negatives have to do with politics, or Elon Musk, which people here tend to feel strongly about. Considering that heated debates are not unwelcome here, I don't think that makes me a terrible person. That being said I'll try to be a little more careful before posting in the future, and I broadly agree with the point you were aiming for with your comment. Apologies to the GP if I caused any unwelcome feelings.
- marxisttemp 2y agoIn what way would you say doctors have largely failed Western society lately?
- tkgally 2y agoAt the end of his post, Simon mentions translation between human languages. While maybe not directly related to token limits, I just did a test in which both R1 and o3-mini got worse at translation in the latter half of a long text. I ran the test on Perplexity Pro, which hosts DeepSeek R1 in the U.S. and which has just added o3-mini as well. The text was a speech I translated a month ago from Japanese to English, preceded by a long prompt specifying the speech’s purpose and audience and the sort of style I wanted. (I am a professional Japanese-English translator with nearly four decades of experience. I have been testing and using LLMs for translation since early 2023.) An initial comparison of the output suggested that, while R1 didn’t seem bad, o3-mini produced a writing style closer to what I asked for in the prompt—smoother and more natural English. But then I noticed that the output length was 5,855 characters for R1, 9,052 characters for o3-mini, and 11,021 characters for my own polished version. Comparing the three translations side-by-side with the original Japanese, I discovered that R1 had omitted entire paragraphs toward the end of the speech, and that o3-mini had switched to a strange abbreviated style (using slashes instead of “and” between noun phrases, for example) toward the end as well. The vanilla versions of ChatGPT, Claude, and Gemini that I ran the same prompt and text through a month ago had had none of those problems.
- nycdatasci 2y agoThis is a great anecdote and I hope others can learn from it. R1, o1, and o3-mini work best on problems that have a “correct” answer (as in code that passes unit tests, or math problems). If multiple professional translators are given the same document to translate, is there a single correct translation?
- jakevoytko 2y agoMy wife is a professional translator and both revises others' work and gets revised. Based on numerous anecdotes from her, I can promise you that "single correct translation" does not exist.
- tkgally 2y agoNo. People’s tastes and judgments vary too much. One fundamental area of disagreement is how closely a translation should reflect the content and structure of the original text versus how smooth and natural it should sound in the target language. With languages like Japanese or Chinese translated into English, for example, the vocabulary, grammar, and rhetoric can be very different between the languages. A close literal translation will usually seem awkward or even strange in English. To make the English seem natural, often you have to depart from what the original text says. Most translators will agree that where to aim on that spectrum should be based on the type of text and the reason for translating it, but they will still disagree about specific word choices. And there are genres for which there is no consensus at all about which approach is best. I have heard heated exchanges between literary scholars about whether or not translations of novels should reflect the original as closely as possible out of respect for the author and the author’s cultural context, even if that means the translation seems awkward and difficult to understand to a casual reader. The ideal, of course, would be translations that are both accurate and natural, but it can be very hard to strike that balance. One way LLMs have been helping me is to suggest multiple rewordings of sentences and paragraphs. Many of their suggestions are no good, but often enough they include wordings that I recognize are better in both fidelity and naturalness compared to what I can come up with on my own.
- johngalt2600 2y agoSo far ive been impressed.. seems to be in the same ballpark as r1 and claude for coding. I will have to gather more samples.. in this past week ive changed from using 100% claude exclusively (since 3.5) to hitting all the big boys: claude, r1, 4o (o3 now), and gemini flash. Then ill do a new chat that includes all of their generated solutions for additional context for a refactored final solution. R1 has upped the ante so Im hoping we continue to get more updates rapidly... they are getting quite good
- submeta 2y ago> The model accepts up to 200,000 tokens of input, an improvement on GPT-4o’s 128,000. So finally ChatGPT catches up with Claude which has a 200,000 token input limit ever since. Claude with its projects feature is my go to tool for working on projects that I have to work on for weeks and months. Now I see a possible alternative.
- lysecret 2y agoHadn’t had much luck with o3. One thing that came to my mind with these test time compute models is that they have a tendency to “overthink” and “overcomplicate” things this is just a feeling for now but has anyone done some study on this? E.g. potentially degrading performance in simpler questions for these types of models?
- deleted 2y ago[deleted]