4 ms·
I understand that things are moving fast and all, but surely the.. 8? models which are currently available is a bit .. overwhelming for users that just want to
by mmsc 1y ago
I understand that things are moving fast and all, but surely the.. 8? models which are currently available is a bit .. overwhelming for users that just want to get answers to their questions of life? What's the end goal with having so many models available?
- Osyris 1y agoThis is a much more expensive model to run and is only available to users who pay the most. I don't see an issue. However, the "plus" plan absolutely could use some trimming.
- djrj477dhsnv 1y agoIf it's better (and newer) than gpt4, it shouldn't have a lower version number.
- bachittle 1y agofree users don't have this model selector, and probably don't care which model they get so 4o is good enough. paid users at 20$/month get more models which are better, like o3. paid users at 200$/month get the best models that are also costing OpenAI the most money, like o3-pro. I think they plan to unify them with GPT-5.
- stavros 1y agoThat doesn't help much when we're asymptotically approaching GPT-5. We're probably going to be at GPT-4.9999 soon.
- rfw300 1y agoNot necessarily true. GPT-4.1 was released after GPT-4.5-preview. Next model might be GPT-3.7.
- nikcub 1y agoI'd be curious what proportion of paid users ever switch models. I'd guess < 10%
- CamperBob2 1y agoI switch to o1-pro on occasion, but it is slow enough that I don't use it as much as some of the others. It is a reasonably-effective last resort when I'm not getting the answer quality that I think should be achievable. It's the best available reasoning model from any provider by a noticeable margin. Sounds like o3-pro is even slower, which is fine as long as it's better. o4-mini-high is my usual go-to model if I need something better than the default GPT4-du jour. I don't see much point in the others and don't understand why they remain available. If o3-pro really is consistently better, it will move o1-pro into that category for me.
- CuriouslyC 1y agoIf you're not at least switching from 4o to 4.1 you're doing it wrong.
- motoxpro 1y ago4o is better than 4.1 for a lot of things that are non-coding/general research.
- blurbleblurble 1y agoOverwhelming yet pretty underwhelming
- nickysielicki 1y agoI just can’t believe nobody at the company has enough courage to tell their leadership that their naming scheme is completely stupid and insane. Four is greater than three, and so four should be better than three. The point of a name is to describe something so that you don’t confuse your users, not to be cute.
- browningstreet 1y agoAt Techcrunch AI last week, the OpenAI guy started his presentation by acknowledging that OpenAI knows their naming is a problem and they're working on it, but it won't be fixed immediately.
- moomin 1y agoI know they have a deep relationship with Microsoft, but perhaps they shouldn’t have used Microsoft’s product naming department.
- simonw 1y agoSam Altman has said the same thing on Twitter a few times. https://x.com/sama/status/1911906570835022319 https://x.com/sama/status/1911906570835022319 > how about we fix our model naming by this summer and everyone gets a few more months to make fun of us (which we very much deserve) until then?
- deleted 1y ago[deleted]
- levocardia 1y agoThere's a humorous version of Poe's law that says "any sufficiently genuine attempt to explain the differences between OpenAI's models is indistinguishable from parody"
- paxys 1y ago> users that just want to get answers to their questions of life Those users go to chat.openai.com (or download the app), type text in the box and click send.
- AtlasBarfed 1y agoI'd like one to do my test use case: Port unix-sed from c to java with a full test suite and all options supported. Somewhere between "it answers questions of life" and "it beats PhDs at math questions", I'd like to see one LLM take this, IMO, rather "pure" language task and succeeed. It is complicated, but it isn't complex. It's string operations with a deep but not that deep expression system and flag set. It is well-described and documented on the internet, and presumably training sets. It is succinctly described as a problem that virtually all computer coders would understand what it entailed if it were assigned to them. It is drudgerous, showing the opportunity for LLMs to show how they would improve true productivity. GPT fails to do anything other than the most basic substitute operations. Claude was only slightly better, but to its detriment hallucinated massive amounts and made fake passing test cases that didn't even test the code. The reaction I get to this test is ambivalence, but IMO if LLMs could help port entire software packages between languages with similar feature sets (aside from Turing Completeness), then software cross-use would explode, and maybe we could port "vulnerable" code to "safe" Rust en masse. I get it, it's not what they are chasing customer-wise. They want to write (in n-gate terms) webcrap.
- CamperBob2 1y agoHow does the latest Gemini 2.5 Pro Ultra Flash Max Hemi XLT release do on that task? It obviously demands a massive context window.
- AtlasBarfed 1y agoI'll check once I get the nitrous tanks and the aftermarket turbos overnighted from Japan arrive.
- nipah 1y agoI have a very simple question with like, 5 lines at best, that basically no model, neither reasoning or simpler could grasp. For obvious reasons I'm not disclosing it here (because I fear data contamination in the long run), but it basically breaks the "reasoning" of those things. Unfortunately, I still can't try the o3-pro because the API version is not easily available, and I'm certainly not willing to pay for it in pro mode, but when it comes to the plus version (if it comes) I'll try. To this date, because of this question (and similar ones) I stand very unimpressed with those models, the marketing is a thousand times larger than reality, and I suspect people in general are surprisingly less capable of detecting intelligence than they think. The normal o3 also managed to break 3 isolated installations of linux I was trying it with, a few days ago. The task was very simple, simply setup ubuntu with btrfs, timeshift and grub-btrfs and it managed to fail every single time (even when searching the web), so it was not impressive either.
- resters 1y agoModels are used for actual tasks where predictable behavior is a benefit. Models are also used on cutting-edge tasks where smarter/better outputs are highly valued. Some applications value speed and so a new, smaller/cheaper model can be just right. I think the naming scheme is just fine and is very straightforward to anyone who pays the slightest bit of attention.