4 ms·
Yes o3 was an inflection point. Both modes of 5 are performing poorly on the IQ test compared to o3 https://www.trackingai.org/home https://www.trackingai.org/h
by CompoundEyes 1y ago
Yes o3 was an inflection point. Both modes of 5 are performing poorly on the IQ test compared to o3 https://www.trackingai.org/home https://www.trackingai.org/home That test best reflects my experience and results from practical use cases with the reasoning models when planning specs, bug finding, ideation, and deep research. It’s great at tool use as well in scripts. What I like least about the release is no transparency about the “routing” taking place. Give me all the options on the system card to pick from https://openai.com/index/gpt-5-system-card/ https://openai.com/index/gpt-5-system-card/ and I don’t want to have to start telling it “ultrathink” or other magic words to affect routing. To be fair though I haven’t tried 5 in reasoning mode beyond Cursor. But now I see o3 is only part of the Pro plan. If 5 reasoning is supposed to be better why would o3 and o3-pro still be a specialized models for Pro customers? I’d like to see some side by side prompts I might go back and test that.