5 ms·
If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20 https://x.com/OpenAI/status/2
by twtw99 7mo ago
If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20 https://x.com/OpenAI/status/2029620619743219811?s=20
- chabes 7mo agoDefinitely don’t want to click in at x either.
- thejarren 7mo agoSolution https://xcancel.com/OpenAI/status/2029620619743219811?s=20 https://xcancel.com/OpenAI/status/2029620619743219811?s=20
- anonym00se1 7mo agoDitto, but I did anyways and enjoyed that OpenAI doesn't include the dogwater that is Grok on their scorecard.
- observationist 7mo ago[flagged]
- Sabinus 7mo agoGet a redirect plugin and set it up to send you to xcancel instead of Twitter. I've done it, and it's very convenient.
- karmasimida 7mo agoIt is a bigger model, confirmed
- Aboutplants 7mo agoIt seems that all frontier models are basically roughly even at this point. One may be slightly better for certain things but in general I think we are approaching a real level playing field field in terms of ability.
- thewebguyd 7mo agoKind of reinforces that a model is not a moat. Products, not models, are what's going to determine who gets to stay in business or not.
- observationist 7mo agoBenchmarks don't capture a lot - relative response times, vibes, what unmeasured capabilities are jagged and which are smooth, etc. I find there's a lot of difference between models - there are things which Grok is better than ChatGPT for that the benchmarks get inverted, and vice versa. There's also the UI and tools at hand - ChatGPT image gen is just straight up better, but Grok Imagine does better videos, and is faster. Gemini and Claude also have their strengths, apparently Claude handles real world software better, but with the extended context and improvements to Codex, ChatGPT might end up taking the lead there as well. I don't think the linear scoring on some of the things being measured is quite applicable in the ways that they're being used, either - a 1% increase for a given benchmark could mean a 50% capabilities jump relative to a human skill level. If this rate of progress is steady, though, this year is gonna be crazy.
- bigyabai 7mo ago> If this rate of progress is steady, though, this year is gonna be crazy. Do you want to make any concrete predictions of what we'll see at this pace? It feels like we're reaching the end of the S-curve, at least to me.
- 7mo ago
- swingboy 7mo agoWhy do so many people in the comments want 4o so bad?
- embedding-shape 7mo agoSomeone correct me if I'm wrong, but seemingly a lot of the people who found a "love interest" in LLMs seems to have preferred 4o for some reason. There was a lot of loud voices about that in the subreddit r/MyBoyfriendIsAI when it initially went away.
- astrange 7mo agoThey have AI psychosis and think it's their boyfriend. The 5.x series have terrible writing styles, which is one way to cut down on sycophancy.
- baq 7mo agoSomebody on Twitter used Claude code to connect… toys… as mcps to Claude chat. We’ve seen nothing yet.
- mikkupikku 7mo agoMy computer ethics teacher was obsessed with 'teledildonics' 30 years ago. There's nothing new under the sun.
- 7mo ago
- dom96 7mo agoWhy do none of the benchmarks test for hallucinations?
- netule 7mo ago[flagged]
- tedsanders 7mo agoIn the text, we did share one hallucination benchmark: Claim-level errors fell by 33% and responses with an error fell by 18%, on a set of error-prone ChatGPT prompts we collected (though of course the rate will vary a lot across different types of prompts). Hallucinations are the #1 problem with language models and we are working hard to keep bringing the rate down. (I work at OpenAI.)
- MarcFrame 7mo agohow does 5.4-thinking have a lower FrontierMath score than 5.4-pro?
- nico1207 7mo agoWell 5.4-pro is the more expensive and more advanced version of 5.4-thinking so why wouldn't it?
- nimchimpsky 7mo ago[dead]
- bicx 7mo agoThat last benchmark seemed like an impressive leg up against Opus until I saw the sneaky footnote that it was actually a Sonnet result. Why even include it then, other than hoping people don't notice?
- conradkay 7mo agoSonnet was pretty close to (or better than) Opus in a lot of benchmarks, I don't think it's a big deal
- jitl 7mo agowat
- 0123456789ABCDE 7mo agomaybe gp's use of the word "lots" is unwarranted https://artificialanalysis.ai https://artificialanalysis.ai indicates that sonnect 4.6 beats opus 4.6 on GDPval-AA, Terminal-Bench Hard, AA Long context Reasoning, IFBench. see: https://artificialanalysis.ai/?models=claude-sonnet-4-6%2Cclaude-sonnet-4-6-adaptive%2Cclaude-sonnet-4-6-non-reasoning-low-effort%2Cclaude-opus-4-6-adaptive%2Cclaude-opus-4-6 https://artificialanalysis.ai/?models=claude-sonnet-4-6%2Ccl...
- conradkay 7mo agoI was basing it off my recollection of this: https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F10b2602771d21378cd6d76628a081c8a76dcf216-2600x2960.png&w=3840&q=75 https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... basically 9/13 are very close
- osti 7mo agoIt's only that one number that is for sonnet.
- 7mo ago