4 ms·
That last benchmark seemed like an impressive leg up against Opus until I saw the sneaky footnote that it was actually a Sonnet result. Why even include it then
by bicx 7mo ago
That last benchmark seemed like an impressive leg up against Opus until I saw the sneaky footnote that it was actually a Sonnet result. Why even include it then, other than hoping people don't notice?
- conradkay 7mo agoSonnet was pretty close to (or better than) Opus in a lot of benchmarks, I don't think it's a big deal
- jitl 7mo agowat
- 0123456789ABCDE 7mo agomaybe gp's use of the word "lots" is unwarranted https://artificialanalysis.ai https://artificialanalysis.ai indicates that sonnect 4.6 beats opus 4.6 on GDPval-AA, Terminal-Bench Hard, AA Long context Reasoning, IFBench. see: https://artificialanalysis.ai/?models=claude-sonnet-4-6%2Cclaude-sonnet-4-6-adaptive%2Cclaude-sonnet-4-6-non-reasoning-low-effort%2Cclaude-opus-4-6-adaptive%2Cclaude-opus-4-6 https://artificialanalysis.ai/?models=claude-sonnet-4-6%2Ccl...
- conradkay 7mo agoI was basing it off my recollection of this: https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F10b2602771d21378cd6d76628a081c8a76dcf216-2600x2960.png&w=3840&q=75 https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... basically 9/13 are very close
- osti 7mo agoIt's only that one number that is for sonnet.
- 0123456789ABCDE 7mo agoexcept for the webarena-verified