7 ms·
What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Rea
by didibus 3mo ago
What's interesting is this:
The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).
Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.
That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?
- theplumber 3mo agoOpus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)
- endorphine 3mo ago"him"? Have we reached that dystopia level?
- squigz 3mo agoI've been seeing a lot more anthropomorphization of these models on HN lately and it's alarming.
- perching_aix 3mo agoIt's a male name, and gendered pronouns can be hard for foreign speakers at times, irrespective of proficiency level. I wonder if you're overthinking this?
- xlii 3mo agoNot all HN visitors are native English speakers and in some languages "it" doesn't construct well with verbs, thus thought frameworks forms through usage of him/her. Nothing more to see I suppose.
- 55555 3mo agoEvery french person I've talked to IRL, for example, calls Claude "him." It's partly a language thing.
- redsocksfan45 3mo ago[dead]
- khimaros 3mo agoi prefer to call her Claudia
- didibus 3mo agoInterestingly, if you filter by the coding index, Sol xhigh is the best one, and only Opus5 max is better than Sol max and Sol high.
- perching_aix 3mo agoI wonder if I'm secretly being routed to some low grade version of Sol, or any of the GPT models really. Their performance is outright insulting at times, even at maximum reasoning, yet if I were to only read HN, I'd never know.
- jchw 3mo agoI've mostly actually stuck to low reasoning for most tasks since it seems to do a surprisingly good job even at low for the stuff I've been throwing at it, and I literally switched directly from Fable 5 to Sol more or less.
- zwaps 3mo agoSol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting
- xrisk 3mo agoMight be an indication that your task unit is too unstructured or your code base is a mess. Is your actual code doing something complex or is this incidental complexity? Relying on the model’s “intelligence” to patch over these issues hasn’t proven to be a reliable strategy for me. Of course, this might not apply to you, just my 2 paisa.
- Gareth321 3mo agoA well structured architecture, repo, and requirements document with concrete small deliverables can be competently delivered with cheap Chinese models. We rely on frontier models so that we don't have to spend several days/weeks on planning/architecture/documentation. That's their value proposition: superior intelligence. If they can't deliver that, they're useless at the current price.
- theplumber 3mo agoIt happened to me as well but in a different direction: i.e adds non library code in a shared library. Another issue I with GPT is that it is chasing too much edge cases/security issues(I.e chasing ghosts). However this makes it also a strong model because it fixes/solves problems that both Opus and Fable are incapable. In reviews it catches bugs that both Fable and Opus are missing to spot. To me the “best of both worlds” is to research the problem with GPT sol, create a plan with Fable and dual review it with both Fable and GPT-SOL and implement it with GPT-SOL. You can see in the code reviews how many times both Fable and opus are sloppy and superficial while GPT-SOL just does its due diligence …I had several problems that Fable just gave up and it was GPT that helped it sort it out.
- colinhb 3mo agoOpus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on
- ffsm8 3mo agoI suspect most of those comments on llms like the parents are generated by anthropic and openai to shape the discussion/mindset They always give off the same astroturfing vibes that reddit became infested with after the early 2010s (just look at it's comment history) Ofc unprovable for users. Ycombinatior could try to, but it'd just become a cat/mouse game which they'd likely lose because of the incentives
- andai 3mo ago> just look at it's comment history I checked one. Old account, nuanced takes, and shitting on everyone equally. The perfect HN user!
- ffsm8 3mo agoAh, I really walked into that one. Yeah, the phrasing + placement of the remark implied that all comments are artificial. That was not my actual intent, it was poorly expressed by me. I was specifically talking about the account which created the comment colinhb responded to. That'd make it the... Grand Grand Grand grandparent now I think?
- rixed 3mo agoThat's not necessary. Sometimes we make choices and we feel the need to justify them in vivid flamewars regardless of how arbitrary. Vi vs emacs or amd vs intel, now anthropic vs openai. We love to take sides, to belong to a small community of peers, no?
- mcintyre1994 3mo agoI don’t know that individuals can really be expected to use a new model enough to develop a proper opinion though. If I try a new model, and it doesn’t seem as good as the one I’m using, I’m just going to stop using it. That’s not enough data to give anyone else a useful view on it, but it’s enough for me to make my mind up. Especially because I’m probably trying it at work, and I can’t really justify using the company’s enterprise plan to develop my understanding of a model that I don’t think is going to be the one I use for my work.
- braebo 3mo agoMy experience is the opposite. GOT 5.6 Sol is the dumbest most dangerous frontier model I’ve ever used. It actively introduces bugs and hacks and lies about what it did. Its code is almost always slop that can’t make it through even a brief review without half a dozen wtf moments. Opus is always cleaning up the terrible mess and GPT models are banned now from work because of how harmful and mind numbing stupid they are. R.e. Opus doing the wrong thing — my environment always gives the required context or leaves a paper trail so I don’t have that issue with Claude models.
- andai 3mo ago> That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is? According to AA's "intelligence vs cost per task" and "intelligence vs time per task" graphs, Opus 5 High and Sol Max are roughly evenly matched on cost and time. On DeepSwe, Opus 5 beats Fable but not Sol. On FrontierCode, it destroys everyone, unless you set it higher than Medium effort, and then it tanks, falling to Sonnet level?