4 ms·
>I thought models are getting dumber, but benchmarks were convincing opposite >Opus 4.8 and Opus 5 seems worse models than Opus 4.6 After all we've heard abou
by unclebucknasty 1mo ago
>I thought models are getting dumber, but benchmarks were convincing opposite
>Opus 4.8 and Opus 5 seems worse models than Opus 4.6
After all we've heard about benchmark cheating, I'm earnestly not sure which or whether benchmarks are reliable anymore. But, beyond the models, I wonder if changes to their harnesses and/or instructions dumb them down. I have noticed models change their behavior, even when using the same version/effort. Sometimes for better. Sometimes for worse.
And, I have noticed a model go from really good to struggling. On 4.8 things were going well for a good stretch, so I did not switch to 5 when it came out. Even after hearing complaints about 5, 4.8 was still going well. Then, suddenly over the last few days, 4.8 seems to have nosedived. It feels similar now to the complaints I hear about 5.
In my case, it suddenly started ignoring my design guide, and introducing new fonts etc. It would even use several different fonts and sizes, as well as different margins for similar elements within the same page. It abandoned classes and started inlining styles. It started feeling random and, even after it realized it needed to go back to the design guide, it just continued with more of the same.
There seems to be something that happens after new model releases in both quality and behavior of previous models. It may not be immediately, but eventually there is frequently some regression.
- Bluestein 1mo agoMaybe compute relocation?
- unclebucknasty 1mo agoWhat do you mean? As in, they are reallocating compute away from previous models? If so, I would think that would result in worsened performance, not quality (unless you are also suggesting they may be quantizing).
- Bluestein 1mo agoSpot on. (I had not considered quantization - that's a thought ...)
- richardfey 1mo agoIt feels lke they replace older models with "optimised" versions which are cheaper to run, but keep the same name.
- unclebucknasty 1mo agoYou mean like quantized?
- richardfey 1mo agoYes; or something which has a similar effect.
- ethbr1 1mo agoIt would describe the observed behavior. Especially if there were an internal quant/efficiency team that wasn't as diligent about regressions as the primary model team.
- unclebucknasty 1mo agoThat's exactly what it feels like and it makes perfect sense, given the known compute constraints.
- pizzafeelsright 1mo agoI cannot speak to the benchmarks but we have many users, who have experience with each model, using it 95% of the week and we observe the same changes in the models as a whole. In addition to that, while yes, 4.6 and 5.0 can solve problems, they do so differently. Sometimes 5.0 does better by a wide margin but that would be expected as they are supposed to be better.