5 ms·
it is barely an improvement according to their own benchmarks. not saying thats a bad thing, but not enough for anybody to notice any difference
by haaz 1y ago
it is barely an improvement according to their own benchmarks. not saying thats a bad thing, but not enough for anybody to notice any difference
- waynenilsen 1y agoi think its probably mostly vibes but that still counts, this is not in the charts > Windsurf reports Opus 4.1 delivers a one standard deviation improvement over Opus 4 on their junior developer benchmark, showing roughly the same performance leap as the jump from Sonnet 3.7 to Sonnet 4.
- esafak 1y agoThat is a big improvement.
- ttoinou 1y agoThat's why they named it 4.1 and not 4.5
- zamadatix 1y agoWhen it's "that's why they incremented the version by a tenth instead of a half" you know things have really started to slow for the large models.
- phonon 1y agoOpus 4 came out 10 weeks ago. So this is basically one new training run improvement.
- zamadatix 1y agoAnd in 52 weeks we've gone 3.5->4.1 with this training improvement, meanwhile the 52 weeks prior to that were Claude -> Claude 3. The absolute jumps per version delta also used to be larger. I.e. it seems we don't get much more than new training run levels of improvement anymore. Which is better than nothing, but a shame compared to the early scaling.
- globalise83 1y agoIs it really a bigger jump to go from plausible to frequently useful, than from frequently useful to indispensable?
- deleted 1y ago[deleted]
- zamadatix 1y agoWhy is there supposed to be no step between frequently useful and indispensable? Quickly going from nothing to frequently useful (which involved many rapid hops between) was certainly surprising, and that's precisely the lost momentum.
- mclau157 1y agoThey released this because competitors are releasing things
- leetharris 1y agoGood! I'm glad they are just giving us small updates. Opus 4 just came out, if you have small improvements, why not just release them? There's no downside for us.
- AstroBen 1y agoI don't think this could even be called an improvement? It's small enough that it could just be random chance
- j_bum 1y agoI’ve always wondered about this actually. My assumption is that they always “pick the best” result from these tests. Instead, ideally they’d run the benchmark tests many times, and share all of the results so we could make statistical determinations.
- gloosx 1y agoThey need to leave some room to release 10 more models. They could crank benchmarks to 100% but then no new model is needed lol? Pretty sure these pretty benchmark graphs are all completely staged marketing numbers since they do solve the same problems they are being trained on – no novel or unknown problematic is presented to them.
- levocardia 1y ago"You pay $20/mo for X, and now I'm giving you 1.05*X for the same price." Outrageous!
- onlyrealcuzzo 1y agoI will only add that it's interesting that in the results graphic, they simply highlighted Opus 4.1 - choosing not to display which models have the best scores - as Opus 4.1 only scored the best on about half of the benchmarks - and was worse than Opus 4.0 on at least one measure.
- Topfi 1y agoI am still very early, but output quality wise, yes, there does not seem to be any noticeable improvement in my limited personal testing suite. What I have noticed though is subjectively better adherence to instructions and documentation provided outside the main prompt, though I have no way to quantify or reliably test that yet. So beyond reliably finding Needles-in-the-Haystack (which Frontier models have done well on lately), Opus 4.1 seems to do better in following those needles even if not explicitly guided to compared to Opus 4.