5 ms·
Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of t
by throwaw12 2mo ago
Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)
- lwansbrough 2mo agoGoing to call it user error if you find Opus 4.5 better than 5, sorry.
- yorwba 2mo agoWell, what kinds of things do you see Opus 4.5 completely fail at? Maybe those are not the ones that newer models have improved on.
- sscaryterry 2mo agoEnshittification.
- stared 2mo agoIt's called frog boiling. We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age. If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice the difference.
- vmg12 2mo agoThe actual term for this is hedonic adaptation.
- ffsm8 2mo agoesp. important to point that correct term because frog boiling is a urban myth. frogs dont stay in a pot even if you slowly increase the heat. they leave. it has reportedly been attempted multiple times and they. always. leave.
- rf15 2mo agoI've worked with these systems for four years now and they have not meaningfully improved in that time frame. We still have: - statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system) - Math completely fails in longer contexts - "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion - smearing of properties between logically distinct objects (a red ball and a green cube can quickly become a red cube and a green ball)
- Barbing 2mo agoLike how toddlers’ skills don’t meaningfully improve on infants’, because either could wake up in a wet bed.
- mdp2021 2mo agoLet us be more clear: there is no structural jump, no architectural overcoming of the original fault. (Edit: and on a similar point, structural properties such as having static ntetworks, as opposed to continuously learning and improving architectures (such as us), will reveal that there is still road ahead.)
- Barbing 2mo agoMaybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.” You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvious show stopping bug yourself. The technology is not a brand new one that fixed everything wrong with the old one, no, but not sure I would’ve noticed your comment if it had been such a bland observation. I genuinely assume good faith here… will say am tempted to assume the standards of someone posting such a thing might be impossibly high. Glad to be having a fun conversation instead of getting your grades on my work product or something :)
- tudelo 2mo agoI honestly just use GPT models nowadays, Claude models are too restrictive and more of a quitter and fable/whatever is just too expensive to be worth it.
- OtherShrezzing 2mo agoIt could be that the set of your day-to-day workload which could feasibly be accelerated by AI just happens to be saturated around Opus4.5, but you can still see lots of “reasoning” which makes you think the model is more performant in the first days of use. That’d mean you couldn’t perceive any meaningful difference in more powerful models’ results, even though you can see a difference in the raw output due to the length of reasoning traces leading up to the result. So for example, if your workload was literally just addition of sets of numbers, you’d never have noticed progress in the result beyond GPT3.x level models. But you would perceive a difference in the now-Tolstoyan length reasoning text accompanying the result.
- SubiculumCode 2mo ago5 seems incredibly smart to me in my conversations today about some pretty niche ideas in.longitudinal modeling. It.felt.like a big step up.from 4.8, to me
- kranke155 2mo ago5 felt both smarter than me and dumber in some ways - it gets stuck to its original ideas. I had never seen a model harder to talk into changing its initial opinions. it continuously hedges.
- submeta 2mo agoSome 20 years ago, the telecommunications sector in Germany was liberalized. Many telephone card providers entered what had previously been a barely competitive market. They advertised their products with aggressive claims like: “Buy our €10 top-up card and get 660 minutes to destination X.” For the first few weeks, they would actually provide those 660 minutes to establish trust in their cards. But after a while, they would quietly start reducing the number of minutes on subsequent top-ups—say, from 660 minutes down to only 300. They wouldn’t do this for every card, so it was difficult to prove. Instead, they relied on averages across their customer base to make the economics work. Lately, I’ve found myself wondering whether something similar may be happening with frontier AI models. Companies launch with an exceptionally strong model and generous compute limits to build adoption. Once the model is established as a market leader, the incentives change, and users may start perceiving the service as becoming more constrained or less capable over time. I don’t have evidence that this is what’s happening with Anthropic—or with any other AI company. It’s simply a pattern that the current situation reminds me of.
- slopinthebag 2mo agoSame, like I prefer 5.3 codex over the “stronger” models.