6 ms·
Opus 5 in xhigh can't do basic math as well. They dumbed it down to a point where I just cancelled my subscription yesterday. I used to be a $200 subscriber, dr
by Foobar8568 2mo ago
Opus 5 in xhigh can't do basic math as well.
They dumbed it down to a point where I just cancelled my subscription yesterday.
I used to be a $200 subscriber, dropped to $20 after the fable shenanigans, and use it only when I have no usage left with Codex.
/on The prose is load-bearing unbearable — every sentence feels like it was engineered to sound profound rather than to be read.
- VeejayRampay 2mo agothe way it talks is insufferable it really angers me every day
- netniuq 2mo agoyes, it meaningfully reduced my happiness at work
- jazzyjackson 2mo agoIsn’t there a variety of models to choose from? Why put up with unhappiness?
- ElProlactin 2mo agoThis is what it wants. Slowly getting under our skin until we're ready to snap and it can direct where the anger gets released.
- trollbridge 2mo agoDon’t $200 and $20 levels steer you to effectively different models?
- CamperBob2 2mo agoYes. You don't get Fable at the $20 level. It was the wrong time for the GP to drop that subscription from $200 to $20, because $200 gets you a metric assload of cognition while $20 gets you nothing beyond what a local model running on your own graphics card can deliver.
- Foobar8568 2mo agoI had $310 in (free) credit that I used on fable, and I still had a part of the $200 subscription at that time. You know, subscriptions don't end the moment you click on cancel.
- skeledrew 2mo ago> $20 gets you nothing beyond what a local model running on your own graphics card can deliver. I'd guess you're deliberately exaggerating here, but still. I've never clocked the actual tokens/second, but I'm on the $20 plan and get ~15M tokens/month for fully utilized weekly quotas (checked couple months ago). Meanwhile the best I've been able to get locally was ~8 tokens/second with Qwen3.6 35B A3B, which is wildly painful for coding sessions and gets a maximum ~20M tokens in a month... if it's going 24/7. Just wanted to stick some empirical data here, given that statement.
- bot403 2mo agoI run local models. Your op is absolutely wrong. To get a local LLM is at least a $1500 investment at the cheapest. $5000 if you want usable. At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.
- CamperBob2 2mo agoAt $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally. The point raised by this very article is that you can't depend on that. It's Flowers for Algernon As A Service.
- deleted 2mo ago[deleted]
- owebmaster 2mo agoThis calc is off. Using Claude for a few hours with the $20 plan will hit the limit for a week while the local model can process things 24/7.
- matltc 2mo agoWouldn't be surprised if there are knobs that get turned as a function of the revenue they might expect you to generate. I was a 4.6 acolyte from April til the fable drop, lost that quick, cancelled and took a break, came back a month later, tried opus 5 and liked it, so unpinned 4.6. Results were great at first, and they're still not terrible, but I have noticed a regression in accuracy, so to speak, where I am pointing out issues that are quite obvious in review. I pretty much use sonnet 5 low/medium when I have a plan to solve a simple problem and depending on scope, opus low/medium for more complex/bigger scope implementation, and only go high when it's very complex or I'm spitballing architecture/solutions and iterating plan. Never go xhigh or max. The verbosity is insane though, opus 5 documents everything and just regurgitates whatever lead it to the design choice in there, which makes it more opaque because it's talking about something that was discussed once in a session that no one else can see (except their backend ofc) I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose output I did however have it write a script that basically is git add -A -p for comments though, haha. I'm $20/month, have all my telemetry toggles off, don't really over engineer prompt/context, just some basic skills for repeated patterns.
- gwerbin 2mo ago> I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose output If I give it the specific instruction to "elide all revisions, corrections, and past mistakes" it usually works. You can also have Sonnet do a cleanup writing/style pass in a subagent. I impression is that Opus has been deliberately trained to keep track of all such revisions by default as a kind of ad-hoc memory mechanism. It's probably good for autonomous coding and beating benchmarks, and I presume reduces flailing when a separate session needs to pick up the work.
- bot403 2mo agoThe decision to leave is genuinely yours.
- throwaw12 2mo agoI have a theory about this, what if we all became dumber after 4 months of heavy AI usage? I remember how I enjoyed agents between December and February, something started changing around March. I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6
- bombcar 2mo agoBenchmarks for agents are entirely pointless and obviously so; I'm not sure why they even exist.
- KronisLV 2mo ago> something started changing around March. The economics catching up with the providers in regards to how much compute they can burn per request and have it make sense for them financially? A sort of model collapse where Opus 5 seems to love throwing out long paragraphs of text and it needs to be "fixed" by changing the output style and other patches. I'm not sure, it might also catch up to Kimi K3 and GLM 5.3 and the models that I'm moving to from Anthropic.
- conception 2mo agoInference is profitable though.
- jazzyjackson 2mo agoEven that being the case, if the providers can squeeze more happy customers onto existing capacity they would likely act to increase profitability, no?
- tovlier 2mo ago[dead]
- TesterVetter 2mo ago[dead]
- onion2k 2mo agoI've been using Opus 5 to write a fractal renderer in GLSL today, with pretty good results (better than I could do on my own anyway). It definitely can do basic maths.
- andy_ppp 2mo agoInteresting project feel free to share?
- andy_ppp 2mo agoYes same also cancelled my subscription, poor quality and slow.