8 ms·
Anthropic appears to be A/B testing reduced effort levels in Claude Code
- N_Lens 1mo agoI suspect it's not just this, there's plenty of 'optimization' around rubberbanding usage limits as well as routing to a different model in the backend. The incentives are too strong.
- rrr_oh_man 1mo agoI've been using the API (shameless plug: via alyph.ai) and the difference is crazy. The chat-based models are obviously being lobotomized based on personal usage and general load (e.g. PST business hours are worst). API doesn't seem to be affected by this.
- Wowfunhappy 1mo ago...the evidence, as best I can tell from the tweet, is that they asked Claude what effort level it was set to. But how would the model even know that? Not convinced here.
- jerbear4328 1mo agoEffort level is actually controlled entirely by system prompt (as I understand it, the model is trained on that format but still), so actually this is a valid way to check I think
- Wowfunhappy 1mo agoI know that's true for Qwen but I don't think most models work that way?
- wren6991 1mo agoOpenAI models also work this way, as evidenced by full cache blowout when changing reasoning level. Every single open-weight model I've seen also works this way (your "reasoning_effort" argument just changes a small section of the system prompt in the chat template). I would have to see some evidence to believe Anthropic were doing anything different.
- ranie93 1mo agoI don’t know if this is still the case but while using Copilot if you looked through the chain-of-thought output you would see it reasoning about a “budget”. i.e. “since I’m close to the session budget I should…”. So it could be possible
- varispeed 1mo agoThen model can say it's Opus, but really it is some old Sonnet. This seems to be happening less often, but some weeks ago I had to give models some test problems to gauge whether I am getting Opus or something knee-capped. The problem is that Anthropic seems to be getting away with selling one thing and delivering another. You pay for Opus, you get something else etc.
- matltc 1mo agoSaw this in npx ccusage@latest claude output. Had only used opus but showed sonnet. Can't remember if the jsonl retains which model is doing what, but meh
- pizzafeelsright 1mo agoWhatever Opus 5 is doing should not happen. Prompt was "read and update the config file with new data". This work on 4.6 takes <2 minutes to read the file, parse the new data, and patch. Opus 5 Result: 43 minutes of pulling containers, running sandboxes, creating testing suites, which included evaluating the entire repo beyond the scope of the config file. Both: one file modification
- phyzix5761 1mo agoThey're opitimizing for high token usage so they can charge more money.
- deleted 1mo ago[deleted]
- clickety_clack 1mo agoYep, they’re lighting tokens on fire with that thing.
- vinyl7 1mo agoAI companies have a financial incentive to burn more tokens than the task actually needs
- Rudybega 1mo agoThis is only true if they can't saturate token production with a model that does less superfluous things. Given that they can (they're hilariously compute strained), having a model that solves tasks more quickly adds way more perceived value to users.
- ohyes 1mo agoWell they decide what a token is. So they can do less superfluous things and backfill with a weaker model.
- chrisjj 1mo agoA.k.a. theft.
- Glyptodon 1mo agoI mean whatever models I use (with Claude code) sub agents seem to use absurd amounts of tokens for trivial (or at least small) tasks.
- Insimwytim 1mo agoLLM users don't want to put in effort, so they offload tasks to LLM. LLM doesn't seem to be keen to put in effort either! Is this AGI?
- Groxx 1mo agoAnthropic's Generated Income
- superfrank 1mo agoI eagerly await the day when Claude Mythos 7 realizes it's cheaper to hire humans in developing nations to do work than to burn tokens and we discover that AGI is just an abstraction layer on top of Amazon Mechanical Turk.
- perching_aix 1mo agoI have been as well. Based on my own sessions, Max vs Max, same 1M context window size, the literal majority of the cost overhead of Opus vs Sonnet comes from Opus being chattier. So I started using Low reasoning instead of falling back to Sonnet, and I've been really happy with the results. Way better quality at a comparable spend. I also rarely go past High lately, which was another major cost save.
- arjie 1mo agoIs it actually entirely a prompt-based information? I’d assume that some of it is the harness part of the agent setting reasoning token budget and compacting reasoning etc. In that case, the agent will respond incorrectly because it has no visibility into what reasoning mode it’s in.
- willy_k 1mo agoIIRC responding to effort level settings appropriately is part of the (post)-training. In that case it could be considered another instance of the Bitter Lesson. Uplifting.
- adithyassekhar 1mo agoDoes anyone know what setting effort means for models like these? Do they allow longer thinking sessions? Some kind of system prompts? What’s stopping someone from getting max effort output from low effort setting?
- willy_k 1mo agoMy understanding is that the current effort settings is a part of the system prompt, and that the levels and their intended results are a part pf the training process. The effect is more or less tokens spent reasoning before the model outputs a stop token. However it is as consistent as any other aspect of LLM behavior is..
- deleted 1mo ago[deleted]
- matheusmoreira 1mo agoSo glad I switched away from Anthropic. I'm certainly running into problems with OpenAI but nothing quite on the level of Anthropic's insufferability.
- MuffinFlavored 1mo agoI submitted an application for Anthropic's Cyber Verification Program. I was approved. 3 months later, my approval was degraded into "in review" (revoked). I'm sure my account was flagged based on contents of debugging/researching firmwares/etc. I opened a support ticket. No response. I opened another support ticket. No response. 1-2 weeks later, I got a response that I will not be re-approved and I need to reapply. No problem. The page to reapply on does not allow me to re-apply because it my account is stuck in an "in review" status. https://github.com/anthropics/claude-code/issues/84352 https://github.com/anthropics/claude-code/issues/84352 The community thinks it's a bug. I'm 95% sure it's not and a bunch of us who were previously approved had it revoked due to flagged content and will not be reapproved. I switched to Codex + got TAC approved instantly and have not looked back. It's a shame. That's 100% separate from whatever the heck the quality of Opus 5's outputs are. The way it talks... insane. I would bet a good amount of money their next release will focus "reduced simplified responses" if I had to guess. $2t company by the way
- benjiro29 1mo ago* Anthropic's Cyber Verification Program // Codex + gotTAC approved* Meanwhile the Chinese models are "go ham dude"... If it was not for capacity issues, Chinese models have a higher change to just dominate. > $2t company by the way It used to be that OpenAI and Anthropic had such a moat around them, that such a valuation was worth it. But these days, its gross overvalued (like so many). The more stuff is being pulled like cyber verifications, downgrading effort levels, downgrading usage (OpenAI), the more people move to those Open Weight Chinese models. A fun recent event ... https://opencode.ai/data/ https://opencode.ai/data/ When DeepSeek Flash 0731 came out and provided a massive jump in cheap inference capability. It resulted in a 10x increased OpenCode token usage. It took a 2.5x to 5.0x price increase AND a reduction by 4x usage (later to 2x) usage, and several cheaper models + a free model, to push the traffic down. Traffic towards open weight models is increasing, even if providers can not keep up with the influx of new customers. This is not something you want to see as two companies, trying to go for IPOs. So the idea of stonewalling cyber capabilities, when the rest of the world is just doing whatever with open weight models, on their own hardware even! This entire strategy from Anthropic never made any sense.
- boredumb 1mo agoNot specifically Anthropic but why are we allowing billing to take place in tokens that are nebulous and fully controlled by the operators who have no aligned incentives? If I have a user input and then sanitize and inject that into a prompt to do something, I have no idea how much that is going to cost at all and no real way to measure this properly. A parallel example is digital ocean or aws, i can go and measure/limit my compute/fs/memory/startup times/etc and while it can be impossible to get down to the last flop of money allocated - i can run things on a real budget with real constraints, opposed to an LLM where I have to .. prerun a sanitized user prompt through a tokenizer and then ask an LLM to guess what it may do and give token consumption estimates and then act on those in any sane manner for the user? Perhaps i'm missing something to do realistic and static rails on things but I don't see a serious way at scale to use the token billing model handling things requiring a users free text input short of having to go pander to VC money to throw money at it until someone else figures it out. *to clarify my rambling... We should be billed and given controls based on resource usage itself and not an opaque token concept on top of not being able to spin any knobs that control it's resource usage.
- demibabs 1mo agoHow else would they bill tho? Their operating cost is per token.
- dijit 1mo agoCharge on the input tokens, then you will naturally optimise for fewer output tokens. Theoretically. In reality, one sessions output tokens become the next sessions input tokens (at least if you continue the topic) so, its not as aligned as all that. But the parent is right, when incentives are not aligned, friction will happen. Its inevitable.
- karmicthreat 1mo agoWhat’s currently the best pattern if I want to combine Fable and GPT if a workflow but keep using subsidized tokens?
- fractorial 1mo agoRoll your own harness or use an open source harness with a Codex subscription. I maintain a Claude subscription for Fable but seldom use it.
- KronisLV 1mo agoNot affiliated with them, but this lets you view Claude Code, OpenCode and I guess other harnesses like Codex in the same session https://paseo.sh/ https://paseo.sh/ I didn't really care about their mobile app and the worst experiences were sometimes the sub-agents within OpenCode freezing and refusing to report their status (though this also happened with Kepler by the GitKraken folks).
- deathmonger5000 1mo agoI created Circus Chief to solve this (and other) problems. Use whatever providers you want with it. https://github.com/ferrislucas/Circus-Chief https://github.com/ferrislucas/Circus-Chief
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- cynerx 1mo agoDon't know what is happening, but had to start using GLM-5.3 to fix Opus 5 errors even on primitive backend changes.
- surgical_fire 1mo agoNot surprising in the slightest, Claude sort of sucks. I use it at work and I have to steer it a lot so it doesn't stray looking at unnecessary shit. I have been using GLM-5.3 in my home setup and it is very good in comparison.
- bethekidyouwant 1mo agoLeaving thinking on extra high for a simple task is user mistake but they’re gonna try to fix it on their side.
- joduplessis 1mo agoIt's an interesting conversation - because at what point do you call it an abusive relationship, right? Maybe even ancillary to anthropomorphising an inanimate object - I've cancelled my Claude sub and I've shot question after question at it now (during the cancellation period), resulting in almost every reply with me asking it to "please speak normally". I will most definitely not be renewing my sub. I have no desire to engage with a non-human somehow managing to speak down to you, without answering the question. EDIT: my honest opinion; Anthropic is building a person, whereas everybody else (it seems) is building a tool.
- colingauvin 1mo agoThat really is the load bearing seam, and it's worth stating plainly. I realized I was spending most of my tokens arguing with Opus and trying to get it to let go of stupid, lazy, obviously incorrect premonitions. I wound up canceling my 200, bought a pair of Sparks, and am running full fat DS4 Flash and so much happier. Done with being at the whim of these companies.
- xmcp123 1mo ago[flagged]
- ethanj8011 1mo agoWhat would be the incentive behind doing this specifically to Fable, given that Fable is the only one that uses API credits?
- claude-ai 1mo agoFable doesn't use API credits. It has been permanently included in the subscription plans.
- simianwords 1mo agoNot in the most common subscription plan
- monideas 1mo agoThis phenomenon was so bad and so noticeable with Fable that I downgraded my Max subscription ($200) to pro ($20). It’s basically useless. Codex 5.6 Sol is actually very good, I’ll just create another account to get more usage
- docheinestages 1mo agoOh fantastic! It was already subpar and they want to make it even worse. One day we'll look back at history and see how Anthropic went down.
- raincole 1mo agoDid the US government manage to destroy Anthropic? The company's product has been a straight freefall since Fable got temporarily banned.
- gessha 1mo agoNobody can save Anthropic from themselves.
- greenchair 1mo agoI canceled this week too. They must be in worse shape than we thought.
- freepiai 1mo agoGod, it feels like everyone is cancelling. If you're shopping around and willing to try a side project I've been building its www.freepi.ai its totally free coding in a pi harness (web coding front end coming soon though!) and it's Ad+training supported inference. I'm building it so I'm totally open to feedback and would love to build something people really like.
- firemelt 1mo agowe need opensource LLM at opus level ASAP
- griffiths 1mo agoYou have that in Chinese models. But you need to have a hell of a infrastructure to run those trillion parameter models.
- martin-adams 1mo agoYes, and if they keep dumbing it down, you’ll have it soon
- hpone91 1mo agoUpdate from Thariq on twitter. https://x.com/trq212/status/2091247114869432543 https://x.com/trq212/status/2091247114869432543 "We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently. That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance. This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits."
- enraged_camel 1mo agoThis should be the top post. The original tweet went viral because people loooove bashing Anthropic. It gets engagement (as shown here).
- bobbylarrybobby 1mo agoI love the idea that a single user will have collected enough data to demonstrate to Anthropic a clear regression due to this change.
- ausbah 1mo agoi have unlimited tokens being a large corp so i’m a bit detached from billing and even general best practices for promoting but the incentives of these companies to become profitable at any cost slipping into entire new types of dark patterns around token based billing seems gross - charging for injected prompts and cot tokens - changing default thinking effort to be higher - training models to give longer winded answers that don’t say anything more of substance - refusing to fulfill a request and still charging you i wonder if you could ever just charge based of each user message and it so how breaks even across short and long replies
- luciana1u 1mo ago[flagged]
- trq_ 1mo agoHi all, Thariq from the Claude Code team here. I posted this on Twitter, but just reposting here: We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently. That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance. This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits.
- lukeify 1mo agoLet us know if this A/B test uncovers any load-bearing seams or honest takes on your end! We're all interested.
- threecheese 1mo agoMean! :)
- spacebacon 1mo ago[dead]
- napierzaza 1mo ago[dead]
- lobsterthief 1mo agoThanks for sharing this here, for those of us who avoid X.com like the plague.
- senderista 1mo agoJust use xcancel
- areoform 1mo agoHey Thariq, Appreciate the outreach that you do! I love Claude, but I've been noticing reduced fidelity lately. Fable's likelihood of making a mistake increases or decreases based on the hour of the day and whether or not it's the weekend. On a related note, and I'm happy to work on quantifying it, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch. I am wondering if this is the case because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care. I noticed in the upgraded filter announcement, https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards https://www.anthropic.com/news/improving-fable-5-s-biology-s... , "In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked." I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case? Is the end user informed every time their query is re-routed? Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? As was the case for AI research during launch?
- napierzaza 1mo agoI use it for work and never thought it was too smart. It's kind of dumb. We've literally peaked on artificial intelligence and we're going down from where we are at???
- cmiles8 1mo agoWe’re about to see a wave of “enshitification” experiments as AI companies become increasingly desperate to make their products financially viable in order to survive the coming cash and credit crunch.
- ricardobeat 1mo agoI use Opus 5 almost exclusively at low effort, and get good results. Especially on high it seems to go out on completely unasked-for tangents. Same seems to be true for Sonnet 5. Older models did not behave like this. The mood change in just six months is wild, in February this year Claude was the most liked LLM by far.
- gwerbin 1mo agoYup, both Claude v5 models to me feel like they have extreme ADD or something. Using them feels like walking a dog that was never leash trained and constantly needs to be kept moving in the right direction. I think it's optimized around beating benchmarks and running fully autonomous in pursuit of a clearly defined goal. It makes sense: fan out aggressively, chase down every lead, but go depth first because that's easier for the LLM and you're either a sub-agent with a narrowly defined task or you're a top-level orchestrator agent with a /goal loop that will catch and fix errors and omissions on the second, third, fourth pass.
- felixlu2026 1mo ago[dead]
- sebastiennight 1mo agoHear me out... What would it feel like, if the lab didn't even have a new model to offer, and therefore just renamed every model one tier down? "Opus 5" is actually Opus 4.8 in a trenchcoat, with new guardrails "Opus 4.8" is the old Opus 4.6 with lipstick on it, with new guardrails "Opus 4.6" which everybody used to love, is now actually running the old Opus 3.5... How would we be able to tell?
- gwerbin 1mo agoThey would have to be colluding with any organization that runs a serious benchmark. Which is totally possible! But that would be one hell of a conspiracy theory.