15 ms·
Claude Code daily benchmarks for degradation tracking
- qwesr123 8mo agoFYI the MarginLab Claude Code degradation tracker is showing a statistically significant ~4% drop in SWE-Bench-Pro accuracy over the past month
- beardsciences 8mo agoVery interesting. I would be curious to understand how granular these updates are being applied to CC + what might be causing things like this. I feel like I can notice a very small degradation but have compensated with more detailed prompts (which I think, perhaps naively, is offsetting this issue).
- chrisjj 8mo ago> more detailed prompts (which I think, perhaps naively, is offsetting this issue). Is exacerbating this issue ... if the load theory is correct.
- goldenarm 8mo agoI really like the idea, but a "±14.0% significance threshold" is meaningless here. The larger monthly scale should be the default, or you should get more samples.
- zacmps 8mo agoCould you elaborate what you think the problems are? I guess they should be using some form of multiple comparison correction?
- goldenarm 8mo agoThe daily scale is not statistically significant and is meaningless. You should lower the confidence interval by either increasing the scale or the evaluations.
- turnsout 8mo agoThis is probably entirely down to subtle changes to CC prompts/tools. I've been using CC more or less 8 hrs/day for the past 2 weeks, and if anything it feels like CC is getting better and better at actual tasks. Edit: Before you downvote, can you explain how the model could degrade WITHOUT changes to the prompts? Is your hypothesis that Opus 4.5, a huge static model, is somehow changing? Master system prompt changing? Safety filters changing?
- fragebogen 8mo agoI was going to ask, are all other variables accounted for? Are we really comparing apples to apples here? Still worth doing obviously, as it serves a good e2e evaluations, just for curiosity's sake.
- FfejL 8mo agoHonest, good-faith question. Is CC getting better, or are you getting better at using it? And how do you know the difference? I'm an occasional user, and I can definitely see improvements in my prompts over the past couple of months.
- turnsout 8mo agoGood-faith answer: I can't be certain. But I've been using CC since its release, and Cursor before that (and actually going all the way back to GPT3 to do codegen in the Playground). After getting used to the CC workflow, the way that I use it has been pretty consistent. To be specific, I use basically the same AGENTS.md with small modifications for each project, and I live almost exclusively in Plan mode and the best model (currently Opus 4.5). My initial prompting is boilerplate at this point, and looks like this: (Explain overall objective / problem without jumping to a solution) (Provide all the detail / file references / past work I can think of) (Ask it "what questions do you have for me before we build a plan?") And then go back and forth until we have a plan. Compared to my work with CC six months ago, it's just much more capable, able to solve more nuanced bugs, and less likely to generate spaghetti code.
- rob 8mo agoI agree with you, it's personally hard to tell. For me I've noticed it getting nothing but better over the past couple months, but I've been working on my workflows and tooling. For example, I used to use plan mode and would put everything in a single file and then ask it to implement it in a new session. Switching to the 'superpowers' plugin with its own skills to brainstorm and write plans and execute plans with batches and tasks seems to have made a big improvement and help catch things I wouldn't have before. There's a "get shit done" plugin that's similar that I want to explore as well. The code output always looks good to me for the most part though and I've never thought that it's getting dumber anything, so I feel like a lot of the improvements I see are because of a skill issue on my part trying to use everything. Obviously it doesn't help there's a new way to do things every two weeks though.
- fragebogen 8mo agoWould love to see this idea expanded to ever alleged SoTA model currently in production. Any speculation as to why this degradation occurs?
- embedding-shape 8mo agoAnecdote, I don't have any proof and it's just a feeling. But around afternoon in GMT+1 compared to the morning/midday, there seems to be a change in the quality of responses, which seems to line up with when the US wakes up. I consistently get (what feels like) worse responses in both Codex and Claude Code in the afternoon/night compared to morning/midday, so much that I usually give up then try the same prompt next morning and get better results. But I guess that might as well be about me being more tired in the night than morning too, as I said, haven't measured this.
- jzig 8mo agoIt’s the afternoon slump. The AI needs a cup of coffee and to doomscroll for half an hour!
- embedding-shape 8mo agoOr a load balancing technique :) Either way, it kicks me off to do other things so maybe it isn't so bad after all.
- sciencejerk 8mo agoWhy is this happening?
- giwook 8mo agohttps://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues https://www.anthropic.com/engineering/a-postmortem-of-three-...
- observationist 8mo ago>>> We never reduce model quality due to demand, time of day, or server load. The problems our users reported were due to infrastructure bugs alone. Just ignore the continual degradation of service day over day, long after the "infrastructure bugs" have reportedly been solved. Oh, and I've got a bridge in Brooklyn to sell ya, it's a great deal!
- alias_neo 8mo ago> We never reduce model quality due to demand, time of day, or server load Forgive me, but as a native English speaker, this sentence says exactly one thing to me; We _do_ reduce model quality, just not for these listed reasons. If they don't do it, they could put a full stop after the fifth word and save some ~~tokens~~ time.
- chrisjj 8mo agoMoreover the assurance re model quality is not re results quality.
- observationist 8mo agoYes, Dario is responsible for some of the weaseliest of corporate weasel wording I've ever seen, and he's got some incredible competition in that arena. Those things aren't the reason, they're just strongly coincidental with the actual reason, which is to slow the burn rate and extend the runway.
- 8mo ago
- Dowwie 8mo agoSimply search user prompts for curse words and then measure hostility sentiment. User hostility rises as agents fail to meet expectations.
- Trufa 8mo agoI'm glad I'm not the only one.
- mrbananagrabber 8mo agoI uh might be skewing that as I generally just use a lot of curse words with Claude by default
- ctxc 8mo agoI feel bad about it but sometimes it's so daft, I can't even xD It's not my fault, they set high standards!
- smotched 8mo agothere are many times where I just do it myself and it thinks it did well.
- preuceian 8mo agoMaybe im overlooking something obvious but how do you 'simply' scan the content of Claude users their prompts?
- 8mo ago
- silverlight 8mo agoThere was a moment about a week ago where Claude went down for about an hour. And right after it came back up it was clear a lot of people had given up and were not using it. It was probably 3x faster than usual. I got more done in the next hour with it than I do in half a day usually. It was definitely a bit of a glimpse into a potential future of “what if these things weren’t resource constrained and could just fly”.
- yoavsha1 8mo agoI had that exact same feeling during the US holidays where I got to enjoy 2x usage limits and everything just seemed to work well
- cmrdporcupine 8mo agoI had terrible results during the holidays -- it wasn't slow but it was clear they were dealing with the load by quantizing in spots because there were entire chunks of days when the results from it were so terrible I gave up and switched to using Gemini or Codex via opencode.
- abathologist 8mo agoI find that if I have my rabbit's foot and lucky socks on, I win working code ~1.2x more often.
- svdr 8mo agoI would also regret it if they become that fast; right now I can really take a moment to enjoy the hard work the model is doing for me.
- asimovDev 8mo agohttps://xkcd.com/303/ https://xkcd.com/303/ the evolution of this xkcd
- nlh 8mo ago
- dajonker 8mo agoWouldn't be surprised if they slowly start quantizing their models over time. Makes it easier to scale and reduce operational cost. Also makes a new release have more impact as it will be more notably "better" than what you've been using the past couple of days/weeks.
- YetAnotherNick 8mo agoBenchmarks like ARG AGI are super price correlated and cheap to run. I think it's very easy to prove that the models are degrading.
- rustyhancock 8mo agoOooff yes I think that is exactly the kind of shenanigans they might pull. Ultimately I can understand if a new model is coming in without as much optimization then it'll add pressure to the older models achieving the same result. Nice plausible deniability for a convenient double effect.
- kilroy123 8mo agoIt sure feels like they do this. They claim they don't, but using it every day for 5-10 hours a day. You notice when something changes. This last week it seems way dumber than before.
- eli 8mo agoI would be surprised tbh. Anthropic does not exactly act like they're constrained by infra costs in other areas, and noticeably degrading a product when you're in tight competition with 1 or 2 other players with similar products seems like a bad place to start. I think people just notice the flaws in these models more the longer they use them. Aka the "honeymoon-hangover effect," a real pattern that has been shown in a variety of real world situations.
- Roark66 8mo agoI haven't noticed much difference in Claude, but I swear gemini 3 pro preview was better in the first week or two and later started feeling like they quantized it down to hell.
- 8mo ago
- ofirpress 8mo ago[SWE-bench co-author here] It seems like they run this test on a subset of 50 tasks, and that they only run the test once per day. So a lot of the movement in accuracy could be attributed to that. I would run on 300 tasks and I'd run the test suite 5 or 10 times per day and average that score. Lots of variance in the score can come from random stuff like even Anthropic's servers being overloaded.
- dana321 8mo ago"Lots of variance in the score can come from random stuff like even Anthropic's servers being overloaded" Aha, so the models do degrade under load.
- mohsen1 8mo agoHope you don't mind the unrelated question: How do you pay for those SWE-bench runs? I am trying to run a benchmark but it is too expensive to run enough runs to get a fair comparison. https://mafia-arena.com https://mafia-arena.com
- ofirpress 8mo agoBenchmarks can get costly to run- you can reach out to frontier model creators to try and get them to give you free credits, but usually they'll only agree to that once your benchmark is pretty popular.
- ghm2199 8mo agoIn medicine there is a concept of reporting adverse effects of medication or interventions which are then collectively studied for Public Health [MedWatch][VAERS][EudraVigilance] and in academia. We should have something like that for all coding agents(and agents in other fields too), given how widely its deployed and affect on "health" in general(not only human). Call it the AI "health" of things benchmark. I would imagine a sort of hybrid qualities of volunteer efforts like wikipedia, new problems like advent of code and benchmarks like this. The goal? It would be to study the collective effort on the affects of usage to so many areas where AI is used. [MedWatch](https://www.fda.gov/safety/medwatch-fda-safety-information-and-adverse-event-reporting-program/reporting-serious-problems-fda https://www.fda.gov/safety/medwatch-fda-safety-information-a...) [VAERS](https://www.cdc.gov/vaccine-safety-systems/vaers/index.html https://www.cdc.gov/vaccine-safety-systems/vaers/index.html) [EudraVigilance](https://www.ema.europa.eu/en/human-regulatory-overview/research-development/pharmacovigilance-research-development/eudravigilance https://www.ema.europa.eu/en/human-regulatory-overview/resea...)
- antirez 8mo agoWhy I do not believe this shows Anthropic serves folks a worse model: 1. The percentage drop is too low and oscillating, it goes up and down. 2. The baseline of Sonnet 4.5 (the obvious choice for when they have GPU busy for the next training) should be established to see Opus at some point goes Sonnet level. This was not done but likely we would see a much sharp decline in certain days / periods. The graph would look like dominated by a "square wave" shape. 3. There are much better explanations for this oscillation: A) They have multiple checkpoints and are A/B testing, CC asks you feedbacks about the session. B) Claude Code itself gets updated, as the exact tools version the agent can use change. In part it is the natural variability due to the token sampling that makes runs not equivalent (sometimes it makes suboptimal decisions compared to T=0) other than not deterministic, but this is the price to pay to have some variability.
- eterm 8mo ago4. The graph starts January 8. Why January 8? Was that an outlier high point? IIRC, Opus 4.5 was released late november.
- littlestymaar 8mo agoOr maybe, juste maybe, that's when they started testing…
- eterm 8mo agoWayback machine has nothing for this site before today, and article is "last updated Jan 29". A benchmark like this ought to start fresh from when it is published. I don't entirely doubt the degradation, but the choice of where they went back to feels a bit cherry-picked to demonstrate the value of the benchmark.
- littlestymaar 8mo agoWhich makes sense, you gotta wait until you get enough data before you can communicate on the said data… If anything it's coherent with the fact that they very likely didn't have data earlier than January the 8th.
- IshKebab 8mo ago> We model tests as Bernoulli random variables and compute 95% confidence intervals around daily, weekly, and monthly pass rates. Statistically significant differences in any of those time horizons are reported. Doesn't really work like that. I'd remove the "statistically significant" labelling because it's misleading.
- sroerick 8mo agoMy personal conspiracy theory is that they choose who to serve a degraded model to based on social graph analysis and sentiment analysis, maximizing for persuasion while minimizing compute.
- arcanemachiner 8mo agoSounds more like a sound business plan than a conspiracy theory.
- copilot_king 8mo agoIt sounds like fraud to me
- arcanemachiner 8mo agoDoes it say anywhere in their terms of service that they guarantee the quality of the model, or promise not to modify it? https://www.anthropic.com/legal/consumer-terms https://www.anthropic.com/legal/consumer-terms https://www.anthropic.com/legal/commercial-terms https://www.anthropic.com/legal/commercial-terms
- deleted 8mo ago[deleted]
- copilot_king 8mo agoIMO this strategy seems inspired by TikTok's approach for retaining new uploaders. TikTok used to give new uploaders a visibility boost (i.e., an inflated number of likes and comments) on their first couple of uploads, to get them hooked on the the service. In Anthropic/Claude's case, the strategy is (allegedly) to give new users access to the premium models on sign-up, and then increasingly cut the product with output from cheaper models. Of course, your suggestion (better service for users who know how to speak Proper English) would be the cherry on top of this strategy. From what I've seen on HackerNews, Anthropic is all-in on social media manipulation and social engineering, so I suspect that your assumption holds water.
- stared 8mo agoDoes it benchmark the underlying code (Opus 4.5) or Claude Code harness? If the second, I would love to see CC versions involved. I would be curious to see on how it fares against a constant harness. There were thread claiming that Claude Code got worse with 2.0.76, with some people going back to 2.0.62. https://github.com/anthropics/claude-code/issues/16157 https://github.com/anthropics/claude-code/issues/16157 So it would be wonderful to measure these.
- Jcampuzano2 8mo agoClaude Code. They mention they are using claude codes CLI in the benchmark, and claude code changes constantly. I wouldn't be surprised if the thing this is actually testing is benchmarking just claude codes constant system prompt changes. I wouldn't really trust this to be able to benchmark opus itself.
- jampa 8mo agoI am using API mode, and it's clear that there are times when the Claude model just gives up. And it is very noticeable because the model just does the most dumb things possible. "You have a bug in line 23." "Oh yes, this solution is bugged, let me delete the whole feature." That one-line fix I could make even with ChatGPT 3.5 can't just happen. Workflows that I use and are very reproducible start to flake and then fail. After a certain number of tokens per day, it becomes unusable. I like Claude, but I don't understand why they would do this.
- arcanemachiner 8mo agoRobbing Peter to pay Paul. They are probably resource-constrained, and have determined that it's better to supply a worse answer to more people than to supply a good answer to some while refusing others. Especially knowing that most people probably don't need the best answer 100% of the time.
- chrisjj 8mo ago> Especially knowing that most people probably don't need the best answer 100% of the time. More: probably don't know if they've got a good answer 100% of the time. It is interesting to note that this trickery is workable only where the best answers are sufficiently poor. Imagine they ran almost any other kind of online service such email, stock prices or internet banking. Occasionally delivering only half the emails would trigger a customer exodus. But if normal service lost a quarter of emails, they'd have only customers who'd likely never notice half missing.
- arresin 8mo agoRight. You can launder quantization that way by muddying the waters of discourse about the model.
- DanielHall 8mo agoI encountered the same situation too; Claude has 'become lazy'.
- WhitneyLand 8mo agoFirst off, this is a cool project, look forward to some interesting insights. I would suggest adding some clarification to note that longer measure like 30 pass rate is raw data only while the statistically significant labels apply only to change. Maybe something like Includes all trials, significance labels apply only to confidence in change vs baseline.
- taf2 8mo agoany chance we can get something like this for codex cli that'd be cool too compare
- esafak 8mo agoFinally someone did it! We need this for all models.
- Topfi 8mo agoI have yet to experience any degradation in coding tasks I use to evaluate Opus 4.5, but I did see a rather strange and reproducible worsening in prompt adherence as part of none coding tasks since the third week of January. Very simple queries, even those easily answered via regular web searching, have begun to consistently not result accurate results with Opus 4.5, despite the same prompts previously yielding accurate results. One of the tasks that I already thought was fully saturated as most recent releases had no issues in solving it was to request a list of material combinations for fabrics used in bag constructions that utilise a specific fabric base. In the last two weeks, Claude has consistently and reproducibly provided results which deviate from the requested fabric base, making the results inaccurate in a way that a person less familiar with the topic may not notice instantly. There are other queries of this type for other topics I am nerdily familiar with to a sufficient degree to notice such deviations from the prompt like motorcycle history specific queries that I can say this behaviour isn't limited to the topic of fabrics and bag construction. Looking at the reasoning traces, Opus 4.5 even writes down the correct information, yet somehow provides an incorrect final output anyways. What makes this so annoying is that in coding tasks, with extensive prompts that require far greater adherence to very specific requirements in a complex code base, Opus 4.5 does not show such a regression. I can only speculate what may lead to such an experience, but for none coding tasks I have seen regression in Opus 4.5 whereas for coding I did not. Not saying there is none, but I wanted to point it out as such discussions are often primarily focused on coding, where I find it can be easier to see potential regressions where their are none as a project goes on and tasks become inherently more complex. My coding benchmarks are a series of very specific prompts modifying a few existing code bases in some rather obscure ways, with which I regularly check whether a model does severely deviate from what I'd seen previously. Each run starts with a fresh code base with some fairly simple tasks, then gets increasingly complex with later prompts not yet being implemented by any LLM I have gotten to test. Partly that originated from my subjective experience with LLMs early on, where I found a lot of things worked very well but then as the project went on and I tried more involved things with which the model struggled, I felt like the model was overall worse when in reality, what had changed were simply the requirements and task complexity as the project grew and easier tasks had been completed already. In this type of testing, Opus 4.5 this week got as far and provided a result as good as the model did in December. Of course, past regressions were limited to specific users, so I am not saying that no one is experiencing reproducible regressions in code output quality, merely that I cannot reproduce them in my specific suite.
- fernvenue 8mo agoThat will be great if there's RSS support.
- rplnt 8mo agoThe chart would benefit from having weekends highlighted. Or have another chart averaged by a weekday.
- copilot_king 8mo agoThis strategy seems inspired by TikTok's approach for retaining new uploaders. TikTok used to give new uploaders a visibility boost (i.e., an inflated number of likes and comments) on their first couple of uploads, to get them hooked on the the service. In Anthropic/Claude's case, the strategy is (allegedly) to give new users access to the premium models on sign-up, and then increasingly cut the product with output from cheaper models.
- chrisjj 8mo agoYes, but the difference is TikTok didn't sell a particular service version. Anthropic did sell a particular model version.
- maximgeorge 8mo ago[dead]
- dmos62 8mo agoLack of transparency as regards "thinking power"-consistency is a big gripe of mine with LLM providers. It's even worse with ChatGPT and the like. E.g. I had to learn the hard way that at >45k input tokens ChatGPT 5.2 Thinking Extended bumps its intelligence down so hard that it can't follow basic instructions (or it somehow truncates the input, losing the instructions). It sucks to lose confidence in an otherwise great tool. I would 100x prefer being forced to back-off, or getting a straight-no, than getting silently downgraded. Transparency is a big deal.
- judahmeek 8mo agoSounds like you ran into the Maximum Effective Context Window: https://arxiv.org/abs/2509.21361?context=cs.AI https://arxiv.org/abs/2509.21361?context=cs.AI
- dmos62 8mo agoInteresting article. Not sure it's the same phenomenon. What I experienced was like a day and night difference when you go from 44.5k to 45.5k. Didn't notice any fluctuation to suggest that it's no a hard 45000 limit. I ran many many queries, similar problem space, but the problems varied a lot.
- parquor 8mo agoDoes this use a claude subscription or key, and has the account been used for anything else that day? On HN a few days ago there was a post suggesting that Claude gets dumber throughout the day: https://bertolami.com/index.php?engine=blog&content=posts&detail=insidious-progressive-intelligence https://bertolami.com/index.php?engine=blog&content=posts&de...
- sd9 8mo agoI’m sure there is not enough data here for this to be statistically significant (it seems to oscillate too much and not show real trends or step changes) - BUT If this measure were hardened up a little, it would be really useful. It feels like an analogue to an employee’s performance over time - you could see in the graphs when Claude is “sick” or “hungover”, when Claude picks up a new side hustle and starts completely phoning it in, or when it’s gunning for a promotion and trying extra hard (significant parameter changes). Pretty neat. Obviously the anthropomorphising is not real, but it is cool to think of the model’s performance as being a fluid thing you have to work with, and that can be measured like this. I’m sure some people, most, would prefer that the model’s performance were fixed over time. But come on, this is way more fun.
- elmean 8mo agoI KNEW I WASNT CRAZY
- wendgeabos 8mo agoCodex is doing better. Why is everyone silent on Codex? https://marginlab.ai/trackers/codex/ https://marginlab.ai/trackers/codex/
- drc500free 8mo agoWhat makes the level they chose a “baseline,” against which it would be appropriate to do statistical tests?
- PlatoIsADisease 8mo agoPretty sure someone at Google, OpenAI, and Anthropic met up at a park, leaving their phones in their car, and had a conversation that January 2026, they were all going to silently degrade their models. They were fighting an arms race that was getting incredibly expensive and realized they could get away with spending less electricity and there was nothing the general population could do about it. Grok/Elon was left out of this because he would leak this idea at 3am after a binge.
- kittikitti 8mo agoThis is why I run my own models. All the inference providers do sneaky things behind the scenes. They will limit the output tokens, turn off attention layers, lower reasoning, or just use a completely different model. I'm actually surprised that Claude Code experienced this, as I've experienced this the least from API and coding agents.
- Rastonbury 8mo agowould be interesting to see what scores it's get when it is actually degraded via the status page, it gets degraded pretty often, so there's at least something to compare or to know at what point Anthropic declares degradation
- mannanj 8mo agoI wonder when I experience noticeably degraded model quality, ie opus, is it because my usage falls in the highest buckets and I’m being shadow limited or served worse versions of opus or is it because of actual server load/burden? It wouldn’t be the first time companies have secret shadow algorithms running to optimize things and wouldn’t it be obvious to limit power users as matter of cost/profit and not tell them. (See history of “Shadow ban” though that’s for different reasons)
- crazygringo 8mo ago> We model tests as Bernoulli random variables and compute 95% confidence intervals around daily, weekly, and monthly pass rates. Statistically significant differences in any of those time horizons are reported. They're going to need to provide a lot more detail on their methodology, because that doesn't make a lot of sense. From their graphs, they seem to be calculating the confidence interval around the previous value, then determining whether the new value falls outside of it. But that's not valid for establishing the statistical significance of a difference. You need to calculate the confidence interval of the difference itself, and then see if all the values within that confidence interval remain positive (if it excludes 0). This is because both the old and new measurement have uncertainty. Their approach seems to be only considering uncertainty for one of them. They should also really be more specific about the time periods. E.g. their graphs only show performance over the past 30 days, but presumably the monthly change is comparing the data from 60 to 31 days ago, to the data from 30 days ago until yesterday? In which case the weekly graph really ought to be displaying the past two months, not one month.
- steveBK123 8mo agoNew to me, but I am starting to infer that for those "in the know" it is common knowledge on HN that LLMs are purposely degraded over time to manage capacity/cost or fudge benchmarks... How do you actually use these in production pipelines in practice then? Are LLMs even well suited for some of the document parsing / data scrubbing automation people are throwing at them now?
- mrandish 8mo agoBenchmark tracking of cloud AI performance is going to be crucial going forward. Vendors are selling a service that by its nature is very difficult for customers to gauge day to day. How will I know if a code revision is ~2.5% less good today than it would have been yesterday? Or if queries during peak load hours use one less 'expert' in their MoE? Yet vendor's costs to deliver these services are skyrocketing, competition is intense and their ability to subsidize with investor capital is going away. The pressure on vendors to reduce costs by dialing back performance a few percent or under-resourcing peak loads will be overwhelming. And I'm just a hobbyist now. If I was an org with dozens or hundreds of devs I'd want credible ways to verify the QoS and minimum service levels I'm paying for are being fulfilled long after a vendor has won the contract.
- MORPHOICES 8mo ago[dead]
- biddit 8mo agoCall it what you will. But the experience is like you have a reliable coworker, but he randomly decides to take bong hits. "No no yeah bro no I'm good like really the work's done and all yeah sorry I missed that let me fix it"
- arresin 8mo agoI hope the author sees this: You have to test inter-day variation. Many have noticed a sudden drop off at certain times.
- trq_ 8mo agoHi everyone, Thariq from the Claude Code team here. Thanks for reporting this. We fixed a Claude Code harness issue that was introduced on 1/26. This was rolled back on 1/28 as soon as we found it. Run `claude update` to make sure you're on the latest version.
- isaacdl 8mo agoAnywhere we can read more about what a "harness issue" means? What was the impact of it?
- airstrike 8mo agoPretty sure they mean the issue is on the agentic loop and related tool calling, not on the model itself In other words, it was the Claude Code _app_ that was busted
- xnorswap 8mo agoOne thing that could be a strong degradation especially for benchmarks is they switched the default "Exit Plan" mode from: "Proceed" to "Clear Context and Proceed" It's rare you'd want to do that unless you're actually near the context window after planning. I pressed it accidentally once, and it managed to forget one of the clarifying questions it asked me because it hadn't properly written that to the plan file. If you're running in yolo mode ( --dangerously-skip-permissions ) then it wouldn't surprise me to see many tasks suddenly do a lot worse. Even in the best case, you've just used a ton of tokens searching your codebase, and it then has to repeat all that to implement because it's been cleared. I'd like to see the option of: "Compact and proceed" because that would be useful, but just proceed should still be the default imo.
- persedes 8mo agoWhat would be cool if this somehow could do a comparison by provider. E.g. in the last outages anthropic models running on vertex were apparently less affected than those deployed elsewhere. (Not saying that one is better than the other, but would be a neat read out).
- account266928 8mo agoPlease try to make this statistically rigorous. There's lots of advice in this thread (intraday variation, etc) but if Im reading this right it looks like the CI includes the baseline value yet you still label this as failing. Wouldn't this just be "our test isn't powerful enough to find a signal if there were one here?" People will see this and derive strong conclusions that the data don't support and you, `qwesr123`, or "JB" from your blogs, will be responsible.
- _zachs 8mo agoThis is super important - even if it's not currently the best measure of degradation yet. Anecdotally, Opus 4.5 has gotten so bad for me it's almost adding time to my workflow instead saving it. It'd be nice to have more 3rd party measurements like this to hold Anthropic accountable.
- snissn 8mo agothey should run their test against a control baseline such as an open source hosted model to see the overall drift in their test
- motoboi 8mo agoI’d love to see, based on the level of non-determinism perfomance on the benchmark how many times you need to run the benchmark for the change to be relevant (or statistically significant if you want). That would be a nice paper.
- willturman 8mo agoCould this be (partially?) explained by Model Collapse [1], i.e. iteratively training on data that includes an ever increasing amount of AI slop? [1] https://thebullshitmachines.com/lesson-16-the-first-step-fallacy/index.html#:~:text=Model%20collapse. https://thebullshitmachines.com/lesson-16-the-first-step-fal...
- jonawesomegreen 8mo agoI’ve noticed Claude has been noticeably worse over the last week. For example, it told me I should pass frozen to make my Enum immutable—that’s not a thing. (It is a thing for dataclasses, but not for Enums.) That’s a pretty basic language feature it was nailing until recently. It also suggested I parse a URL using urlparse in a function that already uses urlparse. These are basic mistakes it wasn’t making before. Something seems to have changed, but I’m not sure what.
- hn_user_9876 8mo ago[dead]
- aorist 8mo agoIf the confidence interval width is 2 * 14.0%, how are you detecting a statistically significant difference between 58% and 50%? The 95% CIs on both timeseries pretty much always cover the baseline number, which is not consistent with the result being statistically significant.
- ed_mercer 8mo agoI would pay 300 for a non-degrading Max plan.
- macinjosh 8mo agoThe degradation does not need to be in the inference it can be in how often inference is used. It is closed source but the algorithms that decide what Claude code does when, could behave differently when the API responses are slower. Maybe it does fewer investigatory greps or performs fewer tasks to get to “an” answer faster and with less load.
- sandeepkd 8mo agoTotally tangential to article, was browsing through the website UI - https://marginlab.ai/explorers/swe-bench-pro/ https://marginlab.ai/explorers/swe-bench-pro/ , the page gives impression that the language, category boxes are selectable. However they are not a dropdown. Not sure if it was intentional design by human or some smart code generation by Claude based on the design sketches.
- stergd 8mo agoI rarely complain about model performance, but Opus 4.5 behaves as Sonnet 4 at best. Need to start testing alternatives asap
- devonkelley 8mo agoRunning agents in production, I've stopped trying to figure out why things degrade. The answer changes weekly. Model drift, provider load, API changes, tool failures - it doesn't matter. What matters is that yesterday's 95% success rate is today's 70%, and by the time you notice, debug, and ship a fix, something else has shifted. The real question isn't "is the model degraded?" It's "what should my agent do right now given current conditions?" We ended up building systems that canary multiple execution paths continuously and route traffic based on what's actually working. When Claude degrades, traffic shifts to the backup path automatically. No alerts, no dashboards, no incident. Treating this as a measurement problem assumes humans will act on the data. At scale, that assumption breaks.
- sd9 8mo agoLLM generated comments are so obvious, please just talk from your personal experience. Nobody cares about this imagined experience.
- carterschonwald 8mo agoive seen degraded reasoning levels that feel like they they might be blur from excess quantization. cause thats what you get from the grid changes
- threethirtytwo 8mo agoDoes this even make sense? Clearly anthropic won't release a model unless it passed a benchmark of some sort that proves it's better than the previous model... or else why would they even release it? It's obvious if this thing shows degradation, than there is another thing that is showing improvement.
- lighthouse1212 8mo ago[dead]
- your_friend 8mo agoThey should add testing from different ips and account countries, that would be fun too see that Americans are getting different models for example
- sreekanth850 8mo agoTried Kimi 2.5 and far ahead of claude for coding.
- cleifer 8mo agoHow much influence have you all found prompting to have on output quality? Generally I've been approaching by just describing my problem and assuming that I'll get the machine's optimal output, but perhaps being explicit in the prompt can impact the output quality?
- foerster 8mo agoIt definitely felt less capable recently, I thought I was imagining it, but it was noticeably more difficult to get it to help on tasks that usually aren't so hard.
- pojzon 8mo agoIm using Claude daily. Mostly delegating boring stuff I can do myself but its a waste of my time now. I store my prompts, so I know I often run the same task multiple times over weeks span. After working with it for pas half a year I have to say the quality pf responses is steadily going down. Feels like cost optimizations. Overall the worse it performs the more stuff I have to do myself, because I won’t waste time tweaking instructions every time it happens. It wpulf waste too much of that time. So seems we are swinging back the pendulum.