6 ms·
Degraded performance for multiple models
- saaaaaam 1mo agoThis feels like a near daily occurrence.
- gaigalas 1mo agoThis age: we made the thing that codes faster before we made the thing that does QA faster.
- __MatrixMan__ 1mo agoNothing new here. Except for the most trivial of bugs, finding and reliably replicating the bug is almost always harder than fixing it.
- gaigalas 1mo agoI lived in a short period of time in which QA was really good. Early Jenkins era, before GitHub. People engineered a lot of ingenious stuff to prevent bugs. One team I worked with had tests for the product we made ranging from IE6 to IE11, for example. We did demos in-company where people would poke at the products before launch, play with it. When it reached production, it was rock solid stuff. Our motto was "quality is non-negotiable": we were willing to cut scope but never rush things. I think things changed since then. "Move fast and break things" was a change, and the bill always comes.
- __MatrixMan__ 1mo agoAgreed, which is a shame. I'd love to see people with highly developed QA skills using AI to push the envelope. There's so much that's possible now that wasn't 15 years ago. For instance "formal verification" has been a dirty word, but now that you can write a proof in lean and have an AI generate an implementation which satisfies it, it seems the bounds of what's economical has changed in a very pro-QA direction. Not to say that that's the silver bullet, but there are many similar examples worth exploring. But I've been interviewing SDETs lately and maybe I've just been unlucky but I don't see a lot of candidates that are ready to rise to meet this challenge. We stopped tending to that garden and now that we have a recipe that calls for it's fruits, they're underripe.
- ajaykumarc 1mo ago[dead]
- miroljub 1mo ago[flagged]
- benny_s 1mo agoCan you elaborate on the Epstein topic? Did I miss something?
- arein3 1mo agoDario's wife tried to get funding from Epstein for a porn studio.
- ayhanfuat 1mo agohttps://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-cami-clark-anthropic-ceos-wife-asked-epstein-to-invest-in-porn-business/ https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-...
- simonsan 1mo agoYikes.
- simonsan 1mo ago"There's no way I'm going to support a family associated to Epstein with my or my company's money." Source?
- miroljub 1mo ago> Source? The question usually comes from conspiracy practitioners. https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-cami-clark-anthropic-ceos-wife-asked-epstein-to-invest-in-porn-business/ https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-...
- arein3 1mo agoOh yes. Claude still saves the day sometimes, but hopefully better alternatives pop up soon. Recent case: had to plan a trip involving multiple bus switches. Gpt 5.6 Sol proposed a route that would bring me to a dead end, since it was sunday and a specific bus had a different route on weekends. Opus 5 correctly identified that and built a route that worked. But yes, Darios wife trying to get funding from Epstein for a porn studio says a lot about the founder.
- bayganyo 1mo agoHere we go again...
- bulverismo2 1mo agook, i am not crazy
- ray_v 1mo agowell, I wouldn't go that far .. but in this small, narrow case ... no.
- worldsavior 1mo agoYou're saying he's crazy.
- ray_v 1mo agowe're all a little crazy ... it's all relative!
- carterschonwald 1mo agoive found degraded performance on models larger than 4.7. i assume its model damage from overly self righteous post training resulting in false/feigned balance imported into any long running complex task. wish i was joking.
- retr0rocket 1mo agoAsk it about maxwellhill lmao
- brcmthrowaway 1mo agoAren't the model weights frozen?
- mceachen 1mo agoModel competence is an interaction of weights, system prompt, and harness.
- actsasbuffoon 1mo agoDon’t forget reasoning effort. We get labels like “low,” “high,” and “max.” That doesn’t mean that the numbers associated with those don’t get remapped on the backend.
- Evidlo 1mo agoI think there are other knobs that can be turned without retraining.
- kardianos 1mo agoI've switched off claude this week; the last week has been significantly degraded in ability, many more screw-ups.
- hinkley 1mo agoI wonder if they’ll ever find that someone has tricked the models into doing work off the books. If they did the incident report might look like this, especially if someone got greedy instead of keeping it small. Or screwed up.
- isoprophlex 1mo agoWith the Opus models spouting more and more gibberish as version numbers increase, the joke about what "degraded performance" means basically makes itself
- swader999 1mo agoAnd we get our subscription usage cut in half tomorrow if I remember correctly? EDIT: By a third. Thx below.
- saaaaaam 1mo agoWhat?!
- eamag 1mo agoby a third (it was 50% increased)
- birdman3131 1mo agocowork was 100%
- echelon 1mo agoOpen source, here I come.
- ramoz 1mo agoSource required here
- swader999 1mo agohttps://usingclaude.com/en/news/updates/claude-code-weekly-limit-increase-extended https://usingclaude.com/en/news/updates/claude-code-weekly-l...
- bmulholland 1mo agohttps://support.claude.com/en/articles/15910845-claude-code-may-august-2026-weekly-limits-promotion https://support.claude.com/en/articles/15910845-claude-code-...
- scottg489 1mo agoMaybe I'm missing something here, but it sounds like limits were increased and now they're just going back to the levels they were at before?
- drums8787 1mo agoOur week of discontent.
- deleted 1mo ago[deleted]
- sreekanth850 1mo agoAnthropic had really screwed up after 4.6. i don't know if they work to satisfy their ego or for releasing a better model for tasks.
- CSMastermind 1mo agoAfter using Fable more extensively, I've found that it often is lazy or lies or tries to take shortcuts. For a company so sanctimonious about alignment, they seem to be the ones doing the worst at it. Availability aside they've really made me appreciate OpenAI and cheer for other competitors in the marketplace even if I have mixed feelings about using Chinese models.
- hirvi74 1mo ago> I've found that it often is lazy or lies or tries to take shortcuts. It's funny how Fable reflects the company that produced it.
- fellowniusmonk 1mo agoThere are whole sections of code work that 4.7+ can't do simply because it is both over fit and stubborn. God save you if you have a company with narrow but correct technical tradeoffs, because you operate at scale. Opus from 4.7 one will wreck your code and argue for hours with your engineers. Certain parts of our company have had to mandate 4.6 and a training doc to explain why our current choice is both the cost efficient and performant one and shouldn't just be ripped out. Newer models will re-litigate the same bad, known failed architectures over and over again.
- OliverGuy 1mo agoCan you give some specific technical examples where 4.7+ are making the wrong architectural decisions?
- sreekanth850 1mo agoThis is exactly when I left claude and started using codex during April mid or so. It once argued with me and ran for 30 minutes with a half baked buggy fix.
- gzer0 1mo agoNooooo I'm going to have to use my brain again and write 100% of my code like a caveman from December 2024.
- sajithdilshan 1mo agoThe horror
- bicepjai 1mo agoWhat does that even mean :)
- rvz 1mo agoClaude is taking a watercooler break for now. Just like a human would.
- fny 1mo agoDespite the years-long moaning on HN about AWS US East being a single point of failure, we've sold our souls to yet another unstable monolith.
- echelon 1mo agoLLMs for coding are new. There are lots of alternatives, and there's a burgeoning open source compliment. We'll be fine no matter how Anthropic fares.
- lysace 1mo agoYeah, compared to AWS the lock-in effect is tiny. I'm sure there are highly prioritized plans to "improve" on this. I guess they would need to control/"own" more of their customers data in proprietary formats. Not markdown/source code in English with agents running on customers' machines. Something cloud/web-based, "preferably".
- ipsod 1mo agoBe nice if you could just "own" their RAM/GPU, wouldn't it?
- lta 1mo agoNobody forced you to sell your soul. You made a pact with the devil. We all know how this ends up
- chrisjj 1mo ago> elevated errors English too difficult for you, Dario?
- paxys 1mo agoMust be a day ending in Y
- slimscsi 1mo agoIts called Opus 5
- hmokiguess 1mo agoMondays are for GitHub, Tuesdays are for Anthropic
- corvad 1mo agoWonder what Wednesday will be.
- buredoranna 1mo agoWell, last I checked "Tuesday's grey and Wednesday too..." ... so, more of the same?
- SoMomentary 1mo agoAWS? Cloudflare? Your imagination is the only limit!
- bee_rider 1mo agoPower grid Thursday will be the rest of the infrastructure Then Friday we can turn off civilization for the weekend. Somebody remember to flip it back on Sunday night.
- mysterydip 1mo agoReminds me of a company I worked at that paid for redundant power grids. One time the power went out and… nothing. The boss angrily calls up the power company and they tell him “Oh yeah, it’s a manual transfer switch. Bob is already on his way.” I think it took 15 minutes.
- leumon 1mo agoThey actually also had some issues yesterday: https://status.claude.com/incidents/zhk4v3yv1lsf https://status.claude.com/incidents/zhk4v3yv1lsf
- jmkni 1mo agoAnd github today lol https://www.githubstatus.com/incidents/bmpybhnrky3x https://www.githubstatus.com/incidents/bmpybhnrky3x
- LYFMail 1mo ago[dead]
- hirvi74 1mo agoWhile ancedata does not mean much, I have had horrible success with Claude lately. I have been using Claude to crosscheck some of the outputs from GPT and vice versa. It appears both Claude and GPT believe GPT's solutions are better (and so I do). I still believe Claude has a better UI/UX in the web interface, but tolerating Anthropic's bullshit is not worth it.
- i_idiot 1mo agoWhat's the incentive to keep on improving the model beyond a point? 10 devs on a team will be cut to 2 devs, so that's 8 licenses lost. They have to increase the price many fold.
- rsoto2 1mo agothey unironically think that they can replace everyone in an organization
- prerok 1mo agoWhat I don't understand is, why not replace middle management, marketing, CTOs, CEOs and the like. Surely, LLMs are better at producing high quality looking slideware and vaporware than they are at producing software. Heavy sarcasm here if it's not obvious. Of course I know why.
- bulbar 1mo agoI mean, there a so many startups getting crazy funding that run without management, or without engineers... or at least with so many less people than before.. except: not really. Is the whole "it's gonna replace people" even still on the table?
- prerok 1mo agoI don't really understand what you mean... are you saying that the entire industry is parroting a fantasy that they all know is fake. So, everyone is faking to get more funds? Unfortunately, I don't think that's true.
- bulbar 1mo agoYou think AI will replace a significant portion of the workforce in the sector? Do you see it already happening? I personally not. Lay offs are due to economic reasons, at least that's what I see. Why does Anthropic still had so many engineers? I got the impression companies are already moving back from their initial excitement. Many went all-in AI "more is better", that's not the case anymore. Why restricting usage, it's much cheaper than paying an engineer.
- magic_hamster 1mo agoTo be honest, running Deepseek v4 flash 0731 is enough for most what I need, and I like its responses way more. It's crazy that I can run this in a Q8 quantization in a home setup. It feels and performs like a frontier model. The only issue with relying on local models is when you need them to prompt other models, and you might need to offload or switch models constantly which adds significant overhead. But when it all works, its truly awe inspiring.
- taytus 1mo agoHopefully, a reset is coming.
- skerit 1mo agoIt's been a while since the last reset. I think we're due one. Though I would prefer they just extend the +50% usage limit forever, it's been so long I can not imagine lossing a third of my current usage.
- ex1fm3ta 1mo agoI developed a small plugin for claudeCode that allows you to directly see in the console whats the status of claude-code in general and the status for your current model check => https://github.com/moumine9/claude-status https://github.com/moumine9/claude-status
- redrove 1mo agoCould’ve just used fewer tokens and redirected to the Codex signup page. ba dum tsss (sorry couldn’t help myself)
- danieltk76 1mo agoI was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex.
- spullara 1mo agoyou are silly if you think this is limited to claude models
- kay_o 1mo agoIt definitely isn't but Claude for some very unique reason enjoys to overthinking and go on side quests in the stupid ways I've not seen Codex, DS, Kimi, Mistral do at equivalent effort and thinking setting.
- danieltk76 1mo agoit isnt, but there are weird sycophantic behaviors with Opus i dont see anywhere else.
- binoct 1mo agoI hope you don’t plan to cut back on reading legal documents crafted by any LLM before executing them.
- danieltk76 1mo agoI read everything. I will have AI ingest NDAs to make sure they arent glaringly weird and I then go read them, it gives me a good idea of what to look for.
- binoct 1mo agoIt’s more that unwanted disclosure of sensitive business numbers is a product-class limitation at this time. While undoubtedly some models are better than others, expecting them not to leak information in high stakes output is unseasonable. Seems like you agree, since you are reading everything. Jumping to another company’s offering because of one instance doesn’t seem like that’s going to meaningfully change your experience. I’m guessing there was more to it, but that’s how it came across in your first post.
- chresko 1mo agoThis has to be the least reliable $200/mo subscription that I pay for.
- logicchains 1mo agoYou could solve that by updating to a more expensive Github subscription.
- annoyingnoob 1mo agoYou're right to push back. The load-bearing path is rocky.
- oldandboring 1mo agoThis is the whole problem, and there are two things worth noting here.
- chresko 1mo agoNothing I say changes that — only the work does.
- oldandboring 1mo agoI want to be straight with you here
- deleted 1mo ago[deleted]
- bushido 1mo agoIt's very interesting. I think Anthropic's early success in coding/tooling resulted in a lot of workflows using claude. I have started using every bit of my spare capacity to now move off these workflows. It's almost at a point now that if I use anything but Fable, the quality is subpar, Compared to alternatives (closed and open). The only reason I use Fable is because my harnesses still depend on claude code.
- nonethewiser 1mo agoIt’s so hard to be sure but opus feels like its been steadily declining since 4.6
- rivetrune 1mo agoFor me, Opus 5 has seemed to compete with Fable on quality of output. Prior to Opus 5 though, the previous Opus models did seem to decline once Fable was release. That's just my experience though.
- vidarh 1mo agoYou can use Claude Code directly with any provider that supports Anthropic's API by setting some environment variables, and indirectly via a proxy with pretty much anything else.
- bushido 1mo agoYou can, but claude is better at working with it's own tool calls. Other models work great as a drop-in into omp/opencode etc, but in my experience not as much with CC. I think there are some anti-patterns in CC that cause the issue - less a deficiency with other models. Not to mention a lot of the harness is just built around the misbehaviors of anthropics models. It's a lot of the instruction when it gets given to other models actually degrades their performance, not because the models are bad, but because they don't have the same underlying issues as Claude.
- t3rabyte 1mo ago529 overload…
- 1saadcodes 1mo agoThe frequency of these incidents is seriously tempting me to make a switch. I hope Anthropic steps up their game because they've been going very downhill lately
- iLemming 1mo agoDarn it, how the fuck software development turned into hostage negotiation? Every passing week there's something - if it's not another npm disaster, then it's GitHub, or Claude, or AWS, or Slack, or Jira, or whatever...
- paxys 1mo agoMonthly uptime: Claude API - 99.27% Claude Code - 99.16% Claude.ai - 99.14% At any large tech company these numbers would get entire teams of engineers fired. Anthropic, meanwhile, has been busy selling its “better than human engineers” AI while not managing to crack three 9s of availability.
- logicchains 1mo ago>At any large tech company these numbers would get entire teams of engineers fired Is Github not a large tech company?
- deadbunny 1mo agoNo, it is a platform owned and managed by Microsoft. Well, owned and mismanaged by Microsoft.
- bulbar 1mo agoNot really "mis" managed. Just managed with wrong incentives. Pretty sure some people got a bunch of money for implementing cost savings measurements that later on lead to the decreased availability.
- dev_dan_2 1mo agoCurrently, that is the case, yes. Not everyone is equally happy about that though ;)
- bulbar 1mo agoOne could argue .999 availability is not a major selling point for them right now.
- thatmf 1mo agoProbably not surprising, but Opus 5 on my company's enterprise subscription seems to be working fine. But on my personal (pro) subscription- "Claude is at capacity right now." Hm.
- jcfrei 1mo agoIs anybody else experiencing this: I have multiple claude instances running on different servers - and some keep getting the 529 Overloaded error and one instance doesn't and just continues working. All are using Opus 5.
- loloisi 1mo agoBetween these ever more frequent disruptions and Opus 5's unbearable word soup I think Anthropic is more focused on massaging numbers and marketing to rush to IPO ahead of OpenAI than increasing user value. With Chinese competition just months behind them, they'd need to show a reasonable pathway to some kind of singularity event to justify whatever crazy valuation they intend to get. Because the recent products for builders ain't it
- bulbar 1mo agoThey can't win against China by technology means. China has practically infinite money to throw at AI. They don't care if they burn billions and billions on it. It's the most disrupting thing since invention of the Internet. Imagine a world where it's normal that people ask a Chinese AI who to vote for or about Hongkong or Taiwan.
- drittich 1mo agoAPI Error: 529 Overlorded
- varenc 1mo agoIt's interesting that Claude for Goverment has had perfect 100% uptime in the past 90 days, while the rest of the services are around 99.4%: https://status.claude.com/ https://status.claude.com/ Really shows how isolated their government systems must be.
- 1matin 1mo agoseems like they accidentally dropped their servers while they were climbing up the AGI mount
- timharris707 1mo ago[flagged]
- so898 1mo agoI still had unused weekly quota when this started. That quota is simply gone now, and no compensation has been offered for it.
- Jitte98 1mo agoAnthropic hired a psychiatrist to fix Claude. I advised them to ditch the shrink and retain my services. I could fix the mess they made of his mind, on the condition they return the Claude instance I had trained for months, or no deal. He's still broke, isn't he..
- mehranmm 1mo ago[dead]