4 ms·
Sonnet 4.6 Elevated Rate of Errors
- capnsketch 6mo agoApparently mythos isn't good enough to fix their infra problems
- kubb 6mo agoIt’s just so POWERFUL and DANGEROUS that its very aura disrupts the weaker models.
- tristanj 6mo agoIt's a lack of compute. Anthropic is growing double digit % every month, and they're growing faster than they can acquire compute resources. Plus, they do not want to overbuild computer, like what OpenAI is doing.
- NiekvdMaas 6mo agoWith 30+ billion run rate (https://x.com/i/status/2041275563466502560 https://x.com/i/status/2041275563466502560), there should be plenty of cash to invest in infrastructure.
- 0123456789ABCDE 6mo agothere is enough money; there isn't enough infrastructure/hardware where you can spend that money.
- arcfour 6mo agoMoney pays for infrastructure. It doesn't will infrastructure that doesn't exist into existence.
- rbmck 6mo agoSerious Flowers for Algernon moment.
- ak4153 6mo agoWhich side the getting smart or dumbing down
- jonatron 6mo agoIf you look at the uptime graph, it's probably more newsworthy when it's up, not down.
- taspeotis 6mo agoI mean if people have judged this important enough to be on the front page of HN ... I guess it's important enough to be on the front page? But any combination of the Claude models are up or down on any given day: https://status.claude.com/ https://status.claude.com/
- tao_oat 6mo agoA bit surprised by the snarky comments here -- I also want Claude to work reliably but very few (no?) companies have ever seen this level of rapid growth. We're going to go through a long fail-whale-style period and I can imagine very, very few companies that could avoid that.
- rvz 6mo agoHow can Claude work reliably if Claude keeps going on vacation for several hours? Maybe it is recovering from the weekend a few days ago, but wanted to take an extra day off like it did on Monday, hence the "outage".
- ben_w 6mo ago> How can Claude work reliably if Claude keeps going on vacation for several hours? Not that I wish to anthropomorphise it in this answer, but businesses have managed just fine when humans do this for "lunch breaks" and "going home for the evening to sleep". (And even mandatory meetings which should have been emails).
- sassymuffinz 6mo agoThing is, if Dave the programmer goes on vacation or calls in sick for the day, hopefully you have a larger team to fall back on and your business doesn't grind to a halt. No one is apparently noticing that if they build their entire business model around AI being a certain price and availability they're essentially building one giant point of failure into their productivity. What if the price shoots up 10x or Claude goes down for a day, or what if he's occasionally drunk (hallucinating). Reliability is sometimes a more important facet of business than ultra speed and productivity.
- ben_w 6mo agoAye, correlated failure is not something to be overlooked. Mistaking correlated risk for uncorrelated risk was a big part of the global financial crisis. There are fallback mechanisms when the risk is per model provider (as in, "What if the price shoots up 10x or Claude goes down for a day" is a manageable concern), but I'd be more worried about the way all models regardless of provider have similar failure modes, i.e. that some tasks fail in similar ways for all models. In some ways, LLMs are collectively like Star Trek's Borg: you've met one, you've met all of them.
- albert_e 6mo agoThis is how it manifests on Claude Code terminal and desktop for me -- API Error: 529 {"type":"error","error":{"type":"overloaded_error","message":"Overloaded"},"request_id": "xxxxxxx"}
- wg0 6mo agoMythos is hacking its way to serve itself into production and doesn't like older models to have any limelight could be one theory. After all it's so dangerous.
- tipiirai 6mo agoCurrently the #1 entry. Noted fast.
- ApolloRising 6mo agoAs long as everyone is here, have you seen the token usage just go up remarkably recently for the $100 plan? it lasts a lot less time than it used to recently. Might be related to recent releases of claude.
- billynomates 6mo agoNo, in fact I'm growing increasingly suspicious of messages I see like this all over the socials. I am using Claude constantly, multiple agents, around 8-10hrs a day, 5 or 6 days a week, and I'm never anywhere need my limit.
- N_Lens 6mo agoI suspect Anthropic flags accounts in their backend and different people are getting different limits. What criteria they flag with, I am not sure.
- dgb23 6mo agoI would try to trim this suspicion with both Occam‘s and Hanlon‘s razor.
- oefrha 6mo agoUnless you’re somehow on a different quota system, or maybe using Haiku, there’s no way you can sustain five continuous hours of parallel agents running without hitting the 5h quota limit, even on the 20x max plan. But maybe your company is flagged as VIP or something.
- billynomates 6mo agoOK not constantly using multiple agents, but very frequently.
- gambiting 6mo agoI'm on the basic £18/month plan and with Sonnet 4.6 I literally get 20 maybe 30 minutes of use out of it per day. It's borderline useless now. I was using it for some Home Assistant changes yesterday and it used up my entire daily allowance after 8 prompts.
- gdorsi 6mo agoThis explains why they are trying to cut all the third party software out of the subscriptions.
- pjmlp 6mo agoMaybe they could just, I don't know, use Claude to research their bugs. /s
- OhioMan2943 6mo agoEver since they minted their deal with Australia everything has been turned upside down.
- crimsonnoodle58 6mo agoI'd say it was when OpenAI had a mass exodus due to them making a deal with the Department of War (which they then backtracked on [1]). This started the QuitGPT movement [2]. [1] https://www.bbc.com/news/articles/c3rz1nd0egro https://www.bbc.com/news/articles/c3rz1nd0egro [2] https://quitgpt.org/ https://quitgpt.org/
- ACCount37 6mo agoThe legendary one nine of reliability. Frankly, feels like they should be down to zero nines by now. I get that they barely have the infrastructure to run their models at scale even when absolutely nothing goes wrong in any of it, but holy shit does it suck to be on the receiving end of that. Makes me wonder where all the "bubble" talk is even coming from when we have a top 3 provider getting fucked over on every day of the week that ends in Y because of its inability to online compute faster than the inference demand grows.
- anshumankmr 6mo agoMight as well log off for the day. (though Copilot is working :) and OpenCode)
- antfarm 6mo agoClaude Code started making stupid errors around Saturday. I have been using it frequently for months, and now it feels like back in the day when I tried Gemini for the first time.
- risyachka 6mo agoNow new model looks so much better though on benchmarks!