8 ms·
Elevated Errors in Claude.ai
- himata4113 7mo agoSeems to be the biggest outage yet. Might be related to power loss events in UAE timing is suspicious as more datacenters appear to be hit.
- lelanthran 7mo ago> Might be related to power loss events in UAE timing is suspicious as more datacenters appear to be hit. More datacenters? I thought it was just one.
- himata4113 7mo agoThe strikes are actually still ongoing afaik.
- lyu07282 7mo agoLatest news I could find on it: https://www.businessinsider.com/amazon-data-centers-middle-east-drone-strukes-us-iran-conflict-2026-3 https://www.businessinsider.com/amazon-data-centers-middle-e... > Two facilities in the United Arab Emirates sustained direct hits, while a third facility in Bahrain was damaged by a drone strike "in close proximity," Also to add context: AWS has contracts with the US military: "The Joint Warfighting Cloud Capability (JWCC) contract enables AWS to continue providing Department of Defense (DoD) customers with secure, reliable, and mission-critical cloud services." https://aws.amazon.com/federal/defense/jwcc/ https://aws.amazon.com/federal/defense/jwcc/ Making them a target for retaliation ofc.
- himata4113 7mo agofriends in the middle east have said that there have been a few missiles flying overhead, possibly reduced media coverage as it is an ongoing operation.
- deleted 7mo ago[deleted]
- kshacker 7mo agoIf you look at their status page, something has been bubbling for the past week https://status.claude.com https://status.claude.com
- himata4113 7mo agoNever noticed it being outright down like this except for today (and yesterday), never had actual downtime except for few failed requests that worked after a retry which coincides with AWS datacenters going offline.
- kube-system 7mo agoA not particularly large AWS region on the other side of the world? Doubt it.
- himata4113 7mo agowell there has been pretty large deals going on in UAE especially when it comes to AI since they can get any power capacity with a flick of their fingers for an unbeatable price and the latency in AI doesn't really matter since the first token is usually seconds anyway. And it's not just AWS it's the entire region.
- upmind 7mo agoJarred (from Bun) said that a lot of the errors are being of how much they've scaled in users recently (i.e., the flock that came from OpenAI)
- fred_is_fred 7mo agoThe first scaling event was after their highly successful Super Bowl ad and the second was being on the right side of history over the weekend.
- dilyevsky 7mo agothis has been an issue for years at this point... other labs are hardly any better tho
- andreagrandi 7mo agoI must have missed something: why are people moving from OpenAI? Since they released gpt-5.3-codex I'be been using it and claude with opus-4.6 and Codex has always been better, more accurate, less prone to allucinations. I can do more with a 20$ OpenAI pland than with a Claude Max 100
- andkenneth 7mo agoPeople are mad at openAI cooperating with the pentagon while anthropic put their foot down over their red lines.
- direwolf20 7mo agoMore specifically OpenAI has agreed to be used for domestic mass surveillance and for autonomous (no human in the loop anywhere) drone attacks. ChatGPT will decide which building to destroy, and then it will be destroyed.
- CSMastermind 7mo agoPolitics, agreed Codex performs significantly better for me.
- rvz 7mo ago“98.92 % uptime” is horrendous and unacceptable. Only one 9 of availability means you are seriously unreliable.
- fred_is_fred 7mo agoThere are 2 9s in 98.92.
- cronelius 7mo agowell actually since 1 == 0.999999… and 98.82 is 98.91999999… there are an infinite number of 9s
- cr125rider 7mo ago“Wait you mean sequential 9s!? Here I was waiting for just the right time to turn it back on…”
- brookst 7mo agoI’m very proud of our 0.999999% uptime. Six nines!
- Tadpole9181 7mo agoOh come on guys, this one is at least funny.
- digitaltrees 7mo agounderrated...
- kshacker 7mo agoI was having an extended incognito chat with claude.ai, and then it stopped responding. I saved the transcript in a notepad and checked in another tab whether it was down. i wonder if the incognito session is gone, and whether by reposting it i can resurrect it. I have done so with Gemini but there it has codes like "Gemini said", which I do not see here. If anyone knows that, appreciate a solution.
- tayo42 7mo agoWho fixes the Ai when the Ai is down? Semi serious since they're pretty big on not writing code?
- brookst 7mo agoMost ops fixes don’t involve writing code though.
- kube-system 7mo agoThe same guy who used to fix stack overflow, presumably
- zvqcMMV6Zcr 7mo agoMaybe network guys can give some hints? I guess they encounter such issue relatively often, when they can't access network equipment by network to fix the network issue. I know management consoles have separate networks on datacenter scale but it isn't that easy with even bigger networks.
- raincole 7mo agoI know you say "semi serious" but you can't seriously think there isn't an LLM for internal usage only in Anthropic, right.
- tayo42 7mo agoI'm not sure what's involved with serving these llms or if the infra could be completely seperate or not for an internal one.
- siliconc0w 7mo agoThey need to keep an emergency backup Claude to fix the production Claude when it goes down. (More seriously I wonder if they'd consider using Openai or Gemini for this purpose)
- bashtoni 7mo agoOpus and Sonnet are still working fine in AWS Bedrock (and probably Google Vertex), so they genuinely do have an emergency backup Claude they can use.
- codegladiator 7mo agoIsnt bedrock and vertex pass thru to anthropic servers ? I didnt know aws/google are deploying the actual models
- etothet 7mo agoAWS actually hosts the models. Security & isolation is part of the proposed value proposition for people and organizations that need to care about that sort of stuff. It also allows for consolidated billing, more control over usage, being able to switch between providers and models easily, and more. I typically don’t use Bedrock, but when I have it’s been fine. You can even use Claude Code with a Bedrock API key if you prefer https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html https://docs.aws.amazon.com/bedrock/latest/userguide/what-is... https://code.claude.com/docs/en/amazon-bedrock https://code.claude.com/docs/en/amazon-bedrock (I am not affiliated with AWS in any way. I’m just a user stuck in their ecosystem!)
- LostMyLogin 7mo agoI’ve been using Claude Code w/ bedrock for the last few weeks and it’s been pretty seamless. Only real friction is authenticating with AWS prior to a session.
- kube-system 7mo agoBedrock runs all their stuff in house and doesn’t send any data elsewhere or train on it which is great for organizations who already have data governance sign off with AWS.
- kelvinjps10 7mo agoBut code is solved?
- digitaltrees 7mo agoWhy do you assume this is a code issue? They were literally banned by DoD and then suddenly go down? There is at least a question to ask there, no?
- iso-logi 7mo agoI switched from OpenAI to Anthropic over the weekend due to the OpenAI fiasco. I haven't been using the service long enough to comment on the quality of the responses/code generation, although the outages are really quite impactful. I feel like half of my attempted times using Claude have been met with an Error or Outage, meanwhile the usage limits seem quite intense on Claude Code. I asked Claude to make a website to search a database. It took about 6 minutes for Claude to make it, meanwhile it used 60% of my 4h quota window. I wasn't able to re-find it past asking it to make some basic font changes until I became limited. Under 30 minutes and my entire 4 hour window was used up. Meanwhile with ChatGPT Codex, a multi-hour coding session would still have 20%+ available at the end of the 4/5 hour window.
- gentleman11 7mo agomight be location based? I've used claude a lot this week and had no downtime at all
- tvink 7mo agoYou're not wrong, for sufficient simple cases it's at a disadvantage. But once things get complicated, it wins by being the only thing that you can get to work without going insane. And yeah, any serious use completely assumes a Max sub.
- digitaltrees 7mo agoI have been using anthropic almost exclusively for a year, while trying other models, and this has literally never happened. I have NEVER experienced a downtime event. At most a random error in a chat but that is immediately solved on the subsequent request. I use the desktop app, the mobile app, the api with several apps in production that I monitor and reliability has never been an issue. I pay about $1500 per month on personal api use fyi.
- tmountain 7mo agoI’ve had semi regular downtime since I stayed using Claude about two months ago. I love it but I find it less reliable than alternatives. This is evidenced on their status page (regularly showing red bars).
- anonnona8878 7mo agokeeps going down. One more time and I'm moving to Codex. Or hell, I better go back to using my actual brain and coding, god forbid. Fml.
- tvink 7mo agoYou'll be back :)
- lambda 7mo agoPlease relearn to use your brain. I cannot imagine how you can properly supervise an LLM agent if you can't effectively do the work yourself, maybe slightly slower. If the agent is going a significant amount faster than you could do it, you're probably not actually supervising it, and all kinds of weird crap could sneak in. Like, I can see how it can be a bit quicker for generating some boilerplate, or iterating on some uninteresting API weirdness that's tedious to do by hand. But if you're fundamentally going so much faster with the agent than by hand, you're not properly supervising it. So yeah, just go back to coding by hand. You should be doing tha probably ~20% of the time anyhow just to keep in practice.
- winwang 7mo agoKind of agreed. I like vibe coding as "just" another tool. It's nice to review code in IDE (well, VSCode), make changes without fully refactoring, and have the AI "autocomplete". Interesting, sometimes way faster + easier to refactor by hand because of IDE tooling. The ways that agents actually make me "faster" are typically: 1. more fun to slog through tedious/annoying parts 2. fast code review iterations 3. parallel agents
- lambda 7mo agoYeah. I've been finding a scary number of people saying that they never write code by hand any more, and I'm having a hard time seeing how they can keep in practice enough to properly supervise. Sure, for a few weeks it will be OK, but skills can atrophy quickly, and I've found it's really easy to get into an addictive loop where you just vibe code without checking anything, and then you have way too much to review so you don't bother or don't do a very good job of it.
- cbracketdash 7mo agoAlready made the switch back to Codex :-)
- rosquillas 7mo agoI'm basing my next projects on the ability of Claude code to write code for me. This disruptions are scary.
- adithyassekhar 7mo agoCongrats you are vendor locked for skills.
- skeledrew 7mo agoThat's a pretty bad idea. No matter how good a product is, never become so reliant on it that it seriously affects things that matter if it becomes unavailable.
- mrguyorama 7mo agoIf your product is made using Claude why would I pay you for it when I can just make it myself? I have Claude access too
- mejutoco 7mo agoWindow cleaners still exist, and drivers, and bakers, etc. You get my point.
- thekid314 7mo agoYeah, the influx of people is disrupting my work, but it brings me joy to witness OpenAI’s decline in consumer support. So much for their Jonny Ive product, whatever it was.
- camillomiller 7mo agoI am so baffled that someone with the stature of Jony Ive fell prey to scam Altman empty promises. I would have expected much more of him.
- chihuahua 7mo agoAltman put all of his attribute points on lying.
- Sammi 7mo agoHe's a Bard with all his points in Charisma. He doesn't do anything except sing fairy tale songs. That's the prettier fantasy version. The other is that he is Gríma Wormtongue.
- rhubarbtree 7mo agoWhat were the empty promises?
- skywhopper 7mo agoSeriously? Jony Ive is in his Cash In era. He long ago stopped being relevant, and was a huge drag on Apple for a decade. He’s perfectly happy to take billions for doing nothing, I’m sure.
- PinkMilkshake 7mo agoI won't hate you for downvoting me, but this is heroin-grade schadenfreude.
- digitaltrees 7mo agoAnyone else find this timing odd given the DoD ban?
- davegardner 7mo agoI hope they improve their incident response comms in the future. 2.5 hours with nothing more than "We are continuing to investigate this issue" is pretty poor form. Their past history of incident handling looks just as bad.
- IsTom 7mo agoThey're waiting for claude to get up so they can use it to investigate why claude is down.
- mrguyorama 7mo agoTheir entire business is a premise of "Humans don't have to know how to do anything anymore" so what did you expect? Their brain is down. https://www.cs.ucdavis.edu/~koehl/Teaching/ECS188/PDF_files/Machine_stops.pdf https://www.cs.ucdavis.edu/~koehl/Teaching/ECS188/PDF_files/...
- adithyassekhar 7mo agoAre employees from Anthropic botting this post now? This should be one of the top most voted posts in this website but it's nowhere on the first 3 pages. Also remember, using claude to code might make the company you're working for richer. But you are forgetting your skills (seen it first hand), and you're not learning anything new. Professionally you are downgrading. Your next interview won't be testing your AI skills.
- AlexeyBelov 7mo ago> Your next interview won't be testing your AI skills Not that I disagree with your overall point, but have you interviewed recently? 90% of companies I interacted with required (!) AI skills, and me telling them how exactly I "leverage" it to increase my productivity.
- adithyassekhar 7mo agoAre they just looking for AI skills? If so that's terrifying.
- tmountain 7mo agoProbably, I think hand coding is going the way of the dodo and the ox cart.
- adithyassekhar 7mo agoSorry but focusing on the hand coding part misses the whole picture and would derail the conversation. Comparisons like that are often dishonest. Hiring someone who writes Rust with Claude but never written anything with it in their lives, never faced the edge cases, never took the wrong decisions feels naive to me. At the end of the day it's still a next token generator, an impressive one. It can hold context but not relate with anything outside that context. Someone needs to take accountability.
- ternwer 7mo ago
- adham-omran 7mo agoThe service has been inconsistent and/or down for the last 12 hours..
- AYBABTME 7mo agoThis right now today is making the case for OSS AI and local inference. 200$/m to get rate limited makes a RTX 6000 Pro look cheap.
- tmountain 7mo agoHow well do local OSS models stack up to Claude?
- sunaookami 7mo agoThey don't, only on meaningless benchmarks.
- Balinares 7mo agoVery well for narrowly scoped purposes. They decohere much faster as the context grows. Which is fine, or not, depending on whether you consider yourself a software engineer amplifying your output by automating the boilerplate, or an LLM cornac.
- wongarsu 7mo agoMuch better than they did half a year ago, but a single RTX 6000 won't get you there Models in the 700B+ category (GLM5, Kimi K2.5) are decent, but running those on your own hardware is a six-figure investment. Realistic for a company, for a private person instead pick someone you like from openrouter's list of inference providers. If you really want local on a realistic budget, Qwen 3.5 35B is ok. But not anywhere near Claude Opus
- Eisenstein 7mo ago> but running those on your own hardware is a six-figure investment GLM-5 is a 744B MoE with 40B active. You can run a Q4_K_M quant on llama.cpp if you can afford 512GB of RAM. An RTX 6000 will help a lot with the prompt processing, and the generations with be relatively fast if you have decent memory bandwidth. llama.cpp's autofit feature is really good at dividing the layers for MoEs to max speed when offloading.
- 7mo ago
- ramon156 7mo agoNext year: Anthropic to buy over OpenAI Datacenters
- re-thc 7mo agoThey have some? Aren’t Oracle and other “friends” running it?
- gdorsi 7mo agoThis comes as reminder that software engineering is way more than generating code. We build systems that can fail in unpredictable ways, and without knowing the system we built deeply is hard to understand what's going on.
- nprateem 7mo agoI've been noticing elevated stupidity. "Do this" "User wants me to [do complete opposite]" Seems not to be as capable as a month ago.
- o10449366 7mo agoNo wonder. It's performance overall was noticeably, like it had regressed to coding models from 1.5 years ago. I've try not to use claude during peak US hours because it tends to struggle more then with reasoning and correctness it seems than off hours.
- FrasiertheLion 7mo agoAI has normalized single 9's of availability, even for non-AI companies such as Github that have to rapidly adapt to AI aided scaleups in patterns of use. Understandably, because GPU capacity is pre-allocated months to years in advance, in large discrete chunks to either inference or training, with a modest buffer that exists mainly so you can cannibalize experimental research jobs during spikes. It's just not financially viable to have spades of reserve capacity. These days in particular when supply chains are already under great strain and we're starting to be bottlenecked on chip production. And if they got around it by serving a quantized or otherwise ablated model (a common strategy in some instances), all the new people would be disappointed and it would damage trust. Less 9's are a reasonable tradeoff for the ability to ship AI to everyone I suppose. That's one way to prove the technology isn't reliable enough to be shipped into autonomous kill chains just yet lol.
- gaigalas 7mo ago"It's fine, everyone does it"
- KronisLV 7mo agoThere's probably a curve of diminishing returns when it comes to how much effort you throw in to improve uptime, which also directly affects the degree of overengineering around it. I'm not saying that it should excuse straight up bad engineering practices, but I'd rather have them iterate on the core product (and maybe even make their Electron app more usable: not to have switching conversations take 2-4 seconds sometimes when those should be stored locally and also to have bare minimum such as some sort of an indicator when something is happening, instead of "Let me write a plan" and then there is nothing else indicating progress vs a silently dropped connection) than pursue near-perfect uptime. Sorry about the usability rant, but my point is that I'd expect medical systems and planes to have amazing uptime, whereas most other things that have lower stakes I wouldn't be so demanding of. The context I've not mentioned so far is that I've seen whole systems get developed poorly, because they overengineered the architecture and crippled their ability to iterate, sometimes thinking they'd need scale when a simpler architecture, but a better developed one would have sufficed! Ofc there's a difference between sometimes having to wait in a queue for a request to be serviced or having a few requests get dropped here and there and needing to retry them vs your system just having a cascading failure that it can't automatically recover from and that brings it down for hours. Having not enough cards feels like it should result in the former, not the latter.
- pmontra 7mo agoEmails with verification codes do not get delivered. > Have a verification code instead? > Enter the code generated from the link sent to [...] > We are experiencing delivery issues with some email providers and are working to resolve this. > Check your junk/spam and quarantine folders and ensure that support@mail.anthropic.com is on your allowed senders list. I'm still waiting for a code from one hour ago. Meanwhile I managed to fix my source code alone, like twelve months ago.
- ruszki 7mo ago> I managed to fix my source code alone, like twelve months ago. I’ve just mentioned to one of my friend yesterday, that you cannot do this anymore properly with new things. I’ve started a new project with some few years old Android libraries, and if I encounter a problem, then there is a high chance that there is nothing about it on the public internet anymore. And yesterday I suffered greatly because of this. I tried to fix a problem, I had a clearly suboptimal solution from myself after several hours, but I hated it, but I couldn’t find any good information about it (multi library AndroidManifest merging in case of instrumented tests). Then I hit Claude Code with a clear example where it fails. It solved it, perfectly. Then I asked in a separate session how this merging works, and why its own solution works. It answered well, then I asked for sources, and it cannot provide me anything. I tried Google and Kagi, and I couldn’t find anything. Even after I knew the solution. The information existed only hidden from the public (or rather deep in AGP’s source code), in the LLM. And I’m quite sure that I wasn’t the only one who had this problem before, yet there is no proper example to solve this on the internet at all, or even anything to suggests how the merging works. The existing information is about a completely separate procedure without instrumented tests. So, you cannot be sure anymore, that you can solve it by yourself. Because people don’t share that much anymore. Just look at StackOverflow.
- mejutoco 7mo agoIt looks like you could write that blogpost and get some traffic, on the other side. Very interesting how the flow has changed direction based on your example.
- 7mo ago
- evara-ai 7mo agoThis is a real operational problem when you're building client-facing automation systems on top of these APIs. I build chatbots, workflow automation, and AI agent systems for clients — and the hardest conversation is explaining that your system's uptime is fundamentally capped by your LLM provider's uptime. Patterns that have helped in production: 1. Multi-provider fallback. For conversational systems, route to Claude by default, fall back to GPT-4 on 5xx errors. The response quality difference is usually acceptable for the 2-3% of requests that hit the fallback. This turns a hard outage into a slight quality degradation. 2. Async queuing for non-real-time workflows. If you're processing documents, generating reports, or running batch analysis — don't call the API synchronously. Queue the work, retry with exponential backoff, and let the system self-heal when the API recovers. Most of our automation pipelines run with a 15-minute SLA, not a 500ms one. 3. Graceful degradation in real-time systems. For chatbots and voice agents, have a scripted fallback path. "I'm having trouble processing that right now — let me transfer you to a human" is infinitely better than a hung connection or error message. The broader issue: we're all building on infrastructure where "four nines" isn't even on the roadmap yet. That's fine if you architect for it — treat LLM APIs like any other unreliable external dependency, not like a database query.
- skillboss9901 7mo ago[dead]