30 ms·
Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both.
by strictnein 1mo ago
Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both.
Also, OpenAI is saying what caused it:
> "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"
Anthropic stated their issue started earlier:
> "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”
I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.
- owaiswiz 1mo ago[flagged]
- deleted 1mo ago[deleted]
- Starlevel004 1mo ago> I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime. And not only that, when one goes down a bunch of API traffic switches over to the other, spiking demand and knocking it down.
- computerex 1mo agoDo you know the probability of all these companies being down at precisely the same time?
- jeffbee 1mo ago"precisely the same time" meaning 80 minutes apart?
- deleted 1mo ago[deleted]
- ceautery 1mo agoYesterday it was 100%
- schiffern 1mo agoIf it's a "thundering herd" problem where everyone's harness falls back to less popular providers that don't normally see that much demand, I'd say the probability is pretty good. Classic cascading failure is consistent with providers failing 80 minutes apart instead of simultaneously.
- computably 1mo agoIf you ballpark it as a single 3 hour downtime window per week and iid Poisson, then overlapping downtime probability of 2 providers is approximately the expected occurrence rate per 3 hours, 1/56. Not particularly surprising at all.
- dgellow 1mo agoPretty high?
- strictnein 1mo ago"All of these" is two. OpenAI had a router issue. Anthropic had a separate issue. Anthropic uses a lot of SpaceX compute, so an Anthropic issue and a SpaceX issue can be one in the same, as was likely the case this time. And they weren't down at precisely the same time. Anthropic's issue started ~1 hour before OpenAI's.
- VBprogrammer 1mo agoIn my experience I've seen plenty of failures caused by user behaviour in these type of cases. Biggest competitor goes down and all of a sudden you have a lot more traffic...
- deleted 1mo ago[deleted]
- DANmode 1mo agoNoting their shared infra feels far from an extraordinary explanation. In fact, it feels pretty ordinary.