10 ms·
Google cloud outage
- qmarchi 7y agoHeyo Googler here. The problem was a mix between another cloud provider and GCP. Dare I say, there should be little customer impact as of 13:37 PST..... The status dashboard is going to be your best idea on information.
- the-dude 7y agoThis can't be real.
- svacko 7y agoIs the another cloud provider AWS? I could see tons of connection timeoutes between GCP & S3/Elasticsearch service. Hope everything is resolved now for good.
- judge2020 7y agoSeems AWS, connection to gmail's smtp relay also started timing out.
- deleted 7y ago[deleted]
- gigatexal 7y agoOh man I had no idea the big cloud providers have dependencies on other clouds like this.
- dodobirdlord 7y agoGiven how much trans-continental/trans-oceanic network cable the major cloud providers own, they almost certainly have special trans-cloud network traffic infrastructure. Especially since so much of "The Cloud" is within a few 10s of square miles in a field in Virginia. I can easily see how one provider could majorly disrupt another provider by accidentally breaking inbound traffic on one of those links.
- qmarchi 7y agoThe bigger issue is that there's a lot of customers where they have split cloud deployments, which means the customers hurt even if they are stable within the clouds themselves.
- thedance 7y agoIf you are deployed in such a way that both GCP and AWS need to be up you're doing it backwards. Multi-cloud strategy is supposed to result in the intersection of cloud failures, not the union of them.
- gigatexal 7y agoYeah, I see that now. Makes total sense.
- lima 7y agoThey do not, according to the dashboard, this issue merely affected connectivity between GCP and other cloud providers. There was a different outage yesterday, which has nothing to do with the one discussed in this thread.
- unbeli 7y ago[removed]
- packetslave 7y agoNext time YOU are about to spout off about something, perhaps think about reading the f'ing page being linked to? "The issue with connectivity between the GCP us-east1, us-east4, and us-central1 regions to other Cloud Providers has been resolved for all affected projects as of Friday, 2020-03-27 13:37 US/Pacific."
- deleted 7y ago[deleted]
- nammi 7y agoWe were seeing timeouts in east-1. I don't know what "normal" looks like, but Pingdom's map seems to show the whole east coast as affected https://livemap.pingdom.com/ https://livemap.pingdom.com/
- svacko 7y agoyeah, our GKE pods running in us-east1 were dying ~90minutes ago like crazy... hope they are gonna resolve this soon. not the luckiest day for Google, nor us
- tagux 7y ago"We had a router failure in Atlanta". WHAT? You kidding us? Urs Hölzle, technical infrastructure at Google Cloud senior vice president, said, "We're very sorry about that! We had a router failure in Atlanta, which affected traffic routed through that region. Things should be back to normal now. Just to make sure: This wasn't related to traffic levels or any kind of overload, our network is not stressed by COVID-19."
- ocdtrekkie 7y agoWas it like... a hardware failure? If you serve more than 100 people you probably should have redundant routers. Was it a configuration issue that replicated over to multiple devices at least, I hope?
- AdamJacobMuller 7y agoNot that simple as you sometimes need to manually isolate the faulty hardware and remove it from service.
- toast0 7y agoHave you worked with redundant routers? They certainly reduce the number of outages, but sometimes the hardware (or software) fails in exciting ways that doesn't engage the redundancy, or doesn't engage it properly, and you still get an outage (or you get an outage that wouldn't have happened). Or sometimes, one circuit is out of service for repair or upgrade, and the other circuit is connected to the router that failed. And routing for the AS that travels on that circuit was set not to fallback to transit because the last time that happened, it caused major issues. I have no specific knowledge of today's events, but this sort of thing happens. You can get the number of incidents down pretty low, but not to zero.
- ocdtrekkie 7y agoI have. I am just highlighting that the problem surely should be more complex than described. Or that their redundancy for these events was not adequately devised.
- kgraves 7y agoThis is extremely concerning as somebody looking to move or build on top of GCP for the long term. I wonder why anyone would choose GCP if outages are occurring on a regular basis.
- x__x 7y agoI was bummed out when Siteground moved all their cloud accounts over G, without telling their customers beforehand
- optimal_alex 7y agoI feel bad for these devs. For years they went around proclaiming they had great engineers, but not many people saw their tech, so nobody could verify this. Then they made one engineering product with public customers and it fails miserably compared to the competition. That must hurt on the inside knowing they are frauds.