5 ms·
Google Cloud Networking Incident
- gtaylor 9y agoJust to verify my understanding: * This is a multi-region (us-east1, us-central1, europe-west1, asia-northeast1 and asia-east1) outage (yeesh) * Incident began at Incident began at 2017-08-29 18:35 PDT. At the time of this comment, we're over 13 hours (??) Correct?
- leesalminen 9y agoFor what it's worth, my load balancer in us-central-1 is unaffected.
- Thaxll 9y agoDoesn't seems to affect that much?
- Danack 9y agoI can't comment as to multi-region, but we saw issues last night at around 6pm UTC, which look similar to the ongoing issues we are seeing.
- jlgaddis 9y agoThe status page linked to here says that the issue spans multiple regions (across multiple continents!).
- wwayer 9y agoIt's now past 12:00 US/Pacific and the incident is still ongoing. I'm thinking that they don't have as many customers as they'd like us to believe. There's been very little discussion and no outrage. If AWS had an incident lasting this long, there would be a lot more noise.
- user5994461 9y agoIf AWS had an incident this long, there would be no noise and no public post about it. And the customer support would keep repeating have you try buying bigger instances?
- cloudcomp 9y agoI think you're alone on this opinion
- tostaki 9y agoFor anyone wanting a quick workaround, try removing all the backend of your load balancer then add them again. No idea why but it did work for us.
- vrobert78 9y agoWe have alerts since 11:30 UTC. We are using GKE. At 15:10 UTC, we duplicated one Kubernetes Service in a new one. Since, we have no alerts. Coincidence ?
- ABS 9y agoI can confirm several people have had luck doing this (not us though :-( )
- ABS 9y agoAccording to Google's support (....) the workaround works but will likely fix it only temporarily. They are currently rolling out a change that should completely fix the issue and somewhere else I read they are rolling back a config change they did recently..
- Danack 9y agoWe've been having 'fun' with ongoing issues for a site since 6pm UTC yesterday, which got dramatically worse this morning...and having been recurring during the day. Having multiple hour outages makes me really want to go back to hiring a couple of physical servers in a rack somewhere.
- kazen44 9y agobut the cloud should have been far more redundant and easier to maintain!
- slackingoff2017 9y agoThe "cloud" is a hilarious failure at the main marketing point. Decentralized just means you have no idea where your servers are. Scaling is the only real selling-point of cloud and far more people think they need it than actually do
- nitinics 9y agoI think it all boils down to how you deploy your stuff. If you think Cloud is so massive that it is never a SPOF, you'll likely not meet your availability aspirations in some time in future. Cloud to me is also shared risk. I read - "Google cloud outage" as "multiple companies that rely on a shared infrastructure is not available ATM". The mindset should be to run your services on a distributed infrastructure with no SPOF. Leverage cloud, fog, racks, PCs, whatever resources you can, but diversify your content/service and be risk averse from failure of one kind.
- joshribakoff 9y ago
- ABS 9y agowe've been suffering from this since 3:07am UK time and it's sad to see google still hasn't learnt how to do support/communicate with their paying customers after all the flack they always take about this. The number of mistakes (silly and otherwise) keeps increasing and I'm not talking about the technical problem itself! ok, end of rant
- mdekkers 9y agogoogle still hasn't learnt how to do support/communicate with their paying customers You are confusing "don't know how" with "absolutely not incentivised to spend money on this"
- ABS 9y agonope I'm not :-) writing poorly on the status dashboard, messing/mixing up timestamps, changing them retroactively hoping people don't notice, forgetting and/or mistakenly swapping the names of affected regions, consistently writing "next update at x o'clock" and then invariably publishing the update several to 10+ minutes after o'clock are all mistakes done by whoever is already paid to update that page and communicate with customers.
- wwayer 9y agoOur load is not being "balanced", as not all backends are being utilized because of this.
- bdimcheff 9y agoThis is our experience as well. On one service with 3 backends, we lost all connections to one at about 0530UTC, then the second at about 1010. The third backend has been able to handle all of our traffic so far, but we're also seeing intermittent connection resets or failed connections.
- ninjakeyboard 9y agoYikes! that's a huge failure.
- notyourday 9y agoWhenever the sun starts shining, the clouds disappear.
- 0xbear 9y ago"A distributed system is one in which the failure of a computer you didn't even know existed can render your own computer unusable." — Leslie Lamport I think Leslie needs to update this for Cloud, however.
- pbarnes_1 9y agoDoes this affect the global load balancer too? We haven't seen any issues, so presumably not? Or is it just a subset?
- manigandham 9y agoStuff goes wrong everywhere. The only problem is thinking that any cloud vendor or service is 100% reliable.