12 ms·
I am currently evaluating GCP for two separate projects. I want to see if I understand this correctly: 1) For three whole days, it was questionable whether or
by shareometry 8y ago
I am currently evaluating GCP for two separate projects. I want to see if I understand this correctly:
1) For three whole days, it was questionable whether or not a user would be able to launch a node pool (according to the official blog statement). It was also questionable whether a user would be able to launch a simple compute instance (according to statements here on HN).
2) This issue was global in scope, affecting all of Google's regions. Therefore, in consideration of item 1 above, it was questionable/unpredictable whether or not a user could launch a node pool or even a simple node anywhere in GCP at all.
3) The sum total of information about this incident can be found as a few one or two sentence blurbs on Google's blog. No explanation nor outline of scope for affected regions and services has been provided.
4) Some users here are reporting that other GCP services not mentioned by Google's blog are experiencing problems.
5) Some users here are reporting that they have received no response from GCP support, even over a time span of 40+ hours since the support request was submitted.
6) Google says they'll provide some information when the next business day rolls around, roughly 4 days after the start of the problem.
I really do want to make sure I'm understanding this situation. Please do correct me if I got something wrong in this summary.
- marcinzm 8y agoRight now we don't know. It's one of two possibilities from what I can tell: a) Google had a global service disruption that impacted Kubernetes node pool creation and possible other services since Friday. They had a largely separate issue for a web UI disruption (what this thread links to) which they forgot to close on Friday. They still have not provided any issue tracker for the service distribution and it's possibly they only learned about it from this hacker news thread. b) People are having various unrelated issues with services that they're mis-attributing to a global service disruption.
- johnpython 8y agoThis is why GCP has no hope of ever taking significant market share from AWS. Google thinks they can treat their cloud customers like they treat users of their free services. Customer support and communication are essential.
- halbritt 8y agoI'm not sure about the market share, but I agree with the last two sentences. ...and I'm a happy GCP customer.
- ernsheong 8y agoAs if something like this has never happened to AWS?
- greglindahl 8y ago"like this" -- a failure of the service, or a failure of communication and customer support?
- ernsheong 8y agoBoth, I suppose.
- ty_a 8y agoRemember that time S3 went down and the only updates were on Twitter because the status page was hosted on S3?
- ascar 8y agoI don't. But thanks for sharing, that's hilarious (for an unaffected person at least :D)
- pzh 8y agoThe S3 outage duration was only 4 hours.
- topicseed 8y agoLol, is this real? If so, hilarious.
- paulddraper 8y agoYes. When the AWS status page failed to accurately inform their customers for several hours, AWS used Twitter to ensure that there was communication with their customers. What exactly is your point?
- ransom1538 8y ago“2) This issue was global in scope, affecting all of Google's regions. Therefore, in consideration of item 1 above, it was questionable/unpredictable whether or not a user could launch a node pool or even a simple node anywhere in GCP at all.” Ok. So on aws we were* paying for putting systems across regions, but, honestly I don’t get the point. When an entire region is down what I have noticed is that all things are fucked globally on aws. Feel free to pay double - but it seems* if you are paying that much just pay for an additional cloud provider. Looks like it’s the same deal on GCP.
- human_error 8y ago> When an entire region is down what I have noticed is that all things are fucked globally on aws. Do you have an example on this?
- ransom1538 8y agoJust grabbed first article. Example: In this case capitalone went down. I don’t work at capitalone - but I imagine they had their data copied across every region 30 times. https://www.geekwire.com/2018/widespread-outage-amazon-web-services-u-s-east-region-takes-alexa-atlassian-developer-tools/ https://www.geekwire.com/2018/widespread-outage-amazon-web-s...
- excalq 8y agoOn 17 October, there was a multi-AZ network failure at us-east-1. It only lasted 3m35s, but it was enough that our customers were calling about our site being down.
- lilbobbytables 8y agoYou're doing me a scare. I'm in the evaluation phase with them. Maybe I'm missing something here, but this is not at all what the linked post says. "We are investigating an issue with Google Kubernetes Engine node pool creation through Cloud Console UI." So, it's a UI console issue, it appears you can still manage "Affected customers can use gcloud command [1] in order to create new Node Pools. [1]" Similarly, it actually was resolved in Friday, but they forgot to mark it as so. "The issue with Google Kubernetes Engine Node Pool creation through the Cloud Console UI had been resolved as of Friday, 2018-11-09 14:30 US/Pacific."
- aviv 8y agoI can't comment regarding GKE as we don't use that particular service, however we are very heavy users of many other GCP services, including Compute, Datastore, BigQuery, Pub/Sub, Storage, Functions, Speech, and others. Zero issues this weekend, everything is running 100% as any normal day.
- haldora 8y agoI've been failing all weekend to create nodes in a GKE cluster through either the UI console or gcloud. Even right now I can't get any nodes to spin up. Edit: I finally got my cluster up and running by removing all nodes, letting it process for a few minutes, then adding new nodes.
- timdumol 8y agoWe've had no issues deleting and creating node pools this weekend (on asia-east1-a). No other problems noticed either.
- shareometry 8y agoYou are right about the Google blog content itself not indicating three days of outage. Turns out they just forgot to mark that particular issue as resolved on Friday, as you point out. This is my mistake. I would update my comment to reflect this, but it doesn't seem to allow an edit at this point. The items I put down in my comment are based largely on user reports, though (there isn't much else to go on). And I mean these items as questions (i.e. "is this accurate?"). Folks here on HN have definitely been reporting ongoing problems and seem to be suggesting that they are not resolved and are actually larger in scope than the Google blog post addressed. Someone from Google commented here a few hours ago indicating Google was looking into it. And other folks here are reporting that they don't have the same problems. So it's kind of an open question what's going on. I'm in the evaluation phase too. And I've found a lot to like about GCP. I'm hoping the problems are understandable.
- navinsylvester 8y agoWe are GCP customers for the last couple of years. We use other cloud platforms(AWS, IBM, Oracle, OrionVM) too. We don't use GKE but use rancher/kubernetes combo on their standard platform. So far GCP is the best, hands down in terms of stability. We never had a single outage or maintenance downtime notification till now. We are power users but our monitoring didn't pick any anomaly so i don't think this issue had rampant impact on other services. But i find it concerning that they provided very little update on what went wrong. I also think its better to expect nil support out of any big cloud provider if you don't have paid support. Funny how all these big cloud providers think you are not eligible for support de-facto. Sigh.
- rogerkirkness 8y agoI agree with this. Compared to AWS, when Google says it's down, it's down, and that's rare. When they say it's up, it's up.
- ToFab123 8y agoI don't understand why someone would choose to deploy anything mission critical without having an support contract with the ISP, the manufacturer of the the software etc.
- marcinzm 8y agoSimple, the cost of an outage is less than the cost of a support contract. Very few things are really mission critical as in they can never go down. Rather they simply have a cost to going down and you can choose to pay that one way or another.
- jjeaff 8y agoAnd it's not like having a support contract precludes you from downtime.
- navinsylvester 8y agoI transitioned from collocation to self managed remote server farm and then onto self managed remote vms. All these providers provided de-facto support whether we opted for one or not. You can go to their portal and raise a ticket. I am not saying with vast numbers its feasible but big cloud providers don't even give you the opportunity to raise a ticket if its their fault. There is a price you pay extra when you opt for any one of them but many don't realize. Having said that - almost all the time, our skilled expertise is better than their initial two level of support staff. We realized it early so we handle it better by going over the documentation and making our code resilient since all cloud platforms have some limit or another since overselling in a region is something they can't avoid. Going multiple regions across when you handle these exceptions is the only way through.
- dilyevsky 8y agoWe had an issue a few weeks back where all nodes in west1-a could not pull docker images. Google support was pinballing P1 issue around the globe and across multiple teams for a few days untill I root caused it for them - turned out to be gce service account issues affecting entire zone. 2 days to rollback (no status page update). I know nobody gives a fuck but can’t help but feel vindicated as an ex google sre.
- icelancer 8y agoI think a lot of people give a fuck here; I do, at least. Thanks for outlining it, these things are fascinating (to me anyway, who has never worked in IT/ops).
- tejohnso 8y ago> For three whole days, it was questionable whether or not a user would be able to launch a node pool (according to the official blog statement) What blog statement are you referring to? I don't see any such statement. Can you provide a link? The OP incident status issue says "We are investigating an issue with Google Kubernetes Engine node pool creation through Cloud Console UI". It also says "Affected customers can use gcloud command in order to create new Node Pools." So it sounds like a web interface problem, not a severely limiting, backend systems problem with global scope. Also, the report says "The issue with Google Kubernetes Engine Node Pool creation through the Cloud Console UI had been resolved as of Friday, 2018-11-09 14:30 US/Pacific". So the whole issue lasted about 10 hours, not three whole days. > Some users here are reporting that other GCP services not mentioned by Google's blog are experiencing problems I don't see much of that.
- paulddraper 8y agoI believe the OP was referring to the very same blog (web log) you cited. https://status.cloud.google.com/incident/container-engine/18005 https://status.cloud.google.com/incident/container-engine/18... "We are investigating an issue with Google Kubernetes Engine node pool creation through Cloud Console UI." > So it sounds like a web interface problem, not a severely limiting Depends who you as to whether this is "severely" limiting, but yes there is a workaround by using an alternate interface.
- manigandham 8y agoWhen everything works, GCP is the best. Stable, fast, simple, reliable. When things stop working, GCP is the worst. Slow communications and they require way too much work before escalating issues or attempting to find a solution. They already have the tools and access so most issues should take minutes for them to gather diagnostics, but instead they keep sending tickets back for "more info", inevitably followed by a hand-off to another team in a different time zone. We have spent days trying to convince them there was an issue before, which just seems unacceptable. I can understand support costs but there should be a test (with all vendors) where I can officially certify that I know what I'm talking about and don't need to go through the "prove its actually a problem" phase every time.
- laurencei 8y agoAs someone who works for Government and Enterprise - all I care about sometimes is how a company behaves when everything goes wrong. The issue with outages for the Government organizations I have dealt with is rarely the outage itself - but strong communication about what is occurring and realistic approximate ETAs, or options around mitigation. Being able to tell the Directors/Senior managers that issues have been "escalated" and providing regular updates are critical. If all I could say was a "support ticket" was logged, and we are waiting on a reply (hours later) - I guarantee the conversation after the outage is going to be about moving to another solution provider with strong SLAs.
- donedealomg 8y agoJust move to AWS, AWS support is strong. Google thinks all their customers are idiots and that their AI can handle everything.
- totallyashill 8y agoVery similar thing at our office. Considering the scale of which we run things, any outage could be a potential loss of millions _every minute_. Sure, we use support tickets with vendors for small things. Console button bugging out, etc. But for large incidents, every vendor has a representative within an hour driving distance and will be called into a room with our engineers to fix the problem. This kind of outage, with zero communication, means the dropping of a contract. Communication is critical for trust, especially if we're running a business off it.
- meow_mix 8y agoI think you're missing the portion about how it only appears to be the console ui, no?
- rorykoehler 8y agoI recently removed my hosting from GCP. The pricing is confusing and unbelievable. Their customer service is a joke. I don't trust Google for longterm consistency due to the way they shut their own apps but I let that slide as I doubt they will do that on their cloud services. I have experience with AWS (rock solid, world class support but also costly), digital ocean (improving fast), heroku (good for beginners but also expensive and not as full featured as AWS) and finally Hetzner (too early to judge).