22 ms·
Google Kubernetes Engine's third consecutive day of service disruption
- 013a 8y agoCloud providers have all of the potential in the world to make each region truly isolated. I shouldn't have to architect my application to be multi-cloud, at least for stability reasons. Yet, somehow every major cloud provider experiences global outages. That old AWS S3 outage in us-east-1 was an interesting one; when it went down, many services which rely on S3 also went down, in other regions beside us-east-1 because they were using us-east-1 buckets. I have a feeling this is more common than you'd think; globally-redundant services which rely on some single point of geographical failure for some small part.
- threeseed 8y agoAWS regions are very much isolated from each other. We know because we are still waiting here in ap-southeast-2 for services such as EKS to be made available. Pretty sure that any reliance within their backend services on us-east-1 was just a temporary bug and nothing systemic.
- AlexB138 8y agoThis has been going on longer than three days. We have been dealing with this exact issue since at least Monday (11/5) morning in us-central1.
- splap 8y agosame here. using gcloud, not web console
- fizzledbits 8y agosame here as well, my us-central1 instance still will not boot
- rlancer 8y agoStatus page is inaccurate as issues doesn't only affect the web UI, the same operations are not functioning via the CLI.
- camhutch 8y agoI just tried turning up my 1-node test cluster via terraform, and it worked fine. I would have thought the gcloud CLI would be using the same API. I did this in the australia-southeast1-a zone.
- base698 8y agoWhat operations? Status just shows node pool creation.
- rlancer 8y agoCan not create a new Clusters or Node Pool and can not resize exiting Node Pools, as far as users are reporting it's happening in all regions too. Error message when creating a new Cluster: Deploy error: Not all instances running in IGM after 35m7.509000994s. Expect 1. Current errors: [ZONE_RESOURCE_POOL_EXHAUSTED]: Instance 'gke-cluster-3-pool-1-41b0abf8-73d7' creation failed: The zone 'projects/url-shortner-218503/zones/us-west2-b' does not have enough resources available to fulfill the request. Try a different zone, or try again later. - ; .
- pm90 8y agoIts kinda strange that HN seems to be the most effective way to give feedback to Google Cloud :/
- deleted 8y ago[deleted]
- regnerba 8y agoIs this just about creating new pools? I haven't noticed an issue with our existing pools scaling.
- rlancer 8y agoYou were able to add more Nodes to you're pool? Are you using any auto scaling?
- closeparen 8y agoDoesn’t GKE “just” run an independent Kubernetes cluster on customer VMs? How is a widespread outage like this possible?
- rlancer 8y agoGKE gives you a fully managed Master Node.
- regnerba 8y agoGKE does the creation of the VMs and setup of them, joining them to the cluster and applying labels for example. The specific issue appears to be about creating new "node pools". Creating standard VMs in GCP works fine however, so this is specific to GKE and their internal tooling that integrates with the rest of GCP. GKE doesn't (at least to my knowledge) allow you to create VMs separately and join them to the cluster in any kind of easy fashion.
- kenan_warren 8y agoIt's actually not just GKE, there have been issues creating normal VMs since late Friday night. It seems anything that required creating VMs gave back resource exhaustion errors. I finally got a cluster for us-east1 setup last night so it looks like the resource issues are clearing up though.
- raincom 8y agoNope, GKE = master/control plane owned by Google. Customers are just tenants, who can schedule workloads.
- thwy12321 8y agoBeen trying to spin up vm instances all day, had to try every single zone just to get one up. Not only is this incredibly harmful to a technology business dependent on this infra, it wasnt obvious to me what the issue was until I tried creating instances. Nothing says, hey resources are constrained here, try this one. Just about ready to bite the bullet and move to aws.
- pfd1986 8y agoSame here. We have spent 2 days trying to create instances and migrate images just to figure out later they can't start. Right when I convinced our project to get migrated from AWS...
- tigershark 8y agoWhy on earth would you do that, unless you had huge problems on aws?
- Masiosare 8y agoSame question... why would you do that? AWS is super stable most of the time. I have been running k8s over EC2 (not eks) for a year and works like a charm. I've even run experiments using spot instances and it's pretty good (no guarantee there of course).
- pfd1986 8y agoWe did. Our customer had a bucket shared to a role they created, we spent a week back and forth trying to mount such bucket in our instance using fuse. Mounting with gcsfuse took 5 minutes (although no role used, so perhaps an unfair comparison). In general, I found Gcloud a lot easier to work with.
- nielsole 8y agoI use preemptible machines in autodialing and for first time did not have any machines available for multiple hours yesterday. I am wondering whether this falls under the normal preemptible behaviour or this service degradation.
- Jedi72 8y ago"The data says engagement is down 46%, I think its time we drop the product." - Someone at Google right now, probably.
- justinsb 8y agoI can assure you that's not the case! Also, while people like to repeat this meme, Google Cloud does have a formal deprecation policy (https://cloud.google.com/terms/ https://cloud.google.com/terms/), whose intent is to give you some assurances. (I work at Google, on GKE, though I am not a lawyer and thus don't work on the deprecation policy)
- whydoineedthis 8y agoim pretty sure he just forget the /s (sarcasm) on his post, but this was pretty cool information anyway, so thanks!
- brian-armstrong 8y agoWhat happens when they suddenly deprecate the deprecation policy?
- omeid2 8y agoNothing, sort of. Subject to "Section 1.7 Modifications" of Terms: b. To the Agreement: Google may make changes to this Agreement, including pricing (and any linked documents) from time to time. .... Google will provide at least 90 days’ advance notice for materially adverse changes to any SLAs by either: (i) sending an email to Customer’s primary point of contact; (ii) posting a notice in the Admin Console; or (iii) posting a notice to the applicable SLA webpage. If Customer does not agree to the revised Agreement, please stop using the Services. Google will post any modification to this Agreement to the Terms URL.
- skrebbel 8y agoSo the deprecation policy is "you got 90 days".
- marcinzm 8y ago>Nov 09, 2018 05:59 >We will provide more information by Monday, 2018-11-12 11:00 US/Pacific. Wait, did the people tasked with fixing this just take the weekend off?
- jasonlotito 8y agoThe people tasked with fixing this aren't the ones providing the updates.
- marcinzm 8y agoFair point but still seems odd that the people providing updates took the weekend off during a large scale customer impacting issue. I'm sure all the people spending the weekend trying to mitigate the impact of this on their infrastructure would love to have timely updates.
- izacus 8y agoIt's weekend, why wouldn't you take it off? It's just silly software.
- jschwartzi 8y agoMore to the point, why would you depend on Google for any critical infrastructure after this?
- user5994461 8y agoBecause you tried to run your own infrastructure and it was so much worse.
- Draiken 8y agoThen in a few months from now, people will be saying the same things after an AWS or Azure outage. Am I the only one that finds this slightly humorous?
- 8y ago
- scarface74 8y agoA generic question: Our company is completely dependent on AWS. Sure we have taken all of the standard precautions for redundancy, but what happened here could just as easily happen with AWS - a needed resource is down globally. What would a small business do as a contingency plan?
- geggam 8y agoInfrastructure as code. Terraform using AMIs plus chef recipes that work in the cloud and bare metal. Dont use AWS specific services. This would allow you to spin over to another cloud provider , vsphere or bare metal with minimal work
- xchaotic 8y agoI think you are downplaying minimal
- Twirrim 8y agoI don't think OP was intending "minimal" to mean it would be easy to get to the stage where it's possible, just that once you've got all your infrastructure-as-code stuff set up correctly, you ought to be able to just be pressing buttons / running scripts and have your infrastructure up and running in another cloud provider. Even when working in small companies with small infrastructure, I've kept recreation of infrastructure as one of my high priorities (one reason it really bugged me in one job to have to depend on Oracle Databases that I couldn't automate to the same degree.) In my mind, it's not different from the importance of having, and testing restoration of, backups. If your infrastructure gets compromised somehow, or you find yourself up the creek with your provider, you've got to be able to rebuild everything from scratch.
- user5994461 8y agoThe infrastructure has always been the easy part, as long as the company is willing to pay for multiple datacenters. Then you realize a lot of software and databases can only run from a single instance, zero support for multi regions, and you're not gonna to rewrite everything and resiliency just can't happen.
- sladey 8y agoSeems to be some weird underlying issue going on at GCP at the moment. Had cloud build webhooks returning a 500 error. Noticed we were at 255 images and deleting some fixed the issue. Created a P2 ticket about the issue before we managed to solve it and haven't had a response in 40+ hours. The timeline of this disruption matches when we started experiencing cloud build errors.
- lstamour 8y agoOutsider here, but I believe Cloud Build runs on GKE Jobs, so if they’re having trouble, it does indeed sound related.
- gigatexal 8y agoOh man must be a tough time to be an SRE at google cloud. But... they’re Google. They have been doing internal cloud for years and years. Borg — which K8s is a reimplementation if — has been the heart of Google for so long now you’d think they’d be able to architect their systems to have no outages whatsoever. I mean nobody is perfect but this looks bad.
- Jedi72 8y agoGoes to show outsourcing infrastructure is more about blame shifting so that when things go wrong its "not our fault" than reducing actual downtime.
- justinsb 8y agoHi - I work at Google on GKE - sorry about the problems you're experiencing. There's a lot of people inside Google looking into this right now! It looks like the UI issue was actually fixed, and that we just didn't update the status dashboard correctly. But we're double checking that and looking into some of the additional things you all have reported here.
- marcinzm 8y agoI appreciate all the effort you're putting in and I understand such situations can be stressful but user's having to depend on someone responding on hacker news for status updates seems really amateur for an organization the size of google.
- ben_jones 8y agoAs much as I love bashing big corps I see HN as a supplementary communication channel for products like GCP - its a luxury we get to access alongside normal customer support channels in the GCP console, twitter, etc.
- thwy12321 8y agoCritical service is failing, minimal information about why, but we should be so happy someone says a few sentences on here? For all of the engineering elitism coming out of google, Amazon is way more on their game across a number of products.
- yeukhon 8y agoLet me put it this way. HackerNews, or in fact, any news outlets are not official. Customers should be getting emails from Google and be informed on its official webpage to explain what's going on. You don't want your neighbor to tell you you owe taxes. You want the government to send a notice to you.
- NicoJuicy 8y agoThe default is : https://status.cloud.google.com/incident/container-engine/18005 https://status.cloud.google.com/incident/container-engine/18... People who respond here could be employees of Google, caring about it and respond here because they know it. What he can mention ( a lot of people are working on it) is what you can suspect when something is going down. All other cloud providers do the same.
- hacknat 8y agoQuestion to Google employees: Why do you guys suffer global outages? This is your 2nd major global outage in less than 5 years. I’m sorry to say this, but it is the equivalent of going bankrupt from a trust perspective. I need to see some blog posts about how you guys are rethinking whatever design can lead to this - twice - or you are never getting a cent of money under my control. You have the most feature rich cloud (particularly your networking products), but down time like this is unacceptable.
- toomuchtodo 8y ago> I’m sorry to say this, but it is the equivalent of going bankrupt from a trust perspective. It's the opposite really: the expectation that service providers have no unexpected downtime is unrealistic, and it's strange this idea persists.
- Twirrim 8y ago(disclaimer: I work for another cloud provider) I agree, in general, outages are almost inevitable, but global outages shouldn't occur. It suggests at least a couple of things: 1) Bad software deployments, without proper validation. A message elsewhere in this post on HN suggest that problems have been occurring for at least 5 days, which makes me think this is the most likely situation. If this is the case, presumably given this is multiple days in to the issue, rolling back isn't an option. That doesn't say good things about their testing or deployment stories, and possibly their monitoring of the product? Even if the deployment validation processes failed to catch it, you'd really hope alarming would have caught it. or: 2) Regions aren't isolated from each other. Cross-region dependencies are bad, for all sorts of obvious reasons.
- toomuchtodo 8y agoThat shouldn't, but they do. S3 goes down [1]. The AWS global console goes down, right after Prime Day outages [2]. Lots of Google Cloud services go down [3, current thread]. Tens of Azure services go down hard [4]. Are software development and release processes improving to mitigate these outages? We don't know. You have to trust the marketing. Will regions ever be fully isolated? We don't know. Will AWS IAM and console ever not be global services? We don't know. Blah blah blah "We'll do better in the future". Right. Sure. Some service credits will get handed out and everyone will forget until the next outage. Disclaimer: Not a software engineer, but have worked in ops most of my career. You will have downtime, I assure you. It is unavoidable, even at global scale. You will never abstract and silo everything per region. [1] https://www.theregister.co.uk/2017/03/01/aws_s3_outage/ https://www.theregister.co.uk/2017/03/01/aws_s3_outage/ [2] https://www.cnbc.com/2018/07/16/aws-hits-snag-after-amazon-prime-day-downtime.html https://www.cnbc.com/2018/07/16/aws-hits-snag-after-amazon-p... [3] https://www.cnet.com/news/google-cloud-issues-causes-outages-for-snapchat-spotify-and-others/ https://www.cnet.com/news/google-cloud-issues-causes-outages... [4] https://www.datacenterknowledge.com/uptime/microsoft-blames-severe-weather-azure-cloud-outage https://www.datacenterknowledge.com/uptime/microsoft-blames-...
- franky_g 8y agoHad it affected all regions or just some? Is there another status page Google? Coz the last update I'm looking at...is dated on the 9th..
- justinsb 8y agoThe general page is at https://status.cloud.google.com/; https://status.cloud.google.com/; you can scroll down to see GKE, and my (unofficial) belief is that https://status.cloud.google.com/incident/container-engine/18006 https://status.cloud.google.com/incident/container-engine/18... should have closed out https://status.cloud.google.com/incident/container-engine/18005 https://status.cloud.google.com/incident/container-engine/18... _If_ that's the case, something else is causing the error messages other people are seeing
- _wmd 8y agoWhen guerilla marketing backfires
- shiftnight 8y agoI have a question. At what point does k8s make sense? I have a feeling that a microservice architecture is overkill for 99% of businesses. You can serve a lot of customers on a single node with the hardware available today. Often times, sharding on customers is rather trivial as well. Monolith for the win! Opinions?
- hacknat 8y agoThis outage really doesn’t have much to do with K8s.
- shiftnight 8y agoMaybe so, but you won't be affected by this outage if you never decided to deploy k8s in the first place. Even if you deploy k8s privately, or over at Amazon, I think there's enough horror stories to make you think twice about the technology. Then, if it isn't going to be k8s for microservices, what's a more reliable alternative?
- kazen44 8y agoi agree with you on major points. Also, migrating to microservices for existing services might not be worth it, especially if you don't operate at a massive scale. Keep it simple stupid is still a solid design decision, despite all the microservice/container hype. Most bussinesses only need a couple of servers that provide the service, spread redundantly with a HA capability.
- polemic 8y agoThere's a huge range between monolith and microservice approach, and even a monolith will have dependent services. A simple web stack these days might include nginx, a database, a caching layer, some sort of task broker and then the 'monolith' web app itself. All of that can be sanely managed in k8s.
- james-mcelwain 8y agoRight... IMO monolith is better understood as a reference to the data model than deployment topology. If you only have a single source of truth, then your application is naturally going to trend towards doing most of its business logic in one place. This still doesn't displace the need for other services like caching, async tasks, etc., that you identify.
- 7ewis 8y agoI honestly don't mind if providers have outages - we can't expect 100.00% accuracy, I know the systems I manage certainly don't achieve that. One thing I do care about though, is root cause analysis. I love reading a good RCA, it restores my faith in the company and makes me trust them more. (I'm not affect by the GKE outage so opinions may differ right now!)
- whatshisface 8y agoWhy do cloud providers have more global outages than major flagship websites like google.com?
- usmannk 8y agoWe had an issue a few weeks ago where the google front-end servers were mangling responses from Pub/Sub and returning 502 responses, making the service completely unusable and knocking over a number of things we have running in production. Despite paying for enterprise support and having in a P1 ticket, we had to spend Friday to Sunday gathering evidence to prove to the support staff that there was indeed a problem, because their monitoring wasn't detecting it. Right now I'm doing something similar (and since Friday!) but for TLS issues they're having. Again, because their support reps don't believe there's a problem. There are so many more problems than they ever show on their status page...
- Jedi72 8y agoThey work for Google so obviously they are much smarter than you. If theres a problem its probably the customers fault. /sarcasm
- deleted 8y ago[deleted]
- mbrumlow 8y agoI was so mad to read that until you said /sarcasm :p That being said I really do think there is a difference between who is working at google today and the google we all fell in love with pre-2008. I am sure there are a amazing people still working at google, but nowhere near like it was. The way I like to think about google is that some amazing people mad ea awesome train that builds tracks in front of it -- you can call them gods maybe -- but those people are gone -- or a least the critical mass required to build such a train has dwindled to just dust. What we have left is a awesome train full of people pulling the many levers left behind. To make things even worse my last interview as a SRE left me wondering if even the people who are there know this as well, and they are actually working hard to keep out those who might expose light on to this. I don't say that because I did not get the job -- I am actually happy I did not get extended a offer. I say this with one exception, the old-timer who was my last interview. I could tell he was dripping in knowledge and eager to share it with any that would listen. I came out of his 45 min session learning many things -- I wold actually pay to work with a guy like that. I would also like to point out that the work ethic was not what I expected. I was told that when on call, my duty was to figure out the root cause was in the segment I was responsible for. I don't know about you, but if my phone rings at night I am going to see through to a resolution and understand the problem in full -- even if it is not on the segment that I was assigned. /end rant
- qaq 8y agoThere is no magic public clouds have incredibly complex control planes and marketing fluff aside you would very likely experience much better uptime at singe top tier DC than @ a cloud provider.
- spiderPig 8y agoOur company is dependent on this as well and the way customer service has been handling this has been abysmal thus far.
- shareometry 8y agoI am currently evaluating GCP for two separate projects. I want to see if I understand this correctly: 1) For three whole days, it was questionable whether or not a user would be able to launch a node pool (according to the official blog statement). It was also questionable whether a user would be able to launch a simple compute instance (according to statements here on HN). 2) This issue was global in scope, affecting all of Google's regions. Therefore, in consideration of item 1 above, it was questionable/unpredictable whether or not a user could launch a node pool or even a simple node anywhere in GCP at all. 3) The sum total of information about this incident can be found as a few one or two sentence blurbs on Google's blog. No explanation nor outline of scope for affected regions and services has been provided. 4) Some users here are reporting that other GCP services not mentioned by Google's blog are experiencing problems. 5) Some users here are reporting that they have received no response from GCP support, even over a time span of 40+ hours since the support request was submitted. 6) Google says they'll provide some information when the next business day rolls around, roughly 4 days after the start of the problem. I really do want to make sure I'm understanding this situation. Please do correct me if I got something wrong in this summary.
- marcinzm 8y agoRight now we don't know. It's one of two possibilities from what I can tell: a) Google had a global service disruption that impacted Kubernetes node pool creation and possible other services since Friday. They had a largely separate issue for a web UI disruption (what this thread links to) which they forgot to close on Friday. They still have not provided any issue tracker for the service distribution and it's possibly they only learned about it from this hacker news thread. b) People are having various unrelated issues with services that they're mis-attributing to a global service disruption.
- johnpython 8y agoThis is why GCP has no hope of ever taking significant market share from AWS. Google thinks they can treat their cloud customers like they treat users of their free services. Customer support and communication are essential.
- spullara 8y agoIf a hosting service is down and nobody uses it, is there really any disruption?
- rlancer 8y agoUPDATE: Got some clarity, these issues are caused by "resource exhaustion" meaning there are no resources left to be allocated.
- halbritt 8y agoI'm curious to see if this is true. I faced some pretty serious resource allocation issues earlier in the year. The us-west1-a region was oversubscribed. I was unable to get any real information from support with regard to capacity. Eventually my rep gave me some qualitative information that I was able to act on.
- locusm 8y agoDo not use GCP without paying for support. We have had resource allocation errors for weeks, as have a lot of other people. Check out the posts in their forum where folk on basic support get zero love. https://groups.google.com/forum/?utm_medium=email&utm_source=footer#!forum/gce-discussion https://groups.google.com/forum/?utm_medium=email&utm_source...
- scarface74 8y agoSay I were a CTO (I’m nowhere near it), why would I choose GCP over AWS or Azure? Even if after doing a technical assessment and I thought that GCP was technically slightly better, if something happened, the first question I would be asked is “why did you choose GCP over AWS?” No one would ever ask why you chose AWS. The old “no one ever got fired for buying IBM”. Even if you chose Azure because you’re a Microsoft shop, no one would question your choice of MS. Besides, MS is known for thier enterprise support. From a developer/architect standpoint, I’ve been focused the last year on learning everything I could about AWS and chose a company that fully embraced it. AWS experience is much more marketable than GCP. It’s more popular than Azure too, but there are plenty of MS shops around that are using Azure.
- halbritt 8y agoGCP has a few features that set it apart from other cloud providers. GKE is head and shoulders above the other offerings from AWS and Azure. GCP can be a fair bit cheaper than AWS and Azure for certain workloads. Raw compute/memory is about the same. Storage can make a big difference. GCP persistent SSD costs a bit more than AWS GP2 with much better performance and way cheaper than IO2. Local SSD is also way, way cheaper than I2 instances. Most folks deploying distributed data stores that need guaranteed performance are using local disk, so this can be a really big deal.
- scarface74 8y agoI have a more detailed post above, but if you are large enough, you’re not paying the listed price for AWS. But even if you are, prices change all of the time. From a completely selfish standpoint, is the price difference worth the cost to bet your reputation on if you are the one that made the final decision? Even if statistically the same could happen with AWS, no one would blame you for choosing AWS. However, I could see doing a multicloud solution where I took advantage of the price difference for one project.
- user5994461 8y agoTo be fair, prices are stable across the providers and Google cloud is very competitive. It's not like 5 years ago when everyone was ramping up their offerings with a yearly price drop and a new generation.
- openloop 8y agoGood luck stopping Kubernetes
- fulafel 8y agoOfftopic but are there some documented exceptions to the "keep the original title" rule?
- fergie 8y agoThings break after everybody has gone home on a Friday? 3 day disruption.
- bdibs 8y agoAs someone currently trying to decide between GCP and AWS for a project, is this a regular occurrence? And for those who have used both, which would you go with today?
- arunoda 8y agoThe is not only GKE. But for GCE as well. I cannot create instance is almost all zones. I tried both preemptible and normal as well. Always saying resource not available. My account is a pretty new account. In contrast, one of my friend is having a pretty old account which is very active. He has no such issue. So I think due to this issue, Google has enabled some resource limitation for new accounts. But they should properly communicate this issue.
- ernsheong 8y ago"third consecutive day of service disruption" is not an accurate statement? Latest update was Nov 11 saying things resolved on Nov 9. https://status.cloud.google.com/incident/container-engine/18005 https://status.cloud.google.com/incident/container-engine/18...
- ernsheong 8y agoIf all nodes in GKE clusters were down for 3 days, I would consider this newsworthy and shocking. This... is not. Come on, people.
- haosdent 8y agoTime to use Mesos.
- aaaaaaaaaab 8y agoDaily reminder that there's no "cloud", just other people's computers. ( ͡° ͜ʖ ͡°)
- thomasfl 8y agoI'd like to upvote, but 666 points seemed relevant.
- wb3tech 8y agoIf anyone is interested, here is my documented experience with this issue. I freaking love GCP and GKE, although I have not production environment as it was a HA cluster in us-central1. Working federation now. https://stackoverflow.com/questions/53244471/gke-cluster-wont-provision-in-any-region https://stackoverflow.com/questions/53244471/gke-cluster-won...
- deleted 8y ago[deleted]
- wijowa 8y agoRight now we're experiencing an issue where a small percentage of end users on our GKE site are getting super slow speeds. The issue is ISP related as they can switch to a 4G hot spot in the same location and get normal speeds... and inside our system the timing looks normal. So there's a slowdown either TO the load balancer or FROM the load balancer. Took a week to convince Google's support contractor to even believe it wasn't an issue with our site and their advice is generally along the lines of Turn it off and Turn it back on again (which might actually fix the problem) though that's easier said than done in GCP.
- fizzledbits 8y agoAs of this morning, I am still unable to reliably start my docker+machine autoscaling instances. In all cases the error is "Error: The zone <my project> does not have enough resources available to fulfill the request" An instance in us-central1-a has refused to start since last Thursday or Friday. I created a new instance in us-west2-c, which worked briefly but began to fail midday Friday, and kept failing through the weekend. On Saturday I created yet another clone in northamerica-northeast1-b. That worked Saturday and Sunday, but this morning, it is failing to start. Fortunately my us-west2-c instance has begun to work again, but I'm having doubts about continuing to use GCE as we scale up. And yet, the status page says all services are available. Is the typical of others' experiences?