20 ms·
Kubernetes Failure Stories
- hjacobs 8y agoChristian already followed the example and created a similar list for Serverless: https://github.com/cristim/serverless-failure-stories https://github.com/cristim/serverless-failure-stories
- SlowRobotAhead 8y agoThe third example on that list is a nice little short story. Simple mistake, gets right to the point. I’m setting up a lambda test right now so I find it perfectly timed!
- gspetr 8y agoIs there also a list for Docker failure stories?
- hjacobs 8y agoIMHO this would be less interesting, some people already run other container runtimes such as containerd with Kubernetes (e.g. Datadog: https://www.youtube.com/watch?v=2dsCwp_j0yQ https://www.youtube.com/watch?v=2dsCwp_j0yQ) --- so Docker might stay as some user interface for local development, but I would not know what "Docker failures" would be in the future.
- pepemon 8y agoDocker is using containerd under the hood as its container runtime component.
- tnolet 8y agoI'd be interested in a related "microservices failure stories". Must be a big overlap with this.
- coredog64 8y agoMicroservices failure stories? “All of them. The End”
- tnolet 8y agoI'm not convinced all microservices ventures are failures. Having worked in the space a bit as a founder/CTO of a vendor in that space I've just seen many examples of misguided attempts due to fashion / CV-driven development / hype driven development. The pattern is valid, just not for everyone at every time.
- hjacobs 8y agoThoughtWorks even has a term for it: "Microservice envy" https://www.thoughtworks.com/radar/techniques/microservice-envy https://www.thoughtworks.com/radar/techniques/microservice-e...
- ex_amazon_sde 8y agoThe good part of "microservices" is "services". The bad part is "micro".
- tyldum 8y agoWhen the definition turns to some kind of dogma, then it will fail.
- lucidone 8y agoI'm consulting on a micro services back end right now with mostly prior experience with monoliths. What is the selling point that drives companies down this direction? It's insane, and my client keeps trying to hire new developers and bring on more consultants to build this thing, but the amount of knowledge required is more than any one person can handle. I have similar issues with their choice of db (nosql) and its inflexibility.
- 8y ago
- awinter-py 8y agoBeyond strictly runtime failures, 2018 feels like the year that most of my friends tried kube but not everybody stayed on. The adoption failures are mostly networking issues specific to their cloud. Performance and box limits vary widely depending on cloud vendor and I still don't quite understand the performance penalty of the different overlay networks / adapters.
- lykr0n 8y agoNetwork is a high performance system, and each layer you add adds latency. Consider a traditional monolithic application. In comes your HTTP request in one end, a bunch of cross thread communication happens, and database queries come out the other end. With that, you have 2 points of network communication. Now with a micro-service, you might have 4 or 5 applications that are needed to replace the above monolith. Throw in a service mesh on top of your cloud providers SDN, you've turned 2 points of network communication into 20 or more. The 5 micro-services talking to each other and the service meshes talking to each other. Add on top the additional processing overhead of maybe 1 to 2ms, you've just added at best 10ms round trip time to get to your databases and some more CPU. And to what benefit? TLS? You can do this in your application, or trust your private network is private. Tracing? You can do this with PID matching and watching the kernel's networking stack.
- devereaux 8y agoSo true. For some low latency applications, anything above the bare minimal virtualization is not acceptable. For what I do, in theory, many things should not impact results. In practice, anything that upon measurement impact results is stripped away. Think A/B testing but for every single component - including the major version of say the python interpreter. That's how you end up running many things baremetal. I'll say the future is not serverless but cloudless
- hjacobs 8y agoI would argue that the longtail of applications does not really care about the impact of overlay networks. For us, the biggest impact on low-latency applications (where 1ms makes a difference) on Kubernetes was disabling CPU throttling in all clusters (you can also remove container limits). Background: a Kernel CFS quota bug leads to throttling even if quota is not yet reached, see https://www.youtube.com/watch?v=eBChCFD9hfs&feature=youtu.be&t=810 https://www.youtube.com/watch?v=eBChCFD9hfs&feature=youtu.be...
- williesleg 8y agoDocker is pretty stable, Kubernetes is a fail pail put out by Google. So we can't catch up.
- peterwwillis 8y agoDang. I wish I had my SRE Wiki up and running already, or I'd add a "public postmortems" section.
- hjacobs 8y agoLooking forward to your public postmortems.. (either yours or whatever you find in the wild)
- alien_ 8y agoJust put it on Github like this and the Serverless one I also created after I saw this.
- alien_ 8y agoJust saw it already exists: https://github.com/danluu/post-mortems https://github.com/danluu/post-mortems
- dvnguyen 8y agoHaving used Docker Compose/Swarm for last two years, I remember having problems with them twice. One of which was an MTU setting which I didn't really understand why, but overall I was relatively happy with them. Since Kubernetes seems to have won, I decided to learn it but got some disappointments. The first disappointment is setting up a local development environment. I failed to get minikube running on a Macbook Air 2013 and a Ubuntu Thinkpad. Both have VTx enabled and Docker and VirtualBox running flawlessly. Their online interactive tutorial was good though, enough for the learning purpose. Production setup is a bigger disappointment. The only easy and reliable ways to have a production grade Kubernetes cluster are to lock yourself into either a big player cloud provider, or an enterprise OS (Redhat/Ubuntu), or introduce a new layer on top of Kubernetes [1]. Locking myself into enterprise Ubuntu/Redhad is expensive, and I'm not comfortable with adding a new, moving, unreliable layer on top of Kubernestes which is built on top of Docker. One thing I like about the Docker movement is that they commoditize infrastructure and reduce lock-ins. I can design my infrastructure so it can utilize an open source based cloud product first and easily move to others or self-host if needed. With Kubernetes, things are going the other way. Even if I never moved out of the big 3 (AWS/Azure/GCloud), the migration process could be painful since their Kubernetes may introduce further lock-ins for logging, monitoring, and so on. [1]: https://kubernetes.io/docs/setup/pick-right-solution/ https://kubernetes.io/docs/setup/pick-right-solution/
- hjacobs 8y agoI never used Docker Swarm (so can't compare), but I don't fully understand your point about Kubernetes cloud lock-in. Certainly there are important differences in networking, load balancing, persistent volumes, and other cloud features, but that's not something any platform can just hide/eliminate (e.g. think about AWS ELB/ALB/NLB vs Google Load Balancer). The Kubernetes concepts (Deployment, Ingress, Service) still work mostly the same for the user across clouds. Some other details like non-standardized Ingress annotations are obviously due to not having them agreed in Kubernetes core API (nginx ingress supports other annotations than say Google LB or Skipper).
- bonesss 8y agoMost of the kubernetes toolchain provides nice support for delineating the requirements from separate cloud providers, too. Compared to most alternatives, a little HELM magic to support hybrid cloud installations is a piece of cake.
- sabareesh 8y agoFor our team size kubernetes was overwhelming so we stayed with docker swarm but we were afraid docker might drop swarm in favor of kubernetes but that didnt happen yet.
- bdcravens 8y agoI've started the planning phase of a Kubernetes course, geared toward developers more so than the enterprise gatekeepers. As I read stories like these, I jump between different thoughts and feelings: 1) no matter what I think I know, there's too many dark corners to create an adequate course 2) K8S is such a dumpster fire that I shouldn't encourage others 3) there's a hell of an opportunity here Thoughts? Worth pursuing? Anything in particular that should be included that usually isn't in this kind of training?
- meddlepal 8y agoThe answer is always 3.
- swampthinker 8y agoIf it was easy, there would be dozens of courses out there already! Sounds like you've found a pain point you can solve.
- zzzcpan 8y agoI think it's ok not to know something to make a course for it. Even if only to learn it better. But I'd be careful given 2, pursuing this could lead to burnout. There is an opportunity for anything infrastructure related.
- parasubvert 8y agoAll three. It’s a gold rush, but as with any gold rush, conditions are hard going - that’s why there’s an opportunity. Best way to think of Kubernetes is that it was designed to be a successful open source project that was widely adopted as a standard foundation to build products. It wasn’t designed to be a useable product on its own. We are at the equivalent stage of Slackware and SLS and Debian Red Hat pre-1.0 stages of GNU/Linux distros circa 1994. Red Hat eventually ran away most of the money by the late 90s, but in the meantime, lots of opportunity to fill an unmet need.
- romeisendcoming 8y agoDon't forget SuSE the sole surviving competitor. Best Buy SuSE Linux gecko box 2.2.14 kernel veteran.
- cygned 8y agoI am a developer and I find k8s frustrating. To me, its documentation is confusing and scattered among too many places (best example: overlay networks). I have read multiple books and gazillions of articles and yet I have the feeling that I am lacking the bigger picture. I was able to set it up successfully a couple of times, with more or less time required. Last time, I gave up after four days because I realized that what I need was a "I just want to run a simple cluster" solution and while k8s might provide that, its flexibility makes it hard for me to use it.
- FridgeSeal 8y agoHave you used other google products? I find their documentation routinely incomprehensible and difficult.
- innocentoldguy 8y agoAgreed! I am an engineer and have written documentation off and on throughout my career. I'm continuously dismayed at the incomprehensible documentation generated by most companies. Google's documentation is particularly bad though.
- abledon 8y agoI have a theory that the type of people who make it past the google interview are smart people who are bad at teaching. Like they get all the concepts, algos etc.. but when it comes to distilling it into an Explain-Like-Im-5 tutorial, it just goes to hell very quickly. What they need to do is hire some people who are great teachers, explainers etc.. Avoid people who rely on already attained technical knowledge, design patterns, algos etc.. to pattern match on new tech to instantly grok it. The 'noob' people who question the engineers who designed the tools and ask a ton of dumb questions about how it works so they can then translate it into everyday tutorial paragraphs.
- hjacobs 8y agoKelsey has the answer for you: "Kubernetes is a platform for building platforms. It's a better place to start; not the endgame." (https://twitter.com/kelseyhightower/status/935252923721793536 https://twitter.com/kelseyhightower/status/93525292372179353...) As an application developer, you probably also don't work with the Kernel and syscalls directly (anymore), so I guess you can expect higher abstractions and a smoother experience for Kubernetes in the future.
- stunt 8y agoKubernetes solves a problem that most of the companies don't have. That is why I don't understand why the hype around it is so big. For the majority, it just adds a little value when you compare to added complexity to infrastructure and the cost of a learning curve and the ongoing operation and maintenance.
- jordanbeiber 8y agoIn my experience most companies lack common conventions and automations. Kubernetes "done right" is almost a part of your application. It becomes this "machine" that you throw stuff into and good stuff happens. You'll need a team to integrate it into the pieces you require (auth, secrets, loadbalancers, permissions/app identities, monitoring and logging) but many places lack bits and pieces, and in my opinion k8s gives you a fast track to create a uniform application delivery platform. What I don't like is that it kind of is the opposite of "the unix philosophy" and in that regard I prefer the hashicorp stack.
- romeisendcoming 8y agoThose are your startups and web/app tier shops. Yes, they suck at sysadmin routinely and they need to be bottle fed a solution that fits the scatter/gather shape of their business. They don't want OPs discipline. They want a programmable solution that performs systems magic with a single toolset to learn.
- jordanbeiber 8y agoNo no, these are your enterprises I’m talking about mainly. Places entrenched in manual processes for release and change management. With true service delivery in a CI/CD fashion (including infrastructure, as it should be codified) many of these manual processes becomes obsolete. Don’t get me wrong, the processes still exist, they are just sped up by a magnitude and automated.
- romeisendcoming 8y ago
- m0zg 8y agoIt's not for everyone and it has significant maintenance overhead if you want to keep it up to date _and_ can't re-create the cluster with a new version every time. This is something most people at Google are completely insulated from in the case of Borg, because SRE's make infrastructure "just work". I wish there was something drastically simpler. I don't need three dozen persistent volume providers, or the ability to e.g. replace my network plugin or DNS provider, or load balancer. I want a sane set of defaults built-in. I want easy access to persistent data (currently a bit of a nightmare to set up in your own cluster). I want a configuration setup that can take command line params without futzing with templating and the like. As horrible and inconsistent as Borg's BCL is, it's, IMO, an improvement over what K8S uses. Most importantly: I want a lot fewer moving parts than it currently has. Being "extensible" is a noble goal, but at some point cognitive overhead begins to dominate. Learn to say "no" to good ideas. Unfortunately there's a lot of K8S configs and specific software already written, so people are unlikely to switch to something more manageable. Fortunately if complexity continues to proliferate, it may collapse under its own weight, leaving no option but to move somewhere else.
- coleifer 8y agoI'm not sure but would something like docker swarm qualify?
- trhway 8y agoAn open source having complexity of an enterprise monster - this is what generates the half million plus salaries. An old enterprise software trick. Simplification of it would serve no interests of anybody in the position to do the simplification.
- m0zg 8y agoI'm going to voice a contrarian viewpoint here and say that "half million plus" salaries are actually a good thing. Rising tide lifts all boats and since a lot of technies live in areas with exorbitant cost of living, that money re-enters the economy at a rapid clip anyway. But such salaries are only good if commensurate value is being delivered for the money. Which in a large IT shop it might be, but as a small business owner K8S is a hard slog, hence my suggestion to simplify. I'm pretty sure 80/20 breakdown still applies, and 80% of K8S complexity could be removed without affecting anything much. One might suggest that I use GKE and bypass the problem entirely, but I need to run a lot of GPUs 24x7, and the pricing on those in any cloud is insane.
- dcomp 8y agoI run a single node cluster at home. In order to handle updates. I just wipe the cluster with kubeadm reset. Then kubeadm init; followed by running a simple bash script. which loops of files in nested subdirectories applying yaml configs. Only have to make sure I only ever edit the yaml files and not mess with kubectl edit etc. for f in /.yaml ... with a directory structure of: drwxrwsrwx+ 1 root 1002 176 Jan 20 21:15 . drwxrwsrwx+ 1 root 1002 194 Nov 17 20:06 .. drwxrwsrwx+ 1 root 1002 68 Jan 20 20:50 0-pod-network drwxrwsrwx+ 1 root 1002 104 Nov 1 11:18 1-cert-manager drwxrwsrwx+ 1 root 1002 34 Jul 11 2018 2-ingress -rwxrwxrwx+ 1 root 1002 93 Jan 20 21:15 apply-config.sh drwxrwsrwx+ 1 root 1002 22 Jul 14 2018 cockpit drwxrwsrwx+ 1 root 1002 36 Jul 3 2018 samba drwxrwsrwx+ 1 root 1002 76 Jul 6 2018 staticfiles
- stonewhite 8y agoI managed multiple mesos+marathon clusters on production a little over 1.5 years, and when I switched over to the K8s the only thing that felt like an improvement was the kubectl cli. I really liked/missed the beauty of simplicity in marathon that everything was a task, the load balancer, autoscaler, app servers everything. I think it failed because provisioning was not easy, lack of first-class integrations with cloud vendors and horrible horrible documentation. Kind of sad to see it lost the hype battle, and since then even Mesosphere had to come up with a K8s offering.
- nisa 8y agoThe k8s hype feels like the Hadoop hype from a few years ago. Both solve problems that most don't have and there is a lot of complexity - some due to the nature of the problem, some because everything is new and moving. Of course it's 2019 and you have to migrate Hadoop to run on k8s now :) My impression is that if you are a small shop and have the money, use k8s on google and be happy, but don't attempt to set it up for yourself. If you only have a few dedicated boxes somewhere just use Docker Swarm and something like Portainer.
- lugg 8y agoDocker swarm is really nice. I wish it had more traction. I fear it's going to be dropped and leave me holding a bag full of bugs.
- BretFisher 8y agoSwarm isn’t going anywhere. It has a growing community and the team is activly working in the repos. See my updates: https://www.bretfisher.com/the-future-of-docker-swarm/ https://www.bretfisher.com/the-future-of-docker-swarm/
- manigandham 8y agoI don't understand all the negative comments here, K8S solves many problems regardless of scale. You get a single platform that can run namespaced applications using simple declarative files with consolidated logging, monitoring, load-balancing, and failover built-in. What company would not want this?
- mirceal 8y agothis is too broad. i think that may actually be the problem: in theory it can do a lot of things, but in the real world it’s hard to get all those theoretical benefits. for me, if you’re in the cloud you don’t need k8s. your favorite cloud provider has already figured out logging and monitoring and the basic things you need to get going. (another story if you run on bare metal) if you’re not running a legacy app you don’t really need containers either. containers are great for legacy apps, for poorly written software or if you like overengineering. the abstraction you need is called a vm. use it. (again if you are in the cloud). your app/service/thing is not as complicated as you think it is (or at least it should not be). I see a lot of people feeling like they need to experiment with new technology, on the job, on whatever they are doing now. actually building something that works and is simple as fuck seems to take a backseat and these types of people will create a narrative around using the new flashy thing. this is how you end up with production systems leveraging tools in beta and you end up closing shop when you finally figure out that you don’t have the resources to understand and maintain what you’ve created. there is a time and place to experiment and learn. on small projects or on your own time. it takes experience to understand the hype cycle and to distinguish good tech from the hype. as for k8s? yes, it solves some problems but it also creates others. do you like basically spending the time you’ve saved on setup and deployment to maintain/troubleshoot/upgrade your cluster? knock yourself out.
- manigandham 8y agoThere is a very big gap between IaaS and PaaS. K8S is an abstraction on top of VMs so you can have a customizable PaaS that runs on YAML code. It has nothing to do with how complex your app is because K8S is about running it with less work in a declarative fashion. I'm currently in and have worked with dozens of startups that have saved lots of time by removing all the ops overhead with K8S because it runs the servers and we can just deploy our apps. It seems like most of the problems are actually about installing and running K8S software itself, but then 95% of companies won't be doing that and using the managed offerings instead. This is no different than companies using the cloud over running their own DCs.
- AaronFriel 8y agoI just went through all of the post-mortems for my own company's purposes of evaluating Kubernetes. I've been running Kubernetes clusters for about a year and a half and have run into a few of these, but here's what I found striking: * About half of the post-mortems involve issues with AWS load balancers (mostly ELB, one with ALB) * Two of the post-mortems involve running control plane components dependent on consensus on Amazon's `t2` series nodes This was pretty surprising to me because I've never run Kubernetes on AWS. I've run it on Azure using acs-engine and more recently AKS since its release, and on Google Cloud Platform using GKE; and it's a good reminder not to to run critical code on T series instances because AWS can and will throttle or pause these instances.
- hjacobs 8y agoNice observation, I haven't done statistics on the linked postportems myself yet. Please note that your observation might also be due to the fact that AWS has a far larger market share and did not provide managed Kubernetes until recently (so people roll their own). We can therefore assume that any random sample of Kubernetes postmortems would be biased towards seeing more incidents with Kubernetes on AWS (compared to other cloud providers).
- AaronFriel 8y agoThat's a good point. In 2017 there weren't widely available managed Kubernetes deployments, and now each platform has their own and much more reliable integrations.
- hjacobs 8y agoThere is now a Kubernetes podcast episode with me about the topic: https://kubernetespodcast.com/episode/038-kubernetes-failure-stories/ https://kubernetespodcast.com/episode/038-kubernetes-failure...