29 ms·
Treat Kubernetes clusters as cattle, not pets
- jeffbee 5y agoRunning a separate cluster for every service assures high overhead and poor utilization. Fine if you can afford it, but be aware that you are paying it.
- q3k 5y agoYeah, especially in production bare metal clusters. If you want N+2 redundancy that's at least 5 physical machines for just the control plane (etcd & apiservers), more if you don't want to colocate worker nodes with that... Even if you have full bare metal automation and thousands of machines that seems like unnecessary waste.
- dharmab 5y agoAdd the cost of administrative services: DNS resolvers, ingress controllers, log forwarders, monitoring (e.g. Prometheus, some exporters), autoscalers, tracing infrastructure, admission controllers, backups/disaster recovery tools (e.g. Velero)... It can add up to millions per year if you aren't auditing your costs.
- dijit 5y agoYeah, that's insane as a concept. One of the larger selling points of Kubernetes was bin packing. Removing that selling point leaves you with... * Orchestration of jobs (restart, start); this can be achieved easily without the complexity of k8s * Sidecar loading; Literally the easiest thing to do with normal VMs. and...?
- dolni 5y agoPacking everything together is a "selling point" until you find that a service can fill up ephemeral storage and take down other services, or consume bandwidth without limit. Let's not forget the potential security implications of not keeping things properly isolated. People who were around when provisioning on bare-metal was still a thing already learned all these lessons. Somehow it seems they have been forgotten by all the people driving hype around Kubernetes.
- dijit 5y agoKubernetes has the concept of limits, especially on ephemeral storage; additionally: if your node becomes unhealthy then the workloads would be rescheduled on another node. I’m super not hypey about kubernetes, mostly because the complexity surrounding networking is opaque and built on a foundation of sand... But let’s not argue things that aren’t true.
- dolni 5y ago> Kubernetes has the concept of limits, especially on ephemeral storage So... why are these issues open? https://github.com/kubernetes/enhancements/issues/1029 https://github.com/kubernetes/enhancements/issues/1029 https://github.com/kubernetes/enhancements/issues/361 https://github.com/kubernetes/enhancements/issues/361 https://github.com/kubernetes/kubernetes/issues/54384 https://github.com/kubernetes/kubernetes/issues/54384 > additionally: if your node becomes unhealthy then the workloads would be rescheduled on another node. Well of course, but you're going to run into that issue (likely) on all of the nodes where the offending service lives. > But let’s not argue things that aren’t true. If what I've said is untrue, looking at open GitHub issues and the Kubernetes documentation is certainly no indication. That's a massive problem all by itself.
- q3k 5y agoThe first issue you've linked concerns quota support for ephemeral storage requests/limits - which is not about the limits themselves, but the ability to set a limit quotas per tenant/namespace. Eg., team A cannot use a total of more than 100G ephemeral storage in total in the cluster. EDIT: No, sorry, it's about using underlying filesystem quotas for limiting ephemeral storage, vs. the current implementation, see the third point below. Also see KEP: https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/1029-ephemeral-storage-quotas https://github.com/kubernetes/enhancements/tree/master/keps/... The second is a tracking issue for a KEP that has been implemented but is still in alpha/beta. This will be closed when all the related features are stable. There's also some discussion about related functionality that might be added as part of this KEP/design. The third issue is about integrating Docker storage quotas with Kubernetes ephemeral quotas - ie., translating ephemeral storage limits into disk quotas (which would result in -ENOSPC to workloads), vs. the standard kubelet implementation which just kills/evicts workloads that run past their limit. I agree these are difficult to understand if you're not familiar with the k8s development/design process. I also had to spend a few minutes on each one of them to understand what the actual state of the issues is. However, they're in a development issue tracker, and the end-user k8s documentation clearly states that Ephemeral Storage requests/limits works, how it works, and what its limitations are: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#local-ephemeral-storage https://kubernetes.io/docs/concepts/configuration/manage-res...
- Spooky23 5y agoRemember in big companies the internal politics rule the day. It’s cheaper to buy more computers than to become the overlord of computing.
- ffo 5y agoYou don’t exactly need to run a cluster per service ;-) Instead you can choose to collocate services who belong together and form a „domain“. But don‘t go the route and build the almighty one Kubernetes cluster where all your domains run in one single place.
- eliodorro 5y agoDecreases utilization but also decreases coordination between teams (no man-bear-pigs). Also weight the long-term costs of poorly maintained platflorms and infrastructure in desaster cases, security issues or when migrating to other providers. High overhead can be automated away, google ORBOS.
- tw600040 5y agoThis idea that cattles are meant to be slaughtered and can't be pets is extremely offensive.
- sweetheart 5y agoPreach. It’s a metaphor that enforces harmful beliefs, like master/slave terminology. Our language defines our world, so we should aim to shy away from language that amplifies a harmful message, even in radically different contexts.
- zachrose 5y agoI’d like to see a vegan alternative to the pets/cattle meme.
- sweetheart 5y agoScreen printing, not painting.
- Zababa 5y agoFungible vs non-fungible seem to be the concept hiding behind pets/cattle. You can substitute a dollar for another, you can't substitute your favorite rock for another.
- mfer 5y agoIf I've read the Google papers on borg right (Kubernetes is conceptually borg v3 with omega being v2) this is different from how Google runs the things. They'll do warehouse scale computing with borg operating large clusters. borg is at the bottom. The workloads spanning dev, test, and prod then run on these clusters. By having large clusters with lots of things running on them they get high utilization of the hardware and need less hardware. It's amusing to see k8s used in such a different way and one that often uses a lot more hardware while driving up costs. Concepts Google used to lower the cost. Or, maybe I read the papers and book wrong. I like the idea of higher utilization and better efficiency because it uses less resources which is more green.
- q3k 5y ago> Or, maybe I read the papers and book wrong. No, that's exactly how it works. You have clusters spanning a datacenter failure domain (~= an AZ), and everything from prod to dev workloads runs on there, with low priority batch jobs bringing up the resource utilization to a sensible level. You can do the same thing with k8s, you just have to trust its multitenancy support. You have RBAC, priority, quotas, preemption, pod security policies, network policies... Use them. You can even force some workloads to use gVisor or separate prod and dev workloads on different worker machines.
- fulafel 5y agoHow do they do version upgrades, isn't that the traditional Achilles heel of K8s that leads people to want to frequently recreate clusters from scratch and/or do blue/green?
- mfer 5y agoFirst, if I understand it right... Google does some smart things in upgrades. They do things like tests and then upgrade their equivalent to AZs in a data center. I'm sure there have been upgrades gone bad that they've had to fix. Kubernetes can be upgraded. I've watched nicely done upgrades happening 4 or 5 years ago. I've watched simple upgrades happen more recently. It's not unheard of. Even in public clouds I've upgrade Kubernetes through many minor versions without issue. I would argue it's more work to create more clusters. You need to migrate workloads and anything pointing to them. It also would cost more as you have to run more hardware.
- AtNightWeCode 5y agoResource utilization is the main reason I would run a cluster in the first place. Immutable infrastructure is also expensive to build and maintain.
- tnisonoff 5y agoFor larger companies, I think a huge benefit of Kubernetes is the shared language for defining and operating services, as well as the well-thought-out abstractions for how these services interact. Costs are generally less of a concern, but having one way of running, operating, and writing services allowed our dev team to move faster, share knowledge, etc.
- ffo 5y agoTrue, cost is not the biggest issue. Separation of teams with different velocity and needs on the other hand is. One API as abstraction with shared processes eases the pain for the people relying on a platform.
- jzelinskie 5y agoI like the idea of "Building on Quicksand" as the analogy for Distributed Systems, but also maintaining your software dependencies. This article basically recommends trying to minimize your dependencies to keep reproducibility/portability high. I generally agree with this, but also carry an "all things within reason" mentality. But just as the article describes coworkers growing into their cluster, the complexity of what they run in their cluster will also grow over time and eventually they'll realize they've just built up their own "distribution". A few years ago, I've written a post asking people to think critically when they hear someone mention "Vanilla" Kubernetes[0]. The real problem they suffered is actually that Kubernetes isn't fundamentally designed for multi-tenancy. Instead, you're forced to make separate clusters to isolate different domains. Google themselves run multiple Borg clusters to isolate different domains, so it's natural that Kubernetes end up with a similar design. [0]: https://jzelinskie.com/posts/youre-not-running-vanilla-kubernetes/ https://jzelinskie.com/posts/youre-not-running-vanilla-kuber... Disclosure: I worked as an engineer and product manager on CoreOS Tectonic, the (now defunct) Kubernetes used in the post.
- ffo 5y agoHey we used tectonic ;-) was a great tool at that time. Tectonic did influence some of the concepts around ORBOS. Just think of Tectonic combined with GitOps, minus the iPXE part. Disclaimer: I am working with ORBOS
- GauntletWizard 5y agoYou're just wrong, unfortunately - Google runs dev and test and prod all on the same clusters. Kubernetes multi-tenancy works just fine, but the conventional definition of multi-tenancy includes things like "network isolation" that are misguided. Multi-tenancy should be set up (and is within Google) by understanding what is and isn't shared with the environment, and through cryptographic assertion of who you're speaking to. If you want to see the latter part nicely integrated, come to a SIGAUTH meeting and help me argue for it
- q3k 5y ago> If you want to see the latter part nicely integrated, come to a SIGAUTH meeting and help me argue for it. Invitation accepted! :) I've been dying to see an ALTS-like [1] thing that works with Kubernetes. I really should be able to talk encrypted and authenticated gRPC-to-gRPC without ever having to set up secrets or manually provision certificates, dammit. [1] - https://cloud.google.com/security/encryption-in-transit/application-layer-transport-security https://cloud.google.com/security/encryption-in-transit/appl...
- psanford 5y ago> It means that we’d rather just replace a “sick” instance by a new healthy one than taking it to a doctor. This analogy really bothers me. Cattle are expensive. They are an investment. You don't put down an investment just because it got sick. If you have a sick cow you will in-fact call your local large animal vet to come and treat it.
- ska 5y agoYes that statement is wrong. I guess the real difference in the pet/cow calculus is for cattle you probably won't pay more for treatment than the cow is worth; with pets people do this all the time.
- psanford 5y agoYup. If the argument is "don't fall in love with your servers", I'm in agreement. However, the idea that whenever anything weird happens you should just kill you server/cluster and move on without doing any sort of investigation seems like a recipe for disaster. That's a great way to mask bugs that may in-fact be systemic in nature, where they are or will eventually cause service degradation for your customers. I would hate to work in an environment where bugs are ignored and worked around instead of understood and fixed.
- lamontcg 5y agoI like the idea of killing it and moving on without paging someone at 2AM in the middle of the night. Ideally that goes into an async queue of issues though and someone finds the root cause and that goes into a queue of issues to fix, which is actually burned down. I suspect what is happening more often is that the whole stack has so many levels and the SREs responsible for it all don't have the visibility into the stack they need to debug it all, so they use their SLOs as a club to ignore issues as long as they're meeting their metrics until it becomes a firefighting drill. A pile of cargo culted best practices and SLOs replacing hands on debugging.
- tene 5y agoWhat you do with the server/cluster after you take it out of service is up to you. Having automation like this to take things out of service means that you can immediately restore production workloads to full functionality. I'm far more likely to ignore and work around a bug instead of doing a proper investigation when I've got pressure to get production back up because this server/cluster is a Special Snowflake that must be fixed in-place. Hardware fails, and bugs happen. There's no getting around it. Automation to handle this case is a good part of any strategy for identifying, understanding, and fixing bugs.
- tnisonoff 5y agoWhen I worked at Asana, we created a small framework that allowed for blue-green deployments of Kubernetes Clusters (and the apps that lived on top of them) called KubeApps[0]. It worked out great for us -- upgrading Kubernetes was easy and testable, never worried about code drift, etc. [0] https://blog.asana.com/2021/02/kubernetes-at-asana/ https://blog.asana.com/2021/02/kubernetes-at-asana/ (Not written by me).
- coding123 5y agoAWS in my mind can quickly lose the kubernetes war amongst cloud providers. This is every cloud providers chance: EKS on AWS is so damn tied into a bunch of other AWS products that it's literally impossible to just delete a cluster now. I tried. It's tied into VPCs and Subnets and EC2 and Load balancers and a bunch of other products that no longer makes sense now that K8s won. In my opinion it needs to be re-engineered completely into a super slim product that is not tied to all these crazy things.
- ffo 5y agoYou mean EKS needs re-engineering?
- kinghajj 5y agoAlso the fact that a new EKS cluster takes at least 20 minutes to come up and be ready, makes AWS' offering the weakest among the Big Three cloud vendors.
- jen20 5y agoThis isn’t true. I’ve provisioned 53 EKS clusters this week, and every one of them has been up in under 11 minutes with all of the accoutrements. I understand it has been substantially slower in the past.
- nickjj 5y ago11 minutes is still a long time IMO. You can spin up a local multi-node cluster using kind[0] in 1 minute on 6+ year old hardware. I know it's not the same but I really have to imagine there's ways to speed this up on the cloud. I haven't spun up a cluster on DO or Linode in a while, does anyone know how long it takes there for a comparison? [0]: https://kind.sigs.k8s.io/ https://kind.sigs.k8s.io/
- mlnj 5y agoI think it is unfair to compare KinD with an actual Kubernetes cluster which comes with Load Balancers, External IP addresses etc. My Terraform scripts get a HA K3s cluster in Google Cloud VMs in less than 9 minutes, which in my opinion is fantastic.
- deleted 5y ago[deleted]
- rllin 5y agoreally you should treat your entire cloud deployment (sans state where impossible) as cattle
- mrweasel 5y agoWhile I don’t disagree, this is also a reminder that debugging Kubernetes can be terribly complicated.
- aliswe 5y ago> ... the GitOps pattern is gaining adoption for Kubernetes workload deployment. Is it really though? I for one am glad I didnt jump on the bandwagon early. A lot of the articles popping up nowadays mentioning the downsides of GitOps make a lot of sense.
- yongjik 5y agoAt this point, why not just drive it up to the logical conclusion? Treat your business model as cattle, not pets. Customers leaving? Fire up another business until capital runs out, and if it does, no worries, jut hop to another job! Sorry, but I feel like I landed in crazy-land. Kubernetes is already an exercise in how many layers you can insert with nobody understanding the whole picture. Ostensibly, it's so that you can isolate those fucking jobs so that different teams can run different tasks in the same cluster without interfering with each other. Hence namespaces, services, resource requirements, port translation, autoscalers, and all those yaml files. It boggles my mind that people look at something like Kubernetes and decide "You know what? We need more layers. On top of this."
- spondyl 5y agoI mean, depending on the business, employees aren't trusted to understand the whole picture regardless. Many employees at traditional business don't even have administrator access on their laptops, depending on their position so it's not logically inconsistent with how things seem to operate. With that lack of big picture overview, it makes things hard to scrutinise since you're only seeing one sliver of an implementation (ie; "It saves money" without seeing, let alone understanding the technical requirements and vice versa)
- aynyc 5y agoLimited admin practice isn’t just to save money. It’s a security practice and it’s a good one. In a large business, no one understand the whole business maybe except legal department and CEO.
- MuffinFlavored 5y ago> In a large business, no one understand the whole business maybe except legal department and CEO. lol... what? A CEO of a 20,000+ person company literally sees projects as revenue sources and a deadline, nothing more. To get it delivered, he/she walks down the chain of managers until they gets answers. It's might as well be a black hole.
- ulzeraj 5y agoI want to go back in time when naming your servers as X-men characters or Dune houses was a thing. I'm not a big fan of this brave new DevOps world.
- exdsq 5y agoMy first job, tech support, had fantasy server names. That was fun :)
- jkmcf 5y agoRocky & Bullwinkle characters. When we ran out we moved on to Underdog and Friends.
- scrollaway 5y agoYou're still free to do that. But honestly having done DevOps/sysadmin in both of those worlds, i vastly prefer the whole one-service-per-machine, cattle treatment with the one off exception for remote workspaces and the like. Losing a server was often devastating, even with backups. "Don't run your builds right now, people are visiting the website" also sounds a lot like "Don't connect to the internet right now, i need to use the phone". I mean there's "correct" ways to implement the whole pet servers concept but .. why? It's fun, sure, but it's also a waste of time, and when you want to be productive it just tends to get in your way.
- deleted 5y ago[deleted]
- rq1 5y ago> It means that we’d rather just replace a “sick” instance by a new healthy one than taking it to a doctor. Oh god! Please treat your cattle better!
- deeblering4 5y agoAll this does is make me want to go vegan and avoid maintaining the entire k8s farm. Truly, if your software team headcount is under 500 why are you running k8s?
- mschuster91 5y agoIt gives you automatic failover and decent-ish (at least when coupled with Rancher, naked k8s is nuts) management compared to a couple of manually (or Puppet) managed servers. A well implemented mini cluster can and will save you so much time in later maintenance and deployments.
- exdsq 5y agoI'm in a team of 1 but using it because my product is based on spinning up dedicated services for users on demand, so it works well
- allyourhorses 5y agoBecause it let me replace Ansible playbooks with a sensibly designed YAML syntax usually needing 40% the number of lines, and got my deploys down to a roughly constant 6 seconds regardless of the size of the application. Kubernetes let me finally stop thinking about deployment, I can get new apps online in 5-10 minutes or less. Don't even get me started on monitoring.. k8s murders everything else here
- LAC-Tech 5y agoRaising cattle is a lot of work. You have to weigh them regularly, apply treatments for intestinal worms, for lice, move them from pasture to pasture so they don't overgraze. It's a fulltime job. Also if a cow dies, people don't just buy a new one. It represents quite a loss of profit. Also represents a big potential problem on the farm that people will want to resolve - they're you're money makers, if they're dying it's an issue.
- deleted 5y ago[deleted]
- benreesman 5y agoWhen I was at BigCo and our requirements became so complex and demanding that we had to migrate onto serious containerization and orchestration software, well, it was necessary but we all pined for the days when it was a dozen services and 20k boxes and we didn’t need that shit.
- beebmam 5y agoNot everyone in the world practices animal husbandry, so the "cattle" metaphor doesn't make a lot of sense to some of us, like me. I have no idea how "cattle" should be treated, other than they are killed/used for resources.
- pyuser583 5y agoSo I was going on vacation, and had to leave my cats at a “cat hotel.” It cost about $50 a night. I looked it up and for the cost of putting them up in the “hotel” I could have them euthanized and buy new cats three times over. I would never do such a thing, but I did use it to guilt them into not complaining about the hotel. It didn’t work very well.
- secnono 5y agoIt's turtles all the way down.
- Clewza313 5y agoObligatory relevant xkcd: https://xkcd.com/1737/ https://xkcd.com/1737/
- midrus 5y agoThe level of madness and overengineering in the kubernetes world is only comparable to the level of madness going on in the React world. Everyone seems to be thinking they're Google or Facebook, or both.
- bfrog 5y agoKubernetes is already insanely complicated. In practice minor version differences of anything in the stack leads to issues. I get why it all exists but at some point I have to ask, are containers really that much better than an rpm/deb and private corporate repo? Every container is a effectively a chroot. Add to that the lack of easy debugging you’d get from simple packages and daemons. I get that doesn’t work for “cloud” scale or whatever, but i think the excitement over this stuff is overblown.