5 ms·
I both agree and disagree with you. In theory, I agree. Tend your garden, do small upgrades often instead of major upgrades less often. In a homelab this is e
by jacurtis 3y ago
I both agree and disagree with you.
In theory, I agree. Tend your garden, do small upgrades often instead of major upgrades less often.
In a homelab this is easy to implement. In a small organization it is too.
But in the facetious "real world", things are a lot more complicated. I work as an SRE Manager and my team is basically ALWAYS upgrading Kubernetes. New releases drop about as fast as we can upgrade the last ones.
When you work on a large cluster, doing an upgrade isn't a simple process. It requires a ton of testing, several steps, and being very slow and methodical to make sure it is all done properly. Where I currently work, we have 2 week sprints and infrastructure changes must align with sprint cycles. So to promote an upgrade at the fastest possible schedule it requires:
- Week 0: Upgrade Dev environment
- Week 2: Upgrade QA Environment
- Week 4: Upgrade Sandbox Environment
- Week 6: Upgrade Prod Environment
That is the fastest possible schedule. That assumes we do a cluster upgrade every sprint, which is 2 weeks. It also ignores other clusters for other business units. We have 4 primary product lines, so multiply all that work times 4. Plus we have supporting products (like self-hosted gitlab, codescene, and custom tools running in k8s clusters).
I say fastest possible schedule because, we can't keep up with this schedule, but even if we could it is the fastest we could go and still maintain our deployment and infrastructure promotion policies.
With new releases every 3-4 months (12-16 weeks), we are essentially in a constant state of always upgrading kubernetes. Right now my team is 2 versions behind. Skipping versions doesn't make sense because you can't guarantee a safe upgrade when skipping versions.
This is why LTS releases are nice. When you run systems at large scale, it is impractical to upgrade that often. I'd prefer to limit upgrades to no more than twice a year and personally I find annual upgrade cycles to be the best balance between "tending the garden" and "not drowning in upgrade work". LTS releases are usually tests to skip upgrades so that companies can go from LTS release to LTS release, without the need to worry about upgrading every minor version in sync.
Remember, upgrading K8s clusters isn't what my bosses want to hear my team spends our time. They want to know that observability is improving, devs are getting infrastructure support, we are building out new systems, deploying hardware for the product team, running our resiliency tests, etc.. Sure upgrading is part of the job, but i can't be ALWAYS upgrading. I have a lot of other responsibilities.
- solatic 3y agoI don't know what your internal architecture looks like, but on the face of it, arguably you should be running Dev and QA loads on the same cluster. Either you have an organization where an SRE team is responsible for running clusters for teams, in which case why not run Dev and QA on the same cluster, or you are responsible for last-line-of-support for teams responsible for running their own clusters, in which case you say, here's a new version of Kubernetes (e.g. a Terraform module) and in this sprint you are responsible for deploying it through dev, QA, pre-prod, and production clusters. Especially if you have separate Kubernetes clusters per product line (do you really need separate dev Kubernetes clusters per product line?).
- dalyons 3y agounless you have some strict security reasons, why even have separate prod clusters per product? We definitely don’t. seems to kinda miss the point of k8s - aka running all kinds of mixed workloads with separation within a cluster if you need it.
- solatic 3y agoAgreed. Separate prod clusters per product usually fixes an organizational problem (lack of trust that running shared workloads would/could be safe, versus giving each team their own servers) first. A lot of organizations, sadly, prefer to pay for separate clusters than to set up RBAC, ResourceQuotas, etc. Hypothetically, since the scaling limit for Kubernetes clusters is ~10,000 nodes (last I checked), you could have multiple product lines that took up 10k+ nodes each. Then there's no reason why not to split by product line. But in the beginning it should be fine. There's also edge cases - Kubernetes doesn't support setting resource requests or limits for networking or I/O, which you usually solve by setting up taints/tolerations/affinity to manually schedule those workloads onto nodes where you've manually run the numbers. But still not usually a reason to prefer separate production clusters.
- p_l 3y agoArguably, the design of kubernetes is for clusters to be based on physical clusters - to the point of a cluster per aisle (or multiple aisles), and resources of those clusters then being used to deploy applications even across clusters. Or in the small, like Chick-Fil-A, with kubernetes running locally at every restaurant.