16 ms·
Learnings from our years of Kubernetes in production
- nostrebored 3y agoMy question when looking at Kubernetes for small teams is always the same. Why? In the blog, there are multiple days of downtime, a complete cluster rebuild, a description of how individual experts have to be crowned as the technology is too complex to jump in and out of in any real production environment, handling versioning of helm and k8s, a description of managing the underlying scripts to rebuild for disaster (I'm assuming there's a data persistence/backup step here that goes unmentioned!), and on, and on and on. When you're already using cloud primitives, why not use your existing expertise there, their serverless offerings, and learn the IaC tooling of choice for that provider? Yes it will be more expensive on your cloud bell. But when you measure the TCO, is it really?
- teaearlgraycold 3y agoWe're starting to use k8s as a small team because the simpler offerings with GPUs available don't meet our needs. It's clear they're either built for someone else or are less reliable than an EKS cluster would be.
- nostrebored 3y agoI'd encourage you to look at the problem space and evaluate if ECS or an external abstraction layer (like Ray) meets your needs. I've seen both work in completely separate domains (e.g. inference on real time video streams vs. model building) -- but obviously ymmv, tech is a big domain and pretending I understand exactly what you're doing would be silly. Sometimes there is a real answer to the why!
- teaearlgraycold 3y agoAh, well today I learned about ECS. I guess we’ll migrate to that once I need to add complexity to our EKS setup. I’m new to this stuff, so it’s hard to dig through all of the possible different solutions. I looked into Ray a bit but it seemed a little too complicated vs. just running a CUDA accelerated docker container. Most of the streamlined solutions in this space are not made for full stack web developers deploying a service that happens to need a GPU. They’re for ML devs who are trying to own the production side of their part of the product.
- belval 3y agoEspecially considering that the author seems to be using some Azure specific features anyway: > While being vendor-agnostic is a great idea, for us, it came with a high opportunity cost. After a while, we decided to go all-in on AKS-related Azure products, like the container registry, security scanning, auth, etc. For us, this resulted in an improved developer experience, simplified security ( centralized access management with Azure Entra Id), and more, which led to faster time-to-market and reduced costs (volume benefits).
- deleted 3y ago[deleted]
- liveoneggs 3y agok8s is half-baked at best but people enjoy copy-paste yaml recipes, which half-baked products lend themselves to, so it is loved
- auspiv 3y agoI work for a US subsidiary of a very large oil company. We are migrating from Azure to AWS for many things (it is deemed "OneCloud"). A very large number of our new EC2 instances, and even our EKS instances, were provisioned within the last 6 months as T2 instances. Some, if we were lucky, were T3. T3 was released 10 years ago. Copy + paste indeed.
- secondcoming 3y agoThink of the cost savings though!
- deleted 3y ago[deleted]
- menschmanfred 3y agoOur setup works very very well. And in smaller setups you would have a shared cluster or fully managed like gke etc.
- cortesoft 3y agoDo people try to push it that strongly for small teams? Lots of us work on bigger teams and enjoy more of the benefits. However, I also still use Kubernetes for my personal projects, because I really appreciate the level of abstraction it supplies. Everyone always points out that you can do all the things k8s does in other ways, but what I like about it defines a common way to do everything. I don't care that there are 50 ways to do it, I just like having one way. What this allows is for tools to seamlessly work together. It is trivial to have all sorts of cool functionality with minimal configuration.
- karolist 3y agoThis. It's the npm install 100 packages and do everything with JS vs Rails arguments all over again.
- politelemon 3y ago> Do people try to push it that strongly for small teams? Yes. You have to understand that a lot of people without the benefit of experience will often base their technology choices on blog posts. K8S has a lot of mindshare and blog attention, so it gets seen as the only way to run a container in a production environment, while all the important aspects of it are ignored.
- cortesoft 3y agoI get that, but I just get frustrated in the same way I get frustrated with all the "you don't need it" responses to any topic... what about all of us that DO work for bigger companies and DO need to use this stuff? Where can we gather to talk about it without being constantly told we don't need the features?
- ehutch79 3y agoThey don’t read those blogs. And if they do, the decision makers have enough experience to know that “your dog blog doesn’t need k8s” doesn’t apply to their 100000 mau app
- dilyevsky 3y agoI would think it's more dependent on technology requirements more than the size of the team. If all you need is some variation of LAMP stack, then you'd probably be better off with a paas like render, fly or the like.
- nostrebored 3y agoTotally a valid point! I think size of team matters as the impact of k8s ownership as a fraction of your development velocity changes immensely as you're able to afford a platform team who can build tooling to remove the cognitive load of deploying to and managing k8s. At an ~400 engineer company I worked at, k8s bugs that actually impacted our team were in the single digits over a year, but a large part of that was the platform team that managed the ecosystem around k8s deployments.
- datadeft 3y agoSame exact question I ask every single time. We just decided against k8s, again, in 2024. We are going to go with AWS ECS and Azure Container Apps (the infra has to exist in both clouds). ECS and Container Apps provides all the benefits of k8s without the cons. What we want is a to be able to execute container (Docker) images with autoscaling and control which group of instances can talk to each other. What we do not want to do: - learn all of the error modes of k8s - learn all the network modes of k8s - learn the tooling of k8s (and the pitfalls) - learn how to embed yaml into yaml the right way (I have seen some of the tools are doing this) - do upgrades of k8s and figuring out what has changed the way that is backward incompatible - learn how to manage certificates for k8s the right way - learn how to debug DNS issues in a distributed system (https://github.com/kubernetes/kubernetes/issues/110550 https://github.com/kubernetes/kubernetes/issues/110550 and many more) I could go on and on but many people and companies figured out the hard way that k8s complexity is not justified.
- ManBeardPc 3y agoMy experience with Kubernetes has been mostly bad. I always see an explosion of complexity and there is something that needs fixing all the time. The knowledge required comes on top of the existing stack. Maybe I'm biased and just have the wrong kind of projects, but so far everything I encountered could be built with a simple tech stack on virtual or native hardware. A reverse proxy/webserver, some frontend library/framework, a backend, database, maybe some queues/logs/caching solutions on any server Linux distribution. Maintenance is minimal, dirt cheap, no vendor lock-in and easy to teach. Is everyone building the next Amazon/Netflix/Goole and needs to scale to infinity? I feel there is such a huge amount of software and companies that will never require or benefit from Kubernetes.
- mhitza 3y agoCompany CTOs in my experience get sold very easily the idea of infinite scalability. In practice not many companies reach that point, but many that go down this road have to build on top of dozens of layers of compute/networking abstractions that only few experts on the team can manage, if any, competently. I think the cost of self-managed Linux VMs and monoliths is smaller than the cloud vendors made it seem. Containers are nice when you have to deal with a language like Python and it's packaging ecosystem, but when Go/Rust/.Net/etc binaries are placed in containers as well... I think sight of what we're trying to solve in real life has been kind of lost.
- ManBeardPc 3y agoMonoliths are so much easier for smaller teams. No additional tooling needed, no service discovery, instead of networks calls you have function calls, can share resources, etc. Much less overhead as well, so you may not even need to scale. The amount of requests a single Go/Rust server can handle on a dedicated machine is insanely high with modern hardware.
- marcinzm 3y ago> But when you measure the TCO, is it really? I'm in the ML space and every small company I try to avoid EKS. Then I hate my life. Sagemaker, for example, is a giant abstracted away mess with random holes (ie: these types of jobs don't work on this GPU type, etc.) compared to just running things on EKS. The same goes to trying to deploy a more complex third party application. I could just deploy their Helm chart or I could spend a lot of time deploying it somehow in our environment.
- nostrebored 3y agoI had to sell Sagemaker and I agree. It is the wrong abstraction layer without the right escape hatches. I am super pro Ray for handling these types of workloads now. Huge shout out to anyone here working on maintaining that project.
- bionsystem 3y agoI see it all the time at different layers of the stack. At some point some knowledge is lost due to people turnover and the solution is to change the technology, instead of paying somebody full time to re-understand it. Why not rewrite XXX part in YYY language as nobody understands XXX anymore ? Linux VMs require a good sysadmin with a taste in digging into existing scripts and playbooks. With Kubernetes we can start from scratch and we only needs a kubernetes expert ! (or so they say). Right now I'm working a lot with an oversized maven configuration that nobody understands ; I'm paid only to dig into it and maybe refactor some parts. It's made way too complicated for the task and does a lot of non-standard stuff to work around problems it created itself. But when I arrived people were blaming jenkins and wanted to move to gitlab because jenkins was becoming too complicated to work around maven (also !). Next thing you know somebody could try kubernetes or moving to the cloud or switching from RedHat to NixOS or whatever, and the problem would still be maven.
- betaby 3y agoIn 2000s we were talking that `snowflake` servers are bad. New generation is re-learning the same with k8s, which can be summarized as 'snowflake k8s clusters are bad'. Fundamentally it's the same problem.
- menschmanfred 3y agoIt's not. The control plane is ha and you can upgrade one after the other independent of your workers. With workers you can do that too. That feels much less like a snowflake and more like snow.
- betaby 3y agoControl panel as HA as your certificates which were expired in the article.
- menschmanfred 3y agoK8s introduced Auto rotation surprisingly late. Even we run into that issue 5 years ago. But k8s is still very young and the problem is solved for 5-6 years
- chupasaurus 3y agoIt was introduced later than the related crash in the article.
- menschmanfred 3y agoYes. But the article describes years of usage and it just isn't a problem if you start with k8s today
- dilyevsky 3y agoEven if control plane is hard down your kubelet wont be evicting anything so unless you need to surge or everything crashes you have some time to fix things. I’ve dealt with multi-hour cp outages and while stressful it didn’t have any customer visible impact whatsoever
- Bassilisk 3y agoNot a native English speaker, but when exactly did "lessons" get replaced by "learnings"? To me the latter always sounds very unsophisticated.
- ojbyrne 3y agoAs a native English speaker, in my opinion it's incredibly pretentious.
- deleted 3y ago[deleted]
- teaearlgraycold 3y agoNative English speaker - I refuse to use "learnings". It's a ridiculous office-speak word.
- PaulStatezny 3y agoSame – just like "asks". "Here's the ask" versus "here's the request".
- arccy 3y agoat least asks lets you differentiate between human requests and http requests
- nequo 3y agoSoon enough we might be talking about HTTP GET asks and POST asks.
- collingreen 3y agoPlease not this
- 3y ago
- zellyn 3y agoI'm curious: what do you do for developer environments? Do you have a need to spin up a partial subgraph of microservices, and have them talk to each other while developing against a slice of the full stack?
- deleted 3y ago[deleted]
- Moto7451 3y agoCan’t speak for everyone but I have worked in this environment. It can work fine if you allocate a sub slice of CPU time (.1 CPU for example) and small amounts of (overcommitted) memory, and explicitly avoid using it for things that are more easily managed by cloud provider sub accounts and managed services. IE don’t force your devs to manage owncloud or a similar stand in for S3 - use something first party to stand in or S3 itself. This doesn’t always work and the failure mode of committing to this can be doubling your hosting bill if it won’t run locally and densely packed small instances can’t handle your app.
- sdwr 3y agoThat's how we do it - micro services run locally in tilt, pointed at staging services / DB for whatever isn't local. When it works it's great.
- mieubrisse 3y agoCan you clarify more about the "when it works"? What are the pain points you're seeing?
- jscheel 3y agoIt's worked great for us. Every developer runs a dev cluster on their own machines. Services like s3 are transparently replaced with mock versions. We have two builds that can be run, which really just determines which set of helm charts to deploy: the full stack or a lightweight one with just the bare necessities.
- 3y ago
- louwrentius 3y agoWhat I really miss in articles like this - and I understand why to some degree - what the actual numbers would be. Admitting that you need at least two full-time engineers working on Kubernetes I wonder how that kind of investment pay’s itself back, especially because of all the added complexity. I desperately would like to rebuild their environment on regular VMs, maybe not even containerized and understand what the infrastructure cost would have been. And what the maintenance burden would have been as compared to kubernetes. Maybe it’s not about pure infrastructure cost but about the development-to-production pipeline. But still. These is just so much context that seems relevant to understand if an investment in kubernetes is warranted or not.
- karolist 3y agok8s is simply a set of bullet proof ideas to run production grade services forcing "hope is not a strategy" as much as possible, it standardises things like rollouts, rolling restarts, canary deployments, failover etc. You can replicate it with a zoo of loosely coupled products but a monolith which you can hire for with impeccable production record and industry certs will always be preferable to orgs. It's Googles way of fighting cloud vendor lockin' when they saw they're losing market share to AWS. Only large companies need it really, a small 5 person startup will do on Digital Ocean VPS just fine with some S3 for blob storage and CDN cache.
- jurschreuder 3y agoThis is always my exact thought with k8. Why not just have auto-scaling servers with a CI/CD pipeline. Seems so much easier and more convenient. But I guess developers are just always drawn to complexity. It's in their nature that's why they became developers in the first place.
- justinclift 3y ago> But I guess developers are just always drawn to complexity. Not sure why you think that? Developers who like simplicity are a thing, and they seem to really think about how the parts of their systems interact.
- 3y ago
- biggestlou 3y agoCan we please put the term "learnings" to rest?
- 0xbadcafebee 3y agoYou aren't a real K8s admin until your self-managed cluster crashes hard and you have to spend 3 days trying to recover/rebuild it. Just dealing with the certs once they start expiring is a nightmare. To avoid chicken-and-egg, your critical services (Drone, Vault, Bind) need to live outside of K8s in something stupid simple, like an ASG or a hot/cold EC2 pair. I've mostly come to think of K8s as a development tool. It makes it quick and easy for devs to mock up a software architecture and run it anywhere, compared to trying to adopt a single cloud vendor's SaaS tools, and giving devs all the Cloud access needed to control it. Give them access to a semi-locked-down K8s cluster instead and they can build pretty much whatever they need without asking anyone for anything. For production, it's kind of crap, but usable. It doesn't have any of the operational intelligence you'd want a resilient production system to have, doesn't have real version control, isn't immutable, and makes it very hard to identify and fix problems. A production alternative to K8s should be much more stripped-down, like Fargate, with more useful operational features, and other aspects handled by external projects.
- throwboatyface 3y agoHonestly in this day and age rolling your own k8s cluster is negligent. I've worked at multiple companies using EKS, AKS, GKE, and we haven't had 10% of the issues I see people complaining about.
- jauntywundrkind 3y agoOnce your team has upgrades down, everything is pretty rote. This submission (Urbit, lol) seemed particularly incompetent at managing cert rotation. The other capital lesson here? Have backups. The team couldnt restore a bunch of their services effectively, cause they didn't have the manifests. Sure, a managed provider may have less disruptions/avoid some fuckups, but the whole point of Kubernetes is Promise Theory, is Desired State Mamagememt. If you can re-state your asks, put the manifests back, most shit should just work again, easy as that. The team had seemingly no operational system so their whole cluster was a vast special pet. They fucked up. Don't do that.
- 3y ago
- datadeft 3y agoThis is insane: The Root CA certificate, etcd certificate, and API server certificate expired, which caused the cluster to stop working and prevented our management of it. The support to resolve this, at that time, in kube-aws was limited. We brought in an expert, but in the end, we had to rebuild the entire cluster from scratch. I can't even imagine how I could explain any of my customers such an outage.
- bdangubic 3y ago“us-east-1 was down” :)
- datadeft 3y agoIf most infra I worked on was a single region one, sure. :) DR is so much easier in the cloud. You can have ECS scale to 0 in the DR site and when us-east-1 goes down just move the traffic there. We did that with amazon.com before AWS even existed. With AWS it became easier. There are still some challenges, like having a replica of the main SQL db if you run a traditional stack for example.
- dilyevsky 3y agoJust in last couple of years I can recall DataDog being down for most of the day and Roblox took something like 72h outage. If huge public companies managed, you probably can too. I'd argue that unless real monetary damage was done it's actually worse for the customer to experience many small-scale outages than a very occasional big outage.
- geodel 3y agoWell the industry analysts and consultants who develop metrics have decided that multiple outages is the way to go as it keeps people on toes more often. And management likes busy people as they are earning their keep.
- badrequest 3y agoIIRC Roblox was using Consul
- dilyevsky 3y ago> During our self-managed time on AWS, we experienced a massive cluster crash that resulted in the majority of our systems and products going down. The Root CA certificate, etcd certificate, and API server certificate expired, which caused the cluster to stop working and prevented our management of it. The support to resolve this, at that time, in kube-aws was limited. We brought in an expert, but in the end, we had to rebuild the entire cluster from scratch. That's crazy, I've personally recovered 1.11-ish kops clusters from this exact fault and it's not that hard when you really understand how it works. Sounds like a case of bad "expert" advice.
- therealfiona 3y agoIf anyone has any tips on keeping up with control plane upgrades, please share them. We're having trouble keeping up with EKS upgrades. But, I think it's self-inflicted and we've got a lot of work to remove the knives that keep us from moving faster. Things on my team's todo list (aka: correct the sins that occurred before therealfiona was hired): - Change manifest files over to Helm. (Managing thousands of lines of yaml sucks, don't do it, use Helm or similar that we have not discovered yet.) - Setup Renovate to help keep Helm chart versions up to date. - Continue improving our process because there was none as of 2 years ago.
- throwboatyface 3y agoIME EKS version upgrades are pretty painless - AWS has a tool that tells you if any of your resources would be affected by an upcoming change even.
- raffraffraff 3y agoIt's not the EKS upgrade part that's a pain, it's the deprecated K8S resources that you mention. Layers of terraform, fluxcd, helm charts getting sifted through and upgraded before the EKS upgrade. You get all your clusters safely upgraded, and in the blink of an eye you have to do it all over again.
- deleted 3y ago[deleted]
- physicles 3y agoWe address this by not using helm, and not using terraform for anything in the cluster. Kustomize doesn't do everything you'd want from a DRY perspective, but at least the output is pure YAML with no surprises. We upgrade everything once a quarter. Usually takes about four hours per cluster. Occasionally we run into something that's deprecated and we lose another day, but not more than once a year.
- 3y ago
- doctor_eval 3y agoI’m in the very unusual situation of being tasked to set up a self-sufficient, local development team for a significant national enterprise in a developing country. We don’t have AWS, Google or any other cloud service here, so getting something running locally, that they can deploy code to, is part of my job. I also want to ensure that my local team is learning about modern engineering environments. And there is a large mix of unrelated applications to build, so a monolith of some sort is out of the question; there will be a mix of applications and languages and different reliability requirements. In a nutshell, I’m looking for a general way to provide compute and storage to future, modern, applications and engineers, while at the same time training them to manage this themselves. It’s a medium-long term thing. The scale is already there - one of our goals is to replace an application with millions of existing users. Importantly, the company wants us to be self sufficient. So a RedHat contract to manage an OpenShift cluster won’t fly (although maybe openshift itself will?) For the specific goals that we have, the broad features of Kubernetes fit the bill - in terms of our ability to launch a set of containers or features into a cluster, run CICD, run tests, provide storage, host long- and short lived applications, etc. But I’m worried about the complexity and durability of such a complex system in our environment - in the medium term, they need to be able to do this without me, that’s the whole point. This article hasn’t helped me feel better about k8s! I personally avoided using k8s until the managed flavours came about, and I’m really concerned about the complexity of deploying this, but I think some kind of cluster management system is critical; I don’t want us to go back to manually installing software on individual machines (using either packaging or just plain docker). I want there to be a bunch of resources that we can consume or grow as we become more proficient. I’ve previously used Nomad in production, which was much simpler than K8s, and I was wondering if this or something else might be a better choice? How hard is k8s to set up today? What is the risk of the kind of failures these guys hit, today? Are there any other environments where I can manage a set of applications on a cluster of say 10 compute VMs? Any other suggestions? Without knowing a lot about their systems, I suspect something like Oxide might be the best bet for us - but I doubt we have the budget for a machine like that. But any other thoughts or ideas would be welcome.
- geodel 3y agoWell Amazon CEO himself said, there is no shortcut to experience. I am sure gaining experience in developing infrastructure solution will give you respectable return in long term. Of course Cloud vendors will be happy to sell turnkey solutions to you though.
- aguacaterojo 3y agoVery similar story for my team, incl. the 2x cert expiry cluster disasters early on requiring a rebuild. We migrated from Kubespray to kOPs (with almost no deviations from a default install) and it's been quite smooth for 4 or 5 years now. I traded ELK for Clickhouse & we use Fluentbit to relay logs, mostly created by our homegrown opentelemetry-like lib. We still use Helm, Quay & Drone. Software architecture is mostly stateless replicas of ~12x mini services with a primary monolith. DBs etc sit off cluster. Full cluster rebuild and switchover takes about 60min-90min, we do it about 1-2x a year and have 3 developers in a team of 5 that can do it (thanks to good documentation, automation and keeping our use simple). We have a single cloud dev environment, local dev is just running the parts of the system you need to affect. Some tradeoffs and yes burned time to get there, but it's great.
- bongripper 3y ago[dead]
- cottsak 3y agoThe challenge with these editorials is that you can never really capture the opportunity cost from the author - they didn't build the monolith without the complex infrastructure as an alternative in parallel. I wonder how much 'sunk cost' and other psychological factors play into the statement: > We started with Kubernetes a bit too early, It feels like a "it wasn't that bad". But if you consider the dedicated resources, costs of pain and suffering that may have been avoided with a simpler architecture and infrastructure, I'm forced to wonder if this short comment hides a lot under the covers. Engineers can have very high pain thresholds sometimes.
- justinclift 3y agoIt also sounds like they have several staff just dedicated to managing their Kubernetes cluster, even though they're using a managed Kubernetes service (AKS) for the last few years. Wonder if they're including those engineers in the cost calculation?
- rwmj 3y agoI also didn't understand what they were building. Is it a simple database-backed website or something truly complicated? How many users do they have? What's the scale of operations? Why wouldn't a handful of dedicated hosts in a datacenter have worked?
- moondev 3y agoI was hoping they were still running the original hyperkube bootstrapped 1.0 release (July 2015) I believe certs were optional back then so no need to rotate!
- tflinton 3y agoCerts expiring is a common occurrence and a source of many RCAs. Not keeping definitions of your configuration separate from running servers (and no baseline) is a big issue. Not keeping secrets in a secret storage and syncing them is another red flag. Thing is, none of these are kubernetes issues. They’re poor practices, these aren’t lessons of running kubernetes they’re poor management of a system.
- hannofcart 3y ago> Also, did we even face the problems Kubernetes solves at that stage? One might argue that we could have initially gone with a sizable monolith and relied on that until scaling and other issues became painful, and then we made a move to Kubernetes (or something else). Probably true for many early stage projects.
- davidelettieri 3y agoAny documentation on this? > This is very context-specific, but depending on the node type, AKS reserves about ~10-30% of the available memory (for internal AKS services)
- jonsson101 3y agoGood point! 25% of the first 4GB of memory, 20% of the next 4GB of memory (up to 8GB), 10% of the next 8GB of memory (up to 16GB), 6% of the next 112GB of memory (up to 128GB), 2% of any memory above 128GB "AKS reserves an additional 2GB for system process in Windows nodes that are not part of the calculated memory." https://learn.microsoft.com/en-us/azure/aks/concepts-clusters-workloads https://learn.microsoft.com/en-us/azure/aks/concepts-cluster...
- oschvr 3y agoThis brings so much painful memories of an entire weekend I had to spend renewing cluster certificates one by one whilst being yelled at that we were not making any money. Good thing our control planes are now managed. I learnt so much about kubeadm and the inner workings of k8s but I'm not sure I wanna go over it again
- hardwaresofton 3y ago> The Root CA certificate, etcd certificate, and API server certificate expired, which caused the cluster to stop working and prevented our management of it. I've run into this and learned my lesson/gained my battle scars, but it just seems like unnecessary pain. Would it have been so bad for k8s to use something simple for securing communications other than the full TLS stack, right from the beginning? It's so cumbersome and so many people run into this footgun that also happens to be proper security practice. A symmetric key setup is simple and if it was available as a fallback all this pain could be avoided. It's not as secure, and you have to be careful with nonces and things but I'll take some somewhat distant insecurity (if someone is already inside your network and reading your asymmetric secrets you have other problems) for the better ergonomics and lower likelihood of blowing off my own foot.
- jnwatson 3y agoRolling your own key management system is not to be taken lightly. I've done it, and you really, really only want to do it when you really know other systems won't work.
- hardwaresofton 3y agoYeah but this isn't rolling your own key management system. This is the stupid simple every machine/program has the same shared secret approach. The difficulty is securing comms between components (assuming they can reach each other, just making sure that the payloads are secret) and making sure you don't leak secrets unintentionally (forgetting nonces) and all the other hard crypto things. But, it's not impossible to make a reasonable to use fallback system that does this, just no one does because of fear of being mocked for not just accepting the pain and bad ergonomics of TLS. Other systems do work, but they have the footguns mentioned in the article that everyone seems to hit.
- roydivision 3y agoThe lessons noted in the article, all valid, can easily be generalised applied to any infrastructure, not just Kubernetes.
- jonsson101 3y agoThanks :)
- wg0 3y agoThe DX of k8s is deceptively simple. You might be forgiven for believing that you have got your own private Heroku in your own backyard which you totally own and control. But - Oh boy! The complexity over complexity of moving parts that themselves are a moving target sometimes. First, there's PKI which you should know all about certificates, signing, expiring, issuing, reissuing or no part of the cluster talks to the other. If you think you can get away with that, the post above had two outages both related to certs. Next is the etcd - you should know how to configure that. A totally different separate product in its own right, a distributed key store that has whole memory of the system. Like what's where and such. Then you have whole DNS running. That again is a whole separate product in its own right whose administration you must master or else. And then comes Networking, the CNI plugin and their internals and if you think you can skip that part, either you have to pay likes of weaveworks (defunct) or Cillium etc or be ready for an incident. And yet I have not talked about ingress controllers, cloud controllers, their configurations and other issues. To top that all, you need to manage all that configuration and package it so now you need Helm and Flux - templated (Helm uses Go templates, if not worst out there) and layered (kustomised) YAML upon YAMl, thousands of lines and that's not just all, hold on! All that configuration language is constantly a moving target from Helm Charts to k8s manifests. Sometimes totally incompatible (like flux to flux2 was almost not so portable) and such so upgrades are going to be so much painful, you just can't imagine even if you're on a managed k8s platform. I say this from my own experience of setting up k8s self managed from scratch across different clouds. I have my scars.
- hasoleju 3y agoMain takeaway from the article and the highly rated comments: Don't use selfmanaged k8s.
- INTPenis 3y agoI felt like I got on the k8s bandwagon relatively late, due to my natural distrust of new things, in this case containers. So I started setting up k8s clusters on-prem at work in 2019. It's 3.5 years later and my takeaway is that k8s should only be used after multiple resource planning meetings have established that it is absolutely necessary, or you're scaling up an existing application. The mistake we made was to first of all take orders from a lead developer who thought that service mesh was the solve all. And secondly to not estimate the load our system would generate, the resources it would require and how to best manage those. In retrospect we now have 3 clusters (dev, staging, prod) for a service that could be hosted on 2 servers, with a job manager in the cloud. And IaC is no excuse because you can achieve the same IaC with container hosts and quadlets set to auto-update.
- louwrentius 3y agoI bet this is probably true (app would have been fine on a hand full of servers) for most Kubernetes installations.
- nucleardog 3y agoCompletely agree. Kubernetes makes some problems easier to solve, but it's rare that anyone asks "do we actually need to solve these problems". It's like buying something you don't need because it's on sale. I've found a happy middle ground to be Hashicorp's nomad. It's a single static go binary you can run so you can define some of your setup in a repeatable way as well as provides things like rolling updates, task monitoring, scheduled jobs, etc. And it's not limited to only running containers, but can run executables directly on the host, VMs with qemu, etc. If I'm running a single server, I usually get a lot of value out of throwing nomad on it.
- Fischgericht 3y agoYes, I am an old boring fart that isn't cloud native. And this post is going to offend a lot of people. I am sorry for that, but I still believe this is an important point to make: From my perspective, the root issue is that people are using the wrong tools for their projects, namely the wrong programming language. The problem started when people began abusing scripting languages for something else than scripting. Python was meant as a teaching language for kids. JavaScript was meant to for some gimmicks on Websites. .net was meant for UI applications. PHP stood for "Personal Home Page Tools". But somehow people started using these tools for something completely different than what they were meant for. Due to people abusing those languages to write server backends, it suddenly became a problem of those languages creating code that's 1000 to 75.000 times slower than native code. That then brought the need for clusters, load-balancers, etc. And that the need for tools to manage those. Then they needed tools to manage those tools. And tools to manage the people who manage those tools. And every year I look at the industry, another layer of complexity is needed. So, the author writes: "Our platform back then was 50% .Net and 50% Python" - and here is the ACTUALY learnings he should have taken away: "Eight years ago, after our 'developers' had drafted our products in scripting languages, we hired a senior C/C++/Pascal/Rust programmer. That developer re-wrote our drafts in a clean way. He used a profiler and checked for memory leaks, and did some optimizations. Afterwards we bought three servers in two different data centers with two upstream providers each. Since then we had zero downtime, our servers are only consuming 800 Watts in total, our operating costs are minimal, and we are contributing towards having a greener planet. Looking at our company growth rate, those six servers will be good for another 8 years." Think that's hyperbole? No, it's not. If your interpreted scripting language is 1000x slower than a compiled one, you'll simply need 1000x the resources. You are burning our planet and your money just because you weren't able to accept the fact that it's OK to use a scripting language for quick hacks and... scripting, but that it's the wrong tool for high workloads. So, for your next project, please consult this checklist: [ ] Is the project about teaching kids how to program? [ ] Is the project about adding a blinking button to your website? [ ] Is the project about running a desktop application on a Windows PC? [ ] Is the project about doing your personal home page? In case you have not checked any of the above boxes, you might want to consider having your code re-written in native code, avoiding 90% of your management layers and dependency hell. And as a bonus, with the CO2 you have saved the planet you are able to buy another two SUVs! ;)
- mdrob2205 3y ago"Lessons" not "Learnings". Spare us all the corporate-speak horseshit. Waves fist at cloud.
- mdrob2205 3y ago"Lessons" not "Learnings". Please please spare the world more corporate-speak h*rsesh*t. Waves fist at cloud.
- YourDadVPN 3y agoWhy are people using "learnings" as a word, when "lessons" already fits all its use cases? It's quite jarring.