29 ms·
Common mistakes using Kubernetes
- xrd 6y agoI wish there was a way to upvote something 10x once a month here. This would be the post I use that on. When I was writing my book my editor asked me to remove any writing about mistakes and changes I made in the project for each chapter. I had a bug that appeared and I wanted to write about how I determined that and fixed it. They said the reader wants to see an expert talking, as if experts never make mistakes or need to shift from one tact to another. But, I find I learn the most from explanations that share how your mental model was wrong initially and how you figured it out and how you did it "more right" the next time. That's really how people build things.
- fnord123 6y ago>They said the reader wants to see an expert talking, as if experts never make mistakes or need to shift from one tact to another. Your editor was very fucking wrong.
- natefox 6y agoSo, so wrong. How did (s)he think experts get so good? Isn't the phrase `the master has failed more times than the apprentice has even tried` well known for a reason?
- emerongi 6y agoIt obviously depends on the type of the book and the reader's expectations. It just might not have been the type of book where you write about things like that.
- ohyeshedid 6y ago>It just might not have been the type of book where you write about things like that. I think that'd be more of the author's choice, instead of the editor.
- a1369209993 6y ago> the master has failed more times than the apprentice has even tried I've never actually heard that one before, but it's so very, very true.
- gridlockd 6y ago> Your editor was very fucking wrong. The editor is completely right in what they were saying. You just want them to be wrong, because you'd prefer to live in the fantasy world where they are wrong. Let's say you go to get a surgery. You don't want the doctor to tell you about all the times they fucked up and what the awful consequences were. It doesn't matter that they're probably a better surgeon now, having learned from their mistakes. Psychologically, you need that person with the sharp tool poking around inside your body to be a superhuman. To a lesser degree, the same is true for any expert. Of course everybody makes mistakes. Notice the de-personalization in the word "everybody". You can talk about the mistakes everybody makes, or those ones that many people make. If you talk about your own mistakes however, you lose the superhuman status. There may be a few situations where that somehow helps you, but not when you want to sell books.
- fnord123 6y agoI wonder if we're on the same page. Should a book on a programming language discuss compilation errors and how to interpret them? Should a book like Effective C++ (afaik the most popular series of books on C++) exist?
- gridlockd 6y agoThe answer (yes) is already in my reply, in the second paragraph.
- sokoloff 6y agoI sure do want that surgeon to have been presented/instructed on the common ways that surgeons before them have made mistakes and how to avoid or overcome them. That they won’t tell me (the patient) is quite a different question from whether they got the material from someone more experienced in their primary or continuing medical education.
- gridlockd 6y agoThe situation is different, the psychological effect is not.
- xrd 6y agoObviously, I agree with you! I also think this is the way the majority of tech books are written. Can you think of another where the author goes from mistake to mistake and then finally gets it right? I believe there is a space in tech writing for this kind of writing, but it is not something traditional book publishers believe. This was an O'Reilly book by the way, with really good editors and a really good editoral process. That editor was right most of the time, IMHO.
- fnord123 6y ago>Can you think of another where the author goes from mistake to mistake and then finally gets it right? Not a book, but Raymond Hettinger often presents in this way and it's fantastic: https://www.youtube.com/watch?v=wf-BqAjZb8M https://www.youtube.com/watch?v=wf-BqAjZb8M
- drvdevd 6y agoMost (in-depth, technical) security books follow this sort of thought process- I think. “A big hunters diary” comes to mind.
- fnord123 6y agoIgnition! A history of liquid rocket propellants. Little Lisper, Little Scheme is a conversation style that discusses a lot of false starts. {Coders|Founders} at Work has a lot of frank talk about early mistakes people make and their pivots.
- drvdevd 6y agoAgreed. The entire expertise of programming is based on the willingness to find these sorts of issues and fix them. Over and over, for your entire career.
- Leonid99 6y agoMost books frame this as "one common mistake is ...". I always understood this phrase to mean that the author has personally made this mistake more than once.
- mgkimsal 6y agopersonally... I'd like to see that sort of info, but typically not in the middle of a chapter/section. make a note or call out in the text pointing to a section on why/how you got to the 'correct' position. That info is often helpful info, but can disrupt the flow of the 'good' information.
- barrkel 6y agos/tact/tack/ Tact means skill in dealing with other people, particularly in sensitive situations. Tack has to do with sailing boats into the wind, and crucially "changing tack" means changing direction.
- xrd 6y agoThanks, I'm glad I had you as an editor this time! I never noticed that before.
- namelosw 6y agoI hate this so much. But it's so common. I was told not to mention the caveats, instead, render a confident image for the team many times in my career. It's like doctors suggest their patients to use some drugs without mentioning any side effects. It reminds me of the video[0] which asks the developer to draw red lines with blue ink, while the project manager keeps pushing the developer like "You're the expert, of course you can do draw red lines with blue ink!". [0] https://www.youtube.com/watch?v=BKorP55Aqvg https://www.youtube.com/watch?v=BKorP55Aqvg
- m463 6y agoyears (and years) ago, if I wanted to learn a new computer technology or language, I would pick up a book and learn it. One language that didn't go how I expected it was applescript. The book on applescript (from o'reilly I believe) was necessarily different. Instead writing "normal" top-down programs, applescript hooks into the OS from the side. Many books can concentrate on what you can do, but the applescript had to do everything by example. This is because all of apple's applications expose interfaces that you have to figure out on the fly. I had to learn a different way from that book, because it necessarily concentrated on "how" instead of "what".
- raverbashing 6y ago> asked me to remove any writing about mistakes and changes I made in the project for each chapter. I had a bug that appeared and I wanted to write about how I determined that and fixed it. Your editor is right. Not because this information is not interesting. But it is distracting and off-topic to go off a tangent every now and then. Especially for someone just learning the stuff, it will confuse them more than anything else.
- lolkube 6y ago#1: Using Kubernetes
- dewey 6y agoYou spent the time signing up just for this?
- toomuchtodo 6y agoIt’s an important point for those who may unknowingly over engineer.
- dewey 6y agoAnd they are probably not going to change their mind by reading an anonymous sarcastic comment on HN.
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]
- Nzen 6y agoFor someone with a short time scale, only trawling this thread, of course not. For someone young in this space, this comment and one hundred others -for and against- sift into hiso'er consciousness as part of the perceived zeitgeist of kubernetes within the larger community. Perhaps after several years, this person may have an intuition to avoid kubernetes in favor of separate docker or lxd containers. That association of kubernetes as a FAANG-level tool builds stronger with the linked article: this hypothetical person can compare the struggles against perceived resources and so on. But not everyone has time to read the article, any given kubernetes article, that we come across. Some of those times, it's enough to take the temperature and move on. So, -over years- that may build an aversion that would not have otherwise formed had commenters avoided denouncing (or endorsing) kubernetes with less than full commitment toward convincing others. Also, in this toneless medium, I can't intuit much of the emotional weight lolkube conveyed the sentence with. Was this person rueful, playful ?
- haolez 6y ago... requiredDuringSchedulingIgnoredDuringExecution: ... This instantly remembered me of this: https://thedailywtf.com/articles/the-longest-method https://thedailywtf.com/articles/the-longest-method Kubernetes sometimes shows its Java roots.
- vsareto 6y agoThat Java link in the article goes to a completely not-Java Chinese site btw
- Igelau 6y agoAnd the Spring link is a 404
- tomnipotent 6y agoTo be fair, I don't see how this could be shorter in any other language without losing readability.
- yongjik 6y agoYou don't have to. Kubernetes should just say "required" here, and the documentation should say "Warning: this is only checked during scheduling." printf isn't named printIntoBufferWhichMayNotFlushUntilLinefeed, and people are fine with it.
- dharmab 6y agoI believe there may have been a plan to allow checks during runtime at some point; although this feature is no longer necessary since https://github.com/kubernetes-sigs/descheduler#removeduplicates https://github.com/kubernetes-sigs/descheduler#removeduplica... can do it.
- kerny 6y agoThey could have used: soft, hard and strict with the documentation explaining the differences.
- bavell 6y agoGreat article, I've learned many of these firsthand and agree with their conclusions. I have some more reading to do on PDBs! K8s is a powerful and complex tool that should only be used when needed. IMO you should be wary of using it if you're trying to host less than a dozen applications - unless it's for learning/testing purposes. It's a complex beast with many footguns for the uninitiated. For those with the right problems, motivations and skillsets, k8s is the holy grail of scale and automation.
- watermelon0 6y agoI'm not necessarily agreeing with you. Kubernetes really is a complex beast, and I wouldn't recommend self-hosting it for companies that don't have people that can focus solely on managing it. I would also not recommend it for hosting a WordPress site, or a simple CRUD app. However, when you get to the level where autoscaling is required, and where you are deploying multiple services, managed Kubernetes is not such a bad idea. Using EKS (especially with Fargate) on AWS is not much harder than figuring out and properly utilizing EC2/ASG/ELB. GKE or DigitalOcean offerings seem to be even easier to use and understand.
- toshk 6y agoWhat would you suggest as an alternative, simpler form for docker deploy, running and managing? Docker-compose?
- ghaff 6y agoWhat do you mean by docker hosting? Kubernetes (and other related tools) are container orchestration/management tools. As if often the case in the management space, if you're just running at small scale, you may not need anything beyond container command line tools and some scripts. You could also use Ansible to automate.
- toshk 6y agoThanks we are running a few node servers, we now deploy command line. Dev we use docker-compose. But we are looking for a way to easily share our servers. We developed it for Amsterdam open source. Around 20-30 cities are in line to start using it. Doesn't have to be scalable, or have high availibility. Ease of deployment, easy way to update and basic security. All sysadmins are pushing for kubernetes, although for the big cities it makes sense, it really starting to feel like an overkill for small cities who will run 1-3 non-critical sites with 0.5-5k users p/m. Heard a lot about ansible, will look into it, thanks!
- config_yml 6y agoWhat would a good liveness and readyness probe do for a rails app? What kind of work and metrics would these 2 endpoints do in my app?
- UK-Al05 6y agoIf it has network problems, kubernetes can take out that instance out of serving traffic. When your doing rolling upgrades it can signal your app is ready to take traffic. Those are the main uses.
- Nextgrid 6y agoFor the readiness probe a simple endpoint that returns 200 is enough. This tests your service’s ability to respond to requests without depending on any other dependencies (sessions which might use Redis or a user auth service which might use a database). For liveness probe I guess you could check if your service is accepting TCP connections? I don’t think there should ever be a reason for your service to outright refuse connections unless the main service process has crashed (in which case it’s best to let Kubernetes restart the container instead of having a recovery mechanism inside the container itself like supervisord or daemon tools).
- kevindong 6y ago> For the readiness probe a simple endpoint that returns 200 is enough. This tests your service’s ability to respond to requests without depending on any other dependencies (sessions which might use Redis or a user auth service which might use a database). If the underlying dependencies aren't working, can a pod actually be considered ready and able to serve traffic? For example, if database calls are essential to a pod being functional and the pod can't communicate with the database, should the pod actually be eligible for traffic?
- Nextgrid 6y agoThe article explicitly warns against that: > Do not fail either of the probes if any of your shared dependencies is down, it would cause cascading failure of all the pods. The idea would be that the downstream dependencies have their own probes and if they fail they will get restarted in isolation without touching the services that depend on them (that are only temporarily degraded because of the dependency failure and will recover as soon as the dependency is fixed).
- kirstenbirgit 6y agoLots of good advice in this article.
- speedgoose 6y agoIn my opinion, the most common mistake is not in the article : using kubernetes when you don't need to. Kubernetes has a lot of pros or the papers but in practice it's not worth it for most small and medium companies.
- balfirevic 6y agoDo you also include managed kubernetes offerings, such as from Digital Ocean, in that assessment?
- blinkingled 6y agoNot the OP but yes if cost is a factor. As far as I know no managed K8S offerings are cheap.
- deleted 6y ago[deleted]
- watermelon0 6y ago- DigitalOcean offers free clusters. - Azure (still?) offers free clusters. - GCP offers one free single-zone cluster. If you need more than DO can offer in terms of compute instances, you can probably afford GKE/EKS, which is around 75$/month. --- Running highly-available control plane (K8s masters & etcd) by yourself is NOT cheaper than using EKS. To achive high availability, EKS runs 3 masters and 3 etcd instances, in different availability zones. Provisioning 3 t3.medium instances (4 GB of memory and 2 CPUs) would cost the same as a completely managed EKS. Not to mention the manual work you need to setup, maintain and upgrade such instances.
- speedgoose 6y agoOh yes. Managed kubernetes is full of various issues. Some major cloud providers sell very poor managed kubernetes. If someone knows about a reliable managed kubernetes, please let me know.
- 6y ago
- fergonco 6y agoShameless plug but on topic. I wrote recently about readiness and liveness probes with Kubenetes. If you look for an educational perspective you can check: https://medium.com/aiincube-engineering/kubernetes-liveness-and-readiness-probes-with-spring-boot-185af0d5b5de https://medium.com/aiincube-engineering/kubernetes-liveness-...
- chx 6y agoIt misses the biggest one: using it. I ranted about the cloud a decade ago http://drupal4hu.com/node/305 http://drupal4hu.com/node/305 and there's nothing new under the Sun. Still most companies doing cloud and Kubernetes doesn't need it... practice YAGNI ferociously.
- tcbasche 6y agoI think you may be on the wrong side of history here
- chx 6y agoNothing new with that one. I still think git was the wrong choice for DVCS and yet, I have been using it since for a decade or more now. I am still still feeling I have an uneasy truce with it but not a friendship. I still think github is a shitty choice for hosted git -- at least most large open source projects have went with gitlab so I am not utterly alone with that. I am using now Kubernetes because my primary client is using it. Doesn't mean I am happy with it or that I think it's necessary by any means. It's fine. I am getting old but I still can learn. Doesn't mean I can't be grumpy about it.
- mlthoughts2018 6y agoI think there is more to the story for some of these points and it can be dangerous to just take this at face value of best practices. For example on the liveness / readiness probe item, the article says, > “ The other one is to tell if during a pod's life the pod becomes too hot handling too much traffic (or an expensive computation) so that we don't send her more work to do and let her cool down, then the readiness probe succeeds and we start sending in more traffic again.” But this is often a very bad idea and masks long term errors in underprovisioning a service. If the contention of readiness / liveness checks vs real traffic is ever resulting in congestion, you need the failure of the checks to surface it so you can increase resources. If you set things up so this failure won’t surface, like allowing the readiness check to take that pod out of service until the congestion subsides, you’re only hurting yourself by masking the issue. It basically means your readiness check is like a latency exception handler outside the application, very bad idea. The other item that is way more complicated than it seems is the issue about IAM roles / service accounts instead of single shared credentials. In cases where your company has an enterprise security team that creates extremely low-friction tools to generate service account credentials and inject them, then sure, I would agree it’s a best practice to ruthlessly split the credentialing of every application to a shared resource, so you can isolate access and revoking. But if you are on some application team and your company doesn’t have a mature enough security tooling setup managed by a separate security team, this can become a bad idea. It can lead to superlinear growth in secrets management as there will be manual service account creation and credential propagation overhead for every separate application. Non-security engineers will store things in a password manager, copy/paste into some CI/CD tool, embed credentials as ENV permanently in a container, etc., all because they can’t create and maintain the end to end service account credential tools in addition to their job as an application team engineer. It’s something they think about twice per year and need off their plate immediately to move on to other work. Across teams it means you end up with 20 different team-specific ways to cope with rapid growth of service accounts, leading to an even worse security surface area, risk of credential-based outages, omission of important testing because ensuring ability to impersonate the right service account at the right place is too hard, etc. Very often it is a real trade-off to consider that one single service account credential that has just one way to be injected for every service is safer in the bigger picture. Yes it means a credential issue for any service becomes an issue for all, and this is a risk and you want automated tooling to mitigate it, but it very often will be less of a risk than insisting on a parochial best practice of individual service account credentials, resulting in much worse and less auditable secrets workflows overall unless it is completely owned and operated by a central security team in such a way that it doesn’t create any approval delays or workflow friction for application teams.
- dang 6y agoI'm glad readers are liking the article, but please read and follow the site guidelines. Note this one: If the title begins with a number or number + gratuitous adjective, we'd appreciate it if you'd crop it. E.g. translate "10 Ways To Do X" to "How To Do X," and "14 Amazing Ys" to "Ys." Exception: when the number is meaningful, e.g. "The 5 Platonic Solids." The submitted title was "10 most common mistakes using kubernetes", HN's software correctly chopped off the "10", but then you added it back. Submitters are welcome to edit titles when the software gets things wrong, but not to reverse things it got right.
- marekaf 6y agoOops, sorry. I thought I made a typo that is why I corrected it. Thanks for pointing that out.
- hinkley 6y ago> You can't expect kubernetes scheduler to enforce anti-affinites for your pods. You have to define them explicitly. Why isn't this the default behavior? Why don't I have to go in and tell it that it's okay to have multiple instances on the same node? Why? So that I somehow feel like I've contributed to the whole process by fixing something that never should break in the first place? I know of a few pieces of code where I definitely want to run N copies on one machine, but for all of the rest? Why am I even running 2 copies if they're just going to compete for resources?
- nielsole 6y agoPod anti affinities did historically dramatically increase scheduling times. Not sure this is the primary reason, but probably one
- jeffbee 6y agoIt's quite possible that you have a machine with 192 CPU cores in it, but it's very unlikely that you are able to write a service that scales to that level ... and if you write it in Go it's really unlikely that you can scale even to 8 CPUs. There's nothing weird about having multiple replicas of the same job on the same node. If you look through the Borg traces that Google recently published you can find lots of jobs with multiple replicas per node.
- hinkley 6y agoThis is not how defaults work. When you are talking about the realm of the possible, you provide settings that allow you to reach the scenarios that you feel are reasonable, desirable, or lucrative (or commonly enough, some happy combination of the three). Defaults are the realm of the probable. And nobody is requisitioning a 192 core machine without a good bit of due diligence, which would include deciding how to set server affinity.
- jeffbee 6y agoYou're suggesting that preventing multiple replicas of the same job to schedule on the same machine as a good default. There's no evidence to support your conclusion, and my experience it quite the opposite. It is much better if people running batch jobs just schedule 100000 tiny replicas, and let the scheduler sort it out. This provides the cluster scheduler with plenty of liquidity. Multiple small processes are more efficient than a shared-nothing single process.
- zegl 6y agoGreat post! If you're in the Kubernetes space for long enough, you'll see all of these configuration mistakes happening over and over again. I've created a static code analyzer for Kubernetes objects, called kube-score, that can identify and prevent many of these issues. It checks for resource limits, probes, podAntiAffinities and much more. 1: https://github.com/zegl/kube-score https://github.com/zegl/kube-score
- kakakiki 6y agoExcellent tool. Can this analyse the result from kustomize files rather than actual k8s YAML?
- zegl 6y agoYes, kustomize is not supported natively, but you can achieve effect by piping the kustomize output to kube-score. kustomize build | kube-score score -
- kakakiki 6y agoThanks. One more question. For Visual Studio Code, Microsoft has a plugin called Kubernetes - which I currently use. Have you done a comparison against that?
- otterley 6y agoI actually disagree with the first recommendation as written - specifically, not to set a CPU resource request to a small amount. It's not always as harmful as it might sound to the novice. It's important to understand that CPU resource requests are used for scheduling and not for limiting. As the author suggests, this can be an issue when there is CPU contention, but on the other hand, it might not be. That's because memory limits are even more important than CPU requests when scheduling: most applications use far more memory as a proportion of overall host resources than CPU. Let's take an example. Suppose we have a 64GB worker node with 8 CPUs in it. Now suppose we have a number of pods to schedule on it, each with a memory limit of 2GB and a CPU request of 1 millicore (0.001CPU). On this node, we will be able to accommodate 32 such pods. Now suppose one of the pods gets busy. This pod can have all the idle CPU it wants! That's because it's a request and not a limit. Now suppose all of the pods become fully CPU contended. The way the Linux scheduler works is that it will use the CPU request as a relative weight with respect to the other processes in the parent cgroup. It doesn't matter that they're small as an absolute value; what matters is their relative proportion. So if they're all 1 millicore, they will all get equal time. In this example, we have 32 pods and 8 CPUs, so under full contention, each will get 0.25 CPU shares. So when I talk to customers about resource planning, I actually usually recommend that they start with low CPU reservation, and optimize for memory consumption until their workloads dictate otherwise. It does happen that particularly greedy pods are out there, but that's not the typical case - and for those that are, they will often allocate all of a worker's CPUs in which case you might as well dedicate nodes to them and forget about how to micromanage the situation.
- jeffbee 6y agoIf you ask for 0.001 CPU share, you might get it. I would advise caution. You that pod gets scheduled on a node with another node that asks for 4 CPUs and 100MB of memory, it's not going to get any time.
- otterley 6y agoIt depends. If the second pod requests 4 CPUs, it doesn't necessarily mean that the first pod can't use all the CPUs in the uncontended case. A lot of this depends on policy and cooperation, which is true for any multitenant system. If the policy is that nobody requests CPU, then the behavior will be like an ordinary shared Linux server under load - the scheduler will manage it as fairly as possible. OTOH, if there are pods that are greedy and pods that are parsimonious in terms of their requests, the greedy pods will get the lion's share of the resources if it needs them. The flip side of overallocating CPU requests is cost. This value is subtracted from the available resources, making the node unavailable to do other useful work. Most of the time I see customers making the opposite mistake - overallocating CPU requests so much that their overall CPU utilization is well under 25% during peak periods.
- apple4ever 6y ago"more tenants or envs in shared cluster" This is what I'm trying to convince my current company about. They want everything in a single cluster (prod, test, stage, qa). Of course self hosting makes this more difficult to justify, since it is additional expenses for more machines.
- LambdaB 6y agoHave you considered using OpenShift instead of Kubernetes? It comes with vastly improved multitenancy features, as well as other aspects, in regards to plain Kubernetes. OKD, the open sourced package of OpenShift allows full self-hosting: https://www.okd.io https://www.okd.io
- apple4ever 6y agoOpenShift comes with its own headaches from my understanding. And we are too deep into Kube to switch now.
- moondev 6y agoThat sounds like a disaster waiting to happen! Seems like a perfect use-case for Cluster API: https://cluster-api.sigs.k8s.io/user/quick-start.html https://cluster-api.sigs.k8s.io/user/quick-start.html Have one global "mgmt cluster" with several workload clusters
- apple4ever 6y agoYou bet it is! Haha Ah interesting! We'll have to look into that.
- gridlockd 6y agoCommon mistake: using Kubernetes
- zomglings 6y agoReally great article. I have used Kubernetes pretty heavily in the past, and didn't know about PodDisruptionBudget.
- sixhobbits 6y agoThis also needs a companion post called "common mistakes: using kubernetes". I feel like it's a weekly occurrence now where I hear of a startup launching their mvp on kubernetes having spent 8 months too long on Dev as a result. The other day in an interview someone bragged to me how he had convinced his team to spend 12 months moving to K8s. Upper management thought it was a waste of time but eventually agreed. I asked him if there were any measurable benefits and he said no. I totally understands why Google needs it. Do you?
- yongjik 6y agoEven Google doesn't need it that much: back when I was there, each Borg cluster had something like 10,000+ cores. Large enough to run a typical SV startup wholescale. The ratio of "cluster management work" vs. "actual work being done on it" was not that high. These days, some people are like "Dude, if you don't have one cluster per AWS availability zone per each environment, you're doing it wrong." Why, just why.
- snupples 6y agoYes we need it. This is absolutely becoming a tiresome trope. K8s is a huge benefit to tons of companies and none of them are Google. Yes some people are using K8s when they don't need it. Just like many are using cloud managed services when they don't need them. Or vms. Or insert any technology here. This article has nothing to do with whether K8s fits some particular use case but may be of help (although I disagree with the entire section on resources which reflects a lack of long term experience with K8s in production) to those who do want to use it. You're the 10th person in this thread saying the same thing and it doesn't appear you even have that much experience with operations in general. Sorry to go off on you but I'm really seeing these types of tropes and quick depthless one liner comments and offtopic snipes lately as the downfall of the hn comment section.
- recursive 6y agoKubernetes doesn't benefit Google? How not?
- garadox 6y agoOne thing to watch for with pod antiAffinity - if you use required vs preferred, and your pod count exceeds the node count, the remainder will be left in Pending and won't spin up anywhere.
- dharmab 6y agoThere's a new feature which does a better job of spreading Pods without blocking scheduling quite as badly: https://kubernetes.io/docs/concepts/workloads/pods/pod-topology-spread-constraints/ https://kubernetes.io/docs/concepts/workloads/pods/pod-topol...
- whatsmyusername 6y agoTBH, for me it's usually "Using Kubernetes." Maybe on GCP (I don't see a lot of companies on GCP) it makes sense, but ECS is AWS native and on bare hardware I immediately go to docker swarm since it ships with the container runtime (instead of a bolted on sidecar container thing). I like the primitives and features of Kubernetes, but the implementation doesn't give me warm fuzzies and it always gets passed over for safer bets for me. Even very early on in it's development I always went to Mesos over Kubernetes (though Mesos is P dead at this point).
- yongjik 6y agoWell, nobody asked me, and I'm no expert, but here's my list of what (not) to do in Kubernetes (if I had the authority). 1. There. Is. No. Machine. (Insert matrix meme here.) Before you open up your cluster to the rest of company, drill it down to them. Maybe even create a Google Form where they have to sign "I hereby acknowledge that there is no machine in k8s and any attempt to tie my job to a particular machine means a broken config by definition." 2. Thanks to 1, don't let anyone use hostNetwork, hostIPC, hostPID, hostPorts, host whatever, unless you have a really good reason to (with explicit approval process). 3. Don't let anybody start a job without memory/CPU limit. Make sure they understand that, if the job goes over the memory limit, it dies, and it's not k8s admin's problem. 4. You can't log anything into the pod - when the pod dies the log is gone. You can't log into the machine, either (see 1). Therefore, you really need some kind of logging framework that takes the log from your pod and saves it, in its raw form, somewhere safe (like S3). I don't know if there's any such framework, but there had better be. 5. Make sure every manual operation is logged (who did what to which job when), unless you like asking "@here Does anybody know who owns fooservice?" every month. 6. Kubernetes is not magic: if it takes thirty minutes to provision your service, fix that, instead of moving thirty minutes of manual provisioning into k8s and somehow expect it to be magically reliable. 7. Don't bring in existing dependencies uncritically. If your job connects to a zookeeper server to find out its peers, don't bring it into k8s, but rewrite it to use k8s service instead. 8. Take extra extra care when writing down your first job specification, because there are a lot of yaml files to write, and people will just copy what's already there. If your first k8s job mounts host /tmp directory just because you were testing something and forgot to delete the line, soon you will have fifty jobs all mounting host /tmp directory. Good luck figuring out which job actually needs it then. Yeah, again, I'm by no means an expert - I'm not even an admin, so just consider the list as a rambling of some poor soul who has seen some stuff. Here be dragons, have fun.
- hvidgaard 6y ago> 8. Take extra extra care when writing down your first job specification, because there are a lot of yaml files to write, and people will just copy what's already there. If your first k8s job mounts host /tmp directory just because you were testing something and forgot to delete the line, soon you will have fifty jobs all mounting host /tmp directory. Good luck figuring out which job actually needs it then. This piece of advice is so underrated, it's hard to put into words. Every senior developer/architect should know that.
- triodan 6y agoThe guaranteed QoS example in the article is wrong. Kubernetes only sets the Guaranteed QoS if the CPU count is an integer (which 0.5 is not). Also, to take full benefit of the QoS you need to configure the Kubelet with "--cpu-manager-policy static"[0]. [0]: https://kubernetes.io/blog/2018/07/24/feature-highlight-cpu-manager/ https://kubernetes.io/blog/2018/07/24/feature-highlight-cpu-...
- marekaf 6y agoThanks to pointing that out! I will edit it.
- aganame 6y agoDon't use Kubernetes unless you Know you need it.
- musicale 6y agoOh I thought this was Common mistakes: using Kubernetes
- darkwater 6y agoReally nice article, it shows a lot of the small details you have to take into account to go from "deploying Kubernetes" to "deploy a production-grade Kubernetes" where production means some real, trafficked site.