3 ms·
How do you prevent someone (or something) from accidentally draining a node hosting a database that does not have a synchronised replica? Kubernetes is neat, b
by chousuke 5y ago
How do you prevent someone (or something) from accidentally draining a node hosting a database that does not have a synchronised replica?
Kubernetes is neat, but destroying stuff is so easy (and the context mechanism is just begging for human error) that I am a bit leery about hosting large amounts of stateful stuff in it.
- paulfurtado 5y agoNo database should lack an unsynchronized replica in general. A kubernetes admin could accidentally drain the host, but we run in AWS and AWS can kill nodes a lot faster than we can haha. We have enough servers in AWS that at least 5 fail per day. If some database is not replicated, someone is getting paged for it. We run MySQL with semi-sync replication for example: if a primary has two replicas, one replica must ack a write before the primary commits it, so you can lose the primary at any time and not lose data. That said, kubernetes has prevention mechanisms for this. PodDisruptionBudgets allow you to specify the minimum number of replicas. The drain command calls evict instead of delete and evict honors these restrictions. This is helpful for node draining, but not necessarily helpful if someone accidentally deletes a StatefulSet. However, nothing in kubernetes deletes PVC objects automatically, so if you delete the pod, the volume remains bound to the PVC until you delete it. preStop hooks are used in all the database pods to ensure that if the current pod is the primary, leadership is correctly changed to a replica before proceeding with the delete. We combine this with extremely long termination grace periods and alerting if termination is taking too long so that an ops team can look into the issue. preStop hooks are great, but they also are somewhat of a time bomb if not careful, so some of our database operators make use of finalizers instead. With finalizers, nothing happens when a pod enters the terminating state until an operator marks the pod's finalizer as completed. This flow is really nice because you end up with behavior where the kubernetes cluster admins or built in controllers effectively are really gently requesting that a pod terminate, and it's up to the operator of that pod to decide exactly how and when that occurs. For non-local volumes, like EBS PVCs, once the PVC is deleted an operator will snapshot the volume during a finalizer before allowing the volume to be fully deleted so that we can recover data if this ever accidentally occurs. This was the first protection we ever implemented when initially adopting kubernetes back when we had no idea what we were doing so that we could always recover data in the event of admin mistakes. We can't do this for local volumes, however, any team making use of local volumes truly needs to stay on top of replication when running in the cloud, whether on kubernetes or not or else you absolutely will lose data. AWS local nvme volumes have limited write cycles just like the ones in on-prem servers. So ideally all databses are configured to not accept writes that cannot be replicated. And even with EBS, at scale we end up with something like 20 unrecoverable volume failures per year
- chousuke 5y agoThanks for the detailed response. I was aware Kubernetes has hooks for pretty much everything, so getting details on how they are used in practice is interesting. Mostly I'm concerned about preventing administrator mistakes; if you have enough privileges, it's distressingly easy to just accidentally delete all kinds of resources in the default configuration, especially since kubectl is context-sensitive. I like my automation software to have checks in it to prevent me from doing stupid things without explicitly disabling a number of safeties first. I feel like Kubernetes is a powertool that requires a competent and trained admin team or managed access so that you simply can't make awful mistakes as a "regular user", but a lot of the hype around it seems to be focused on how "easy" it is. To my eyes, it's indeed easy to just throw manifests from the internet at Kubernetes, but that is often akin to disabling SELinux because some blog says so; it may get things working, but it's not the competent choice most of the time. I'm also a big proponent of storing everything in git repositories, so I do like how Kubernetes enables declarative configuration, though it seems much of the tooling in the ecosystem works against that by just creating resources in a cluster with no version control...
- paulfurtado 5y agoCool thing about finalizers is that there's no straightforward way to delete them from the CLI, you need to explicitly patch the object to delete the finalizer so that goes a long way towards preventing instant mass deletion, they really do help a lot here. The additional protection we have is via admission controllers which hook all API calls to prevent mistakes, of course, you have to foresee those mistakes and come up with a validation that prevents it. But even something as simple as "only 5 pods are allowed to be terminating at once" can help. I absolutely agree about people claiming kubernetes is easy. It is easy to throw some helm chart at the kube API and suddenly have an app running with all it's databses configured. It is hard to keep them all running. The pattern of operators seeks to solve this. Instead of a helm chart creating a MySQL statefulset, it creates a MySQLCluster that gets managed by a well-written operator following the best practices and correctly implementing all the hooks. The ecosystem is starting to converge on this finally and the beauty is that the sum of the industry's operational knowledge can be coded into these operators and we'll finally arrive at nearly fully self managing databases with backups and all. The industry isn't there quite yet though, and so at least for stateful workloads you need significant operational experience. For background: I am the Tech Lead of our team that runs the kubernetes clusters and we provide all of the primitives our users need and control access. Each database type then has dedicated teams operating them and codifying operations into operators. There is a lot of work being done. But prior to kubernetes, all of these teams separately were coding against the EC2 API with no standardization, no unified view, ad-hoc failure detection, ad-hoc auditing, etc. Kubernetes is a very substantial improvement over that, but 90% of companies never reach a scale where this is necessary. But once extremely solid open source operators exist we may truly hit the dream of just applying some self managing operator manifest to an EKS or GKE cluster and getting an actually production grade database setup. Version control is an interesting topic because it's hard to express transitive dependencies in that code. For example, the MySQL operator creating a StatefulSet and the Statefulset creating pods. It doesn't make sense to commit those lower level resources to git, they're not fully independent. However, for top level things, our build system produces artifacts from git and the deploy system creates them in the cluster. With this setup, only admins could directly apply them, which really helps with keeping things in sync with git.