17 ms·
The Horrors of Upgrading Etcd Beneath Kubernetes
- akeck 8y agoTo sidestep upgrade issues, we're pursuing stateless immutable K8S clusters as much as possible. If we need new K8S, etcd, etc., we'll spin a new cluster and move the apps. Data at rest (prod DBs, prod Posix FS, prod object stores, etc.) is outside the clusters.
- SteveNuts 8y agoWhere do you run your persistent apps?
- brianwawok 8y agoNot the OP, but I have two kinds of persistent data. 1) Images / files / etc. It all lives in cloud storage ("s3"), outside of K8s 2) RDBMS data. You can just run as hosted sql (say CloudSQL) or a not-in-k8s VM. I have found no compelling reason to move my RDBMS into my k8s cluster.
- zbentley 8y agoThat's a bit distressing. Most everywhere I've worked, the infrastructure that matters to the business has fallen into roughly two categories: Category 1: stateless-"ish" workloads. More than 90% of hosts/containers used . . . less than 25% of operations headaches and time. Issues that happen here are solvable with narrow solutions: add caches, scale out, do very targeted, transparent fixes to poorly-performing application code. Category 2: stateful workloads. Less than 10% of hosts/containers. 75% or more of operations headaches and time. Issues that happen here have less visibility, fewer short-term fixes ("just add an index and turn off the bad queries" only works so many times before you're out of low-hanging fruit), and require more expertise to solve in a way that doesn't require the application/clients to change. If k8s and other next-gen technologies are only easing the first category, that makes me sad. It's like we have this sedan (off-the-shelf web technologies) that we have to take off-roading and it falls apart all the time. I don't want a better air conditioning system and more cushions in my seats; I want the vehicle to not break.
- wereHamster 8y agoIn k8s you can easily host stateful services. You have persistent disks that you can attach to containers, and you also have StatefulSets if you have a stateful service that you want to have automatically scaled (https://kubernetes.io/docs/concepts/workloads/controllers/statefulset/ https://kubernetes.io/docs/concepts/workloads/controllers/st...). You can use both to run a database (postgres) for example.
- zaat 8y agoI thought that this was the case too, but OP had a link (https://gravitational.com/blog/running-postgresql-on-kubernetes/ https://gravitational.com/blog/running-postgresql-on-kuberne...) to previous post that got me worried.
- outworlder 8y agoFrom TFA > Kubernetes is not aware of the deployment details of Postgres. A naive deployment could lead to complete data loss. That sounds ominous, but is actually a tautology. You have the exact same challenges anywhere else, but since K8s makes some operations so easy to do, you need to be careful. RDBMs are specially tricky because most of them expect a single "master" which holds special status. And it so happens to hold all your data too (as do your replicas, provided they are up to date).
- thraxil 8y agoK8s and similar represent a current view on how systems should be designed and run from the ground up (largely based on the same observations you have made that the stateless workloads can be run with drastically less operational problems). If your architecture lines up with this approach (eg, following the 12-factor approach), there really are a lot of advantages. But no, legacy architected applications are just not going to benefit in the same way. I assume that there must be consultants and "thought leaders" out there who are pushing k8s and containerization as silver bullets and that's unfortunate.
- mdekkers 8y ago
- bonesss 8y ago> ... not-in-k8s VM One of the cooler innovations we've seen, and I think we're going to see more of, is the ability to take a non-k8s VM and expose it to the cluster as-if it were just another pod. This would let you schedule and expose RDBMs and other specialized servers through kubernetes while keeping them on a standard VM. I think that's the win/win approach to bridge the gap.
- outworlder 8y agoIndeed. But I think that doesn't go far enough: I want to provide K8s with a YAML file and have Kubernetes itself go and create and provision a VM for me. I.e, I want to "kill" Terraform. I don't care about a 'fabric' network (although that could be convenient), just give me an IP that I can reach even it if is external to the cluster. VMs could be just one more resource that can be created or destroyed by k8s, just like network-based storage today. I haven't found this exact use-case implemented yet. If noone else builds it in the coming months I'll probably start doing it.
- smarterclayton 8y agoWe did something simple for GCP and CI jobs, when we needed to have VMs in a certain project to launch Kubernetes single node testing - https://github.com/openshift/ci-vm-operator/ https://github.com/openshift/ci-vm-operator/ The base pattern should be pretty easy to modify, although the use case here is very specific.
- hunter_n 8y agoHave you looked at Virtlet or Kubevirt?
- bjslade 8y agoThis is the approach we’ve taken with other similar systems too. Things can go so horribly wrong underneath your containers that the concept of having only one cluster in a prod scenario and maintaining it mid-air would be unthinkable. Also agree on keeping state outside. Maybe the relevant tech will be mature sometime soon but we’ve seen orchestration bugs do nasty things to stateless containers that would have been a nightmare if state had been involved.
- philips 8y agoHey thanks for the article. I know etcd upgrades can look complex but upgrading distributed databases live is always going to be quite non-trivial. That said for many people taking some downtime on their Kube API server isn't the end of the world. The system, by design, can work OK for sometime in a degraded state: workloads keep running. A few things that I do want to try to clarify: 1) The strict documented upgrade path for etcd is because testing matrixes just get too complicated. There aren't really technical limitations as much as wanting to ensure recommendations are made based on things that have been tested. The documentation is all here: https://github.com/coreos/etcd/tree/master/Documentation/upgrades https://github.com/coreos/etcd/tree/master/Documentation/upg... 2) Live etcd v2 API -> etcd v3 API migration for Kubernetes was never a priority for the etcd team at CoreOS because we never shipped a supported product that used Kube + etcd v2. Several community members volunteered to make it better but it never really came together. We feel bad about the mess but it is a consequence of not having that itch to scratch as they say. 3) Several contributors, notably Joe Betz of Google, have been working to keep older minor versions of etcd patched. For example 3.1.17 was released 30 days ago and the first release of 3.1 was 1.5 years ago. These longer lived branches intend to be bug fix only.
- kevin_nisbet 8y agoAwesome, thanks for the clarifying items. When going through this, I did test it and see that the upgrade appeared to work, until I think it was 3.3 where it would panic, but didn't want to rely on the undefined / untested behaviour, even if it seemed to work in a lab. The interview was long, so as we were cutting it down I think this aspect got lost. And thanks for the hard work on etcd.
- geertj 8y ago> upgrading distributed databases live is always going to be quite non-trivial. I understand that the implementation of live upgrades for a distributed database will be complex but this post is about the user experience. Given enough resources, is there a reason that it can't be a single "upgrade now" command? Or maybe slightly more real-world, a 3 step process like: "stage update" -> "test update" -> "start update".
- peterwwillis 8y agoI don't know about you, but my application is tested on a single platform/stack with a specific set of operations. When the operation of the thing I'm running on changes, my application has changed. It just can't be expected to run the same way. Upgrade means your app is going to work differently. Not only is the app now different, but the upgrade itself is going to be dangerous. The idea that you can just "upgrade a running cluster" is a bit like saying you can "perform maintenance on a rolling car". It is physically possible. It is also a terrible idea. You can do some maintenance while the car is running. Mainly things that are inside the car that don't affect its operation or safety. But if you want to make significant changes, you should probably stop the thing in order to make the change. If you're in the 1994 film Speed and you literally can't stop the vehicle, you do the next best thing: get another bus running alongside the first bus and move people over. Just, uh, be careful of flat tires. (https://www.youtube.com/watch?v=rxQI2vBCDHo https://www.youtube.com/watch?v=rxQI2vBCDHo)
- viraptor 8y agoIt really depends on what your scenario/usage is. Sometimes it's a terrible idea: you wouldn't do a live code update on a server handling your blog. Sometimes it's just what you're aiming for: you have expensive hardware plugged into physical lines, there's not enough capacity to migrate data flow uninterrupted, you have to update in place without clients having more than X ms delay. Real life is virtually always the first case... But if you really need it, tech like Java hotswap or elang hot code swap is there.
- taeric 8y agoPicking a dependency for a system that does not have a live method of updating, is a terrible idea for a software system. Physical systems, like a car, have obvious limitations on what can be modified when. Similarly, software will have some limitations on what happens when you are updating. But accepting "upgrades can't be done easily" for software is putting much more limitations on the software than makes sense.
- peterwwillis 8y agoWith the exception of like, binary patching of executables that use versioned symbols or some craziness like that, virtually all software cannot be upgraded while it is running and expected not to produce errors. I mean, if you use a plug-in style system, you can program it to block operations while a module is reloaded or something. But most software is not designed this way. Especially with ancient monolithic models like Go programs. Upgrades just can't be done easily in a complex system. You can do them without concern for their consequences, but that doesn't mean they're safe or reliable methods.
- segmondy 8y agoI'll have to read this later this weekend, my home k8s cluster that broke did so because of etcd. Grrr
- djb_hackernews 8y agoThe clustering story for etcd is pretty lacking in general. The discovery mechanisms are not built for cattle type infrastructure or public clouds. ie it is difficult to bootstrap a cluster on a public cloud without first knowing the network interfaces your nodes will have or it requires you to already have an etcd cluster OR use SRV records. From my experience etcd makes it hard to use auto scaling groups for healing and rolling updates. From my experience consul seems to have a better clustering story but I'd be curious why etcd won out over other technologies as the k8s datastore of choice.
- nvarsj 8y ago> From my experience consul seems to have a better clustering story but I'd be curious why etcd won out over other technologies as the k8s datastore of choice. That'd be some interesting history. That choice had a big impact in making etcd relevant, I think. As far as I know, etcd was chosen before kubernetes ever went public, pre-2014? So it must have been really bleeding edge at the time. I don't think consul was even out then - it might have been they were just too late to the game. The only other reasonable option was probably ZooKeeper.
- robszumski 8y agoI was around at CoreOS before Kubernetes existed. I don't recall exactly when etcd was chosen at the data store, but the Google team valued focus for this very important part of the system. etcd didn't have an embedded DNS server, etc. Of course, these things can be built on top of etcd easily. Upstream has taken advantage of this by swapping the DNS server used in Kubernetes twice, IIRC. Contrast this with Consul which contains a DNS server and is now moving into service mesh territory. This isn't a fault of Consul at all, just a desire to be a full solution vs a building block.
- otterley 8y agoMy understanding is that Google valued the fact that etcd was willing to support gRPC and Consul wasn't -- i.e., raw performance/latency was the gating factor. etcd was historically far less stable and less well documented than Consul, even though Consul had more functionality. etcd may have caught up in the last couple years, though.
- mirceal 8y agoThis may be an unpopular opinion, but I’m not a big fan of containers and K8S. If your app needs a container to run properly, it’s already a mess. While what K8s has done for containers is freaking impressive, to me it does not make a lot of sense unless you run your own bare metal servers. Even then, the complexity it adds may not be worth it. Did I mention that the tech is not mature enough to just run on autopilot and now instead of worrying about the “devops” for your app/service you are playing catch-up with upgrading your K8s cluster? If you’re in the cloud, VMs + autoscalling or fully managed services (eg S3, lambda, etc) make more sense and allow you to focus on your app. Yes there is lock-in. Yes, if not properly arhitected it can be a mess. I wish we would live in a world where people pick simple over complex and think long term vs chasing the latest hotness.
- shoo 8y ago> If you’re in the cloud, VMs + autoscalling or fully managed services (eg S3, lambda, etc) make more sense and allow you to focus on your app. Yes there is lock-in. Yes, if not properly arhitected it can be a mess. I've just rolled off a project (line of business web app + cluster of workers for background job processing), this more or less describes what we have running. Some of the architecture is a bit wrong (it wasn't designed for AWS but shifted there after it had been running for a year or so) but the system works well enough to deliver value to the business. As a developer I hate state, things that aren't properly isolated, ill-defined system boundaries -- but it's not obvious to me what the business case would be to containerise everything.
- threeseed 8y ago> it's not obvious to me what the business case would be to containerise everything Containers allow you to move apps trivially between environments and guarantee that they will just work. It allows you to isolate dependencies between apps e.g. Python 2 versus Python 3. It allows you to move apps between cloud providers or between on premise and cloud. With platforms like Kubernetes it allows you to easily scale and self heal when nodes die. And compared to rewriting your app in Lambda which is expensive and complex it is simple to build a container as almost every language has automated tooling.
- jacques_chester 8y agoEtcd misbehaving during upgrades or when a VM was replaced was a massive source of bugs for Cloud Foundry. There is no longer an etcd anywhere in Cloud Foundry.
- ec109685 8y agoAren’t they introducing Kubernetes as part of Cloud Foundry: https://techcrunch.com/2018/04/20/kubernetes-and-cloud-foundry-grow-closer/ https://techcrunch.com/2018/04/20/kubernetes-and-cloud-found...
- jacques_chester 8y agoSort of. Kinda. It's complicated. But insofar as Kubernetes becomes the container orchestrator, I imagine we will encounter some or all of those problems again. Cloud Foundry's current orchestrator, Diego, is of similar vintage to Kubernetes. It now relies entirely on relational databases for tracking cluster state. Ditto other subsystems (eg, Loggregator). It scales just fine. MySQL, while not my personal favourite, has proved more reliable in practice than etcd. Some folks use PostgreSQL. Also more reliable in practice. Paying customers care more about being more reliable in practice than being more reliable in theory. I've ha-ha-only-seriously suggested we throw engineering support behind non-etcd cluster state. For example: https://github.com/rancher/k8s-sql https://github.com/rancher/k8s-sql
- a2tech 8y agoWhat problems exactly is etcd trying to solve?
- aberoham 8y agoInterviewer here -- "If you want to form a multi-node cluster, you pass it a little bit of configuration and off you go." .. "Raft [etcd] gets you past this old single-node database mantra, which isn’t really particularly old or even a wrong way to go. Raft gives you a new generation of systems where, if you have any type of hardware problem, any type of software instability problem, memory leaks, crashes, etc., another node is ready to take over that workload because you are reliably replicating your entire state to multiple servers."
- outworlder 8y agoDistributed storage + consensus, basically. For the purposes of this thread, etcd is the underlying Kubernetes storage mechanism. For many practical purposes, that's all one needs to know, unless you are in charge of maintaining the ETCD cluster.
- fen4o 8y agoFrom my experience, running etcd in cluster mode simply creates too many problems. It can scale vertically very well and if you run etcd (and other Kubernetes control plane components) on top of Kubernetes you can get away with running only a single instance.
- williesleg 8y agoNo kidding. We actually totally abandoned Kubernetes because of a botched upgrade. Too many moving parts, not a lot of coordination. Maybe we'll try again in a few years, maybe not.
- outworlder 8y agoThis article really hits home. A K8s cluster can survive just about anything. Worker nodes destroyed, meh, scheduler will take care of bringing stuff up. Master nodes destroyed. Meh. It doesn't care. ETCD issues though? Prepare for a whole lot of pain. They are very uncommon though. Upgrading is the most frequent operation.
- king_nothing 8y agoLol. CoreOS and Hashicorp products often throw “cloud” and “discoverability” around but lack crucial features for ops supportability found in solutions that came before. Zookeeper, Cassandra, Couchbase didn’t evolve in a development vacuum chamber. New != better.
- amorousf00p 8y agoI'm old school. I look at containers as jails and all the work to isolate applications in containers as of indifferent value given a flat plane process scope with MAC and application resource controls in well designed applications. That is I default to good design and testing rather than boilerplate orchestration and external control planes. All containers have done (popularly) in my opinion is add complexity and insecurity to the OS environment and encouraged bad behavior in terms of software development and systems administration.