10 ms·
Self-hosting a high-availability Postgres cluster on Kubernetes
- mailcheap 3y agoWhy does the author not use a native HA database such as Cassandra or ScyllaDB?
- andrewmunsell 3y agoMy holiday project was doing another pass at my Homelab Kubernetes cluster, part of which involved switching to a proper operator to manage Postgres. Coincidentally, I setup cloudnative-pg (https://github.com/cloudnative-pg/cloudnative-pg https://github.com/cloudnative-pg/cloudnative-pg) yesterday.
- x86hacker1010 3y agoAny reason you landed on that Operator compared to what OP is using (Zalando)?
- andrewmunsell 3y agoHonestly no, it's mostly due to inexperience with operators and not really understanding what the "best" way to find operators is. I did also look at the Crunch Data one (I was having some issues setting that one up), but didn't even find Zalando during my search. OperatorHub is currently the main resource I use, but GitHub stars aren't exposed in the search so I have been looking at the "Capability Level" chart and checking for Github popularity when I find one with the feature support I want. I'm facing this exact same issue now when trying to find an operator for Redis. I am not sure if I am just missing out on the "right" option by limiting myself to Googling and Operator Hub and looking for the one with the most Github stars, so I am open to tips.
- turtles3 3y agoA subtle advantage of cnpg is that it doesn't use statefulsets, instead the operator handles things like mapping storage volumes and stable identities. Regular kubernetes statefulsets have some tricky sharp edges for failure recovery. I don't know if all of these alternatives use statefulsets but I remember several doing so. I've personally found cnpg to be pretty robust, and supports everything you will eventually need once you're locked into a solution (eg. Robust backups, CDC, replica clusters). I'm yet to find anything of a similar standard for mysql. EDIT: it should also be noted that CrunchyData is a proprietary solution and requires a license to use in production. This is not particularly obvious from their docs.
- activescott 3y agoWhat sharp edges are you referring to with statefulsets?
- turtles3 3y agoCnpg's docs articulate this better than I could: https://cloudnative-pg.io/documentation/1.16/controller/ https://cloudnative-pg.io/documentation/1.16/controller/ Statefulsets have their place but are surprisingly inconvenient for database workloads.
- Szpadel 3y agoI was setting fairly important database with Zalando pg operator and after first good impressions it went downhill. after like a month of use WAL files used for point in time recovery started failing to offload to dedicated nodes and kept growing on database pods filling up all the space. I firstly assumed that maybe there is not enough space for some scheduled work (I do not really know details how this process work, I assumed that operator should handle all implementation details for me) but even after upscaling database 2.5x it just kept failing with full storage and requiring manual recovery to bigger storage, where most of it was WAL files. HA didn't handled this case at all whole cluster went in crash loop there was also issue of huge pages caused crashing and not easy way to disable those without some dirty injecting of config files at runtime there could be some my fault at misconfiguration on by side, but I wasn't able to figure anything better from docs
- ahachete 3y agoI'm the founder of OnGres [1] the company behind StackGres [2]. I'd love to hear your feedback if you'd be interested in also trying StackGres. It's one of the most feature-full operators available, has a complete Web Console and REST API and supports close to 200 extensions. Hope it would be interesting for you. [1]: https://ongres.com https://ongres.com [2]: https://stackgres.io https://stackgres.io
- bo0tzz 3y agoI've been using CNPG on my home cluster since it came out, and it's been an absolute pleasure to use. I haven't done a full comparison, but I get the sense that it's learned from (and improved on) the other postgres operators like Zalando and Crunchy.
- peterbecich 3y agoI also recommend the Kubegres operator: https://www.kubegres.io/ https://www.kubegres.io/
- gchamonlive 3y agoJust like with cloud providers, it seems to me like with kubernetes, it is not a matter of if but when orchestration problems will arise. Specially in this case of hosting databases, a composition of provisioning complexities (db operational complexities on top of k8s operational complexities) is really scary. Is there any way to overcome hidden complexity biting your hand other than studying k8s extensively?
- renegade-otter 3y agoThis kind of resilience is a form of art, and it's also kind of a full-time job. I would not advise trying this for a "side project at work". Generally we all agree that we move to the cloud and it's "fully managed". If it goes down - that's the price we pay. If you have to ask the question "what if the RDS goes down", then you are really in a different universe. That last guarantee of uptime requires a ton of work, testing, and money, because you are all the way up and to the right on the curve of diminishing returns.
- candiddevmike 3y ago> That last guarantee of uptime requires a ton of work, testing, and money, because you are all the way up and to the right on the curve of diminishing returns. Not when it's a core part of your business...?
- williamdclt 3y ago> If you have to ask the question "what if the RDS goes down", then you are really in a different universe. It does go down though, don’t neglect the possibility because it likely will happen. With very average workloads, I’ve seen RDS databases restart unexpectedly, read replicas being completely out of service, and even databases being completely frozen (can’t even connect as root). I’d still go with managed, but it certainly doesn’t give full reliability :) you still have to consider “what if it goes down” - it will!
- debarshri 3y agoProblem is that RDS comes at a price. It is purely about operation cost. When you have 1500+ databases these cost add up. At that point, this kind of techniques are required to self host the databases. Price per DB with HPA na VPA is way lower than what you would pay for managed databases as well as you can hire a full time devops+dbadmin and still be cheaper.
- ysofunny 3y agoonce upon a time I set up an elastic search cluster in kubernetes after a lot of tweaking I made it so that the pods would be as big as the underlying hardware nodes. one pod one node. once that was working I realized that I was using the wrong tool for the job. the kubernetes tooling added nothing but complexity. needless to say I let it run like that having had wasted about a week getting it to work
- jhgg 3y agoOn the other hand, at a certain scale (running hundreds of ES nodes across 80 or so ES clusters), Kubernetes actually does make a lot of sense. At work, we moved from hosting elastic search on bare VMs to kubernetes. By leveraging scheduler policies we are able to pack / over-provision ES node pods of different clusters onto the same Kubernetes nodes, allowing for far greater resource efficiency, while being able to handle node failure while maintaining availability across all clusters. Additionally, this simplified operations significantly as we can now leverage the operator to do cluster wide operations (e.g rolling restarts, node OS upgrades, ES version upgrades, etc...) fairly easily. We did, however, go 6 years (and several hundred million users and trillions of documents indexed) without needing to use Kubernetes! We will blog about this at some point this year.
- jen20 3y agoA big part of the problem with Kubernetes is it doesn't make a ton of sense at small scale, and it just plain doesn't work at large scale. Nomad is generally speaking a much more appropriate technology when you hit the point of needing such a system.
- jhgg 3y agoI would consider our scale pretty large here, and it works just fine.
- marcosdumay 3y agoDid you keep adding and removing replicas into your cluster based on a scheduler policy? How often did you adjust the number of nodes (and how long did it take to make a node available)? At the high-level you are describing your setup, it doesn't make sense. You'd spend way more resources managing any cluster than what you would gain from a normal-looking policy. I seem to be missing some important detail.
- ssijak 3y agoIs Kubernetes still hard in 2024?
- _joel 3y agoI don't know about hard but it's fairly straightforward to spin up clusters and maintenance seems to have become less of a headache (at least imhe). It depends on what you will be doing with the cluster and how you use it.
- azlev 3y agoYes. Orchestration is not easy.
- jamesu 3y agoFrom recent experience I'd say it's the sort of tech that starts off simple enough with the right distribution, but then gets more complicated the deeper down you dive. Probably the most hard thing I found was wrapping my head around the way the storage works.
- danielvaughn 3y agoGranted I'm new to devops, only been on a platform team for about 9 months now, but I still feel incredibly dumb every time I try to work with it. That being said, we're also layering a bunch of stuff on top - helm, nginx, GKE, terraform, as well as a mountain of other things, and then to top it off we have a bunch of shell scripts doing random things to help tie it all together. Normally I can pick things up pretty quickly. I just built a parser with tree-sitter, despite knowing virtually nothing about language design. Didn't take very long. But the modern devops stack is a learning curve like I've never seen before. It's taking me more energy to learn it than it did for me to learn programming itself. Then again, maybe I'm just getting old.
- cjaro 3y agoI appreciate your shared experience as I have a rather similar one. I've been on a platform SRE team for about eight months myself, with seven years of SWE experience before that, and feel as though I'm just able keep my head above water. It's one thing to learn about k8s, the cloud, terraform, etc, then quite another to pile it all together, particularly since it all becomes heavily customized. It's a different job than code-writing software engineering, that's for sure. To me, it feels less like there's a 'stack' so much as there's a word cloud of DevOps buzzwords to start throwing at problems. Even when directed by architects & principals, it overwhelms. I resonate with feeling incredibly dumb whenever I pick up a new ticket from our backlog. It feels like gaining deep knowledge of these systems will be a nearly insurmountable challenge. It's been eight months, and while I know far, far more than I did on day one, I feel that every day is a day one of sorts.
- justanotheratom 3y agodumb question - where is the storage kept?
- iamgopal 3y agoMe who have never used kubernetes, what if node crash ? Will I lost everything ?
- stanac 3y agoNo, attached storage is not part of the node (not directly). It's something like attaching external (host) directory to a docker container. Your can kill the node/pod and storage is not affected, later you can attach new pod to the same storage.
- rad_gruchalski 3y agoThe answer is: it depends. It depends on if you use persistent volumes, and how well is your pv isolated from the failed node. If done right, no data loss.
- marcosdumay 3y agoA big emphasis to the "if done right" part. You should test your setup, because it's very often not done right, and it's easy to overlook a problem.
- siamese_puff 3y agoI think people over think how hard this actually is. Data replication isn't a new concept. You can use RAID with a NAS or setup async replication with operators for things like Postgres/SQL. Obviously it's worth doing simulated disaster recovery to ensure you would recover if there is hardware failure. The larger the scale and throughput with parallel writes against the same keys, etc then the more complicated the setup will be. I hope to write more on this topic, but setting up a persistent volume with a NAS is a great way to ensure high durability.
- bo0tzz 3y ago
- xenic 3y ago”Zalando is a Postgres operator that facilitates the deployment of a highly available (HA) Postgres cluster.” Zalando is the company. ”Postgres Operator” is the software. Happy user here, not much complaints about the operator come to mind.
- siliconc0w 3y agoI wonder why more don't take advantage of native k8s and just rely on it to move over the persistent volume and start the new pod. This may have some small amount of downtime but it's a lot less complicated.
- siamese_puff 3y agoThat gets complicated when you're running an HA cluster and need to worry about write conflicts.
- oxfordmale 3y agoNo, just no. K8s shouldn't be used to host database systems. Its main function is micro services. Cloud provider provide managed versions of Postgres that are highly available. Even if you self host, Kubernetes isn't the answer.
- dilyevsky 3y agoPerfectly fit for running databases. HA databases that have built in replication (cockroach, tidb, foundationdb, etc) are easier, pg is just not built for proper clustering. Still doable tho
- Havoc 3y agoI wonder how much of big cloud one can replicate on a diy cluster. Database works via this. S3 via minio. Redis. And maybe openFAAS? That seems like a lot of the key building blocks already.
- siamese_puff 3y agoI'm biased, but I agree! This is a great article that further illustrates how we are a bit conditioned to use off-the-shelf vendor tooling that we can run ourselves. https://kiwiziti.com/~matt/wireguard/ https://kiwiziti.com/~matt/wireguard/ Of course, there are tradeoffs you have to make (security, uptime, criticality of the system), but homelabs exist as perfect experimentation frameworks. At $DAYJOB we run a global scale Ceph cluster (S3-like API/object storage), so even at a large scale it's not impossible to imagine.
- yeseniajimenez1 3y ago[dead]