3 ms·
Why shouldn't you run databases on it? Ive been running larger databases(10tb+) on it for some time. If your database can deal with failure of single node and
by eicnix 9y ago
Why shouldn't you run databases on it? Ive been running larger databases(10tb+) on it for some time.
If your database can deal with failure of single node and uses network storage it's totally doable to run databases on Kubernetes.
Take a look at Patroni on how to run Postgres on IaaS or Kubernetes.
- eropple 9y agoYou shouldn't run databases on it unless you can express exactly why you should. Me, I already have resource segmentation throughout my infrastructure. They have their own network stacks and predictable, understandable resource utilization. I call those segments "virtual machines." It's rather harder to blog about an RDS instance (or even a self-hosted EC2 one), but you minimize moving parts and the way your bleeding-edge systems can screw you.
- eicnix 9y agoMy main reason to run basically everything in Kubernetes is abstraction between resources and applications and standardization of deployments. While you can achieve the same with various other ways like good configuration management and virtual machines I have found that the overall cost of creating/running and monitoring deployments on Kubernetes is smaller than on a traditional VM based solutions.
- cookiecaper 9y agoDatabase servers are generally designed from the ground-up to utilize basically the entire machine and to be left running for a long time, with caching and memory utilization techniques designed around these assumptions. VMs muck with this somewhat but at least the memory region is reserved for that VM's usage. Docker eliminates that, and k8s eliminates the knowledge even of what type of hardware/memory the system is being executed on (without taking several contrived steps to restore this). Database servers are also designed on the assumption that each node will be available on a consistent basis, at a consistent address, and that the data directory will be available to the server as soon as it starts up (except for rare node bootstrap operations). This is the exact opposite of what Kubernetes/Docker seek to provide, and in fact, such things can only be provided within Kubernetes by extensive special configuration that leans heavily on experimental features. There could not be a worse ideological match. Yes, you can try really hard to ram that square peg through that round hole, eventually pushing it through with significant damage to both the peg and the hole, but why would you?
- xkarga00 9y ago> This is the exact opposite of what Kubernetes seek to provide Not true. You can have dedicated machines for your db instances, check out 1. node affinity 2. Statefulsets. Statefulsets are designed to provide consistency (no "split brain"), stable network identity ("at a consistent address"), stable storage identity ("and that the data directory will be available to the server as soon as it starts up"). Those features are beta but it's a matter of time (hardening and the likes) before we tut them stable. > why would you? For all the same reasons you would run any other workload on the same platform.
- cookiecaper 9y agoThe reason to run any other workload on Kubernetes is to dynamically schedule your containers across a cluster of anonymous hardware resources, and to provide automatic monitoring, recovery, and control of those containers when certain events occur. The goal, essentially, is to abstract the hardware/system-level element from the deployment element. That's all well and good, but applications that make certain architectural assumptions do not lend themselves well to random scheduling across an anonymous array of system resources. Databases are absolutely among that class of applications, as are many other application types (which are now called "stateful" applications). For k8s to work well for an application, that application has to be anonymous, without any masters or controllers. It has to be able to tolerate the sudden vaporization of any member. It has to be willing and able to share host hardware with any number of other services, including some which may be bad neighbors/CPU hogs, and it has to be content to be rescheduled in the event that a node goes down, that a pod is killed to ease the transition, etc. Databases fail on virtually every point of this. If the database and Kubernetes start from opposite design paradigms, why does it make sense to run a database inside Kubernetes? I am still not getting it. `kubectl delete pod/my-postgres-pod` is not a smart thing; you don't want any of that scheduling magic that Kubernetes provides. At most you would want Kubernetes to tell you that your database is failing health checks and execute the STONITH process to failover to replica, but you hardly need Kubernetes if all you care about is process monitoring. So can you elaborate on the reasons? Databases are simply not designed for this kind of infrastructure and I see no value in trying to pretend they are. Isn't this the reason that CockroachDB exists, so that people can finally run their DBs in something like Kubernetes without endless headaches? I think it would be very interesting to analyze the stability and performance features of a PgSQL k8s deployment, PgSQL VM deployment, and PgSQL bare metal deployment. The only issue is that you can't expect the failures to be open with their data.
- otterley 9y agoAre you actually hosting customer data in this system? How are you ensuring that the scheduler never shuts down or relocates the database server without your explicit consent? And when the scheduler does relocate the service, how do you ensure clients can connect to the new instance without interruption?