5 ms·
Who is deploying databases in containers?
by danappelxx 3y ago
Who is deploying databases in containers?
- huahaiy 3y agoEmbedded DB
- orbz 3y agoA disturbingly large number of deployments I’ve seen using Kubernetes or docker compose have databases deployed as such.
- danappelxx 3y agoIMO if you’re concerned about performance and yet are deploying databases this way — mmap should not even be on the radar.
- charcircuit 3y agoHow would containers even hurt performance? How does the database no longer having the ability to see other processes on the machine somehow make it slower?
- danappelxx 3y agoI’ll assume the worst case: - lots of containers running on a single host - containers are each isolated in a VM (aka virtualized) - workloads are not homogenous and change often (your neighbor today may not be your neighbor tomorrow) I believe these are fair assumptions if you’re running on generic infrastructure with kubernetes. In this setup, my concerns are pretty much noisy neighbors + throttling. You may get latency spikes out of nowhere and the cause could be any of: - your neighbor is hogging IO (disk or network) - your database spawned too many threads and got throttled by CFS - CFS scheduled your DBs threads on a different CPU and you lost your cache lines In short, the DB does not have stable, predictable performance, which are exactly the characteristics you want it to have. If you ran the DB on a dedicated host you avoid this whole suite of issues. You can alleviate most of this if you make sure the DB’s container gets the entire host’s resources and doesn’t have neighbors.
- gcoakes 3y ago> - containers are each isolated in a VM (aka virtualized) Why are you assuming containers are virtualized? Is there some container runtime that does that as an added security measure? I thought they all use namespaces on Linux.
- danappelxx 3y agoIt’s becoming standard as a security measure. See: Kata containers, Firecracker VM
- otterley 3y agoNot so; neither Kata containers nor Firecracker are in widespread public use today. (Source: I work for AWS and consult regularly with container services customers, who both use AWS and run on premise.)
- danappelxx 3y agoAh, good to know!
- charcircuit 3y agoNone of those are the fault of containers. You can do all of what you said without containers.
- crabbone 3y agoThere are many "holes" in these containers. 1. fsync. You cannot "divide" it between containers. Whoever does it, stalls I/O for everyone else. 2. Context switches. Unless you do a lot of configurations outside of container runtime, you cannot ensure exclusive access to the number of CPU cores you need. 3. Networking has the same problem. You would either have to dedicate a whole NIC or SRI-OV-style virtual NIC to your database server. Otherwise just the amount of chatter that goes on through the control plane of something like Kubernetes will be a noticeable disadvantage. Again, containers don't help here, they only get in the way as to get that kind of exclusive network access you need more configuration on the host, and, possible an CNI to deal with it. 4. kubelet is not optimized to get out of your way. It needs a lot of resources and may spike, hindering or outright stalling database process. 5. Kubernetes sucks at managing memory-intensive processes. It doesn't work (well or at all) with swap (which, again, cannot be properly divided between containers). It doesn't integrate well with OOM killer (it cannot replace it, so any configurations you make inside Kubernetes are kind of irrelevant, because system's OOM killer will do how it pleases, ignoring Kubernetes). --- Bottom line... Kubernetes is lame from infrastructure perspective. It's written for Web developers. To make things appear simpler for them, while sacrificing a lot of resources and hiding a lot of actual complexity... which is impossible to hide, and which, in an even of failure will come to bite you. You don't want that kind of program near your database.
- FridgeSeal 3y agoAs these are obviously very real issues, and Kubernetes also isn’t going away imminently, how many of these can be fixed/improved with different design on the application front? Would using direct-Io API’s fix most of the fsync issues? If workloads pin their stuff to specific cores can we incite some of the overhead here? (Assuming we’re only running a single dedicated workload + kubelet on the node). > You would either have to dedicate a whole NIC or SRI-OV-style virtual NIC to your database server Tbh I’ve no idea we could do this with commodity cloud servers, nor do I know how, but I’m terribly interested in knowing how, do you know if there’s like a “dummy’s guide to better networking”? Haha > kubelet is not optimized to get out of your way...Kubernetes sucks at managing memory-intensive processes Definitely agree on both these issues, I’ve blown up the kubelet by overallocating memory before, which basically borked the node until some watchdog process kicked in. Sounds like the better solution here is a kubelet rebuilt to operate more efficiently and more predictably? Is the solution a db-optimised kubelet/K8s?
- spockz 3y agoGiven the ability to deploy pods to dedicated nodes based on label selectors, what is the actual performance impact of running a database in a container on a bare metal host with mounted volume versus running that same process with say systemd on that same node? Basically, shouldn’t the overhead of running a container be minimal?
- crabbone 3y agoThe problem is kubelet likes to spike in memory / CPU / network usage. It's not a well-behaved program to put alongside a database. It's not written with an eye for resource utilization. Also, it brings nothing of value to the table, but requires a lot of dance around it to keep it going. I.e. if you are a decent DBA, you don't have a problem setting up a node to run your database of choice, you would be probably opposed to using pre-packaged Docker images anyways. Also, Kubernetes sucks at managing storage... basically, it doesn't offer anything that'd be useful to a DBA. Things that might be useful come as CSI... and, obviously, it's better / easier to not use a CSI, but to interface directly with the storage you want instead. That's not to say that storage products don't offer these CSI... so, a legitimate question would be why would anyone do that? -- and the answer is -- not because it's useful, but because a lot of people think they need / want it. Instead of fighting stupidity, why not make an extra buck?
- FridgeSeal 3y agoI run DB’s on K8s, not because I don’t know what I’m doing, but because most of the trade offs are worth it. If I run a db workload in K8s, it’s a tiny fraction of the operational overhead, and not a massively noticeable performance loss. I would absolutely love a way to deploy and manage db’s as easily as K8s with fewer of the quite significant issues that have mentioned, so if you know of something that is better behaved around singular workloads, but keeps the simple deploys, the resiliency, the ease of networking and config deployments, the ease of monitoring, etc, I am all ears.
- crabbone 3y agoIf you think that deploying anything with Kubernetes is simple... well, I have bad news for you. It's simple, until you hit a problem. And then it becomes a lot worse than if you had never touched it. You are now in the stage of a person who'd never made backups and never had a failure that required them to restore from backups, and you are wondering why would anyone do it. Adverse events are rare, and you may go like this for years, or, perhaps the rest of your life... unfortunately, your experience will not translate into a general advice. But, again, you just might be in the camp where performance doesn't matter. Nor does uptime matter, nor does your data have very high value... and in that case it's OK to use tools that don't offer any of that, and save you some time. But, you cannot advise others based on that perspective. Or, at least, not w/o mentioning the downsides.
- crabbone 3y agoNobody who matters. Those who do that don't know what they are doing (even if they outnumber the other side hundred to one, they "don't count" because they aren't aiming for good performance anyways). Well, maybe not quite... of course it's possible that someone would want to deploy a database in a container because of the convenience of assembling all dependencies in a single "package", however, they would never run database on the same node as applications -- that's insanity. But, even the idea of deploying a database alongside something like kubelet service is cringe... This service is very "fat" and can spike in memory / CPU usage. I would be very strongly opposed to an idea of running a database on the same VM that runs Kubernetes or any container runtime that requires a service to run it. Obviously, it says nothing about the number of processes that will run on the database node. At the minimum, you'd want to run some stuff for monitoring, that's beside all the system services... but I don't think GP meant "one process" literally. Neither that is realistic nor is it necessary.
- hyc_symas 3y ago>but I don't think GP meant "one process" literally. Neither that is realistic nor is it necessary. The point was simply about other processes that could be competing for resources - CPU, memory, or I/O. It is expensive for a user-level process to perform accounting for all of these resources, and without such accounting you can't optimally allocate them. If there are other apps that can suddenly spike memory usage then any careful buffer tuning you've done goes out the window. Likewise for any I/O scheduling you've done, etc.
- morelisp 3y agoI'm running prod databases in containers so the server infra team doesn't have to know anything about how that specific database works or how to upgrade it, they just need to know how to issue generic container start/stop commands if they want to do some maintenance. (But just in containers, not in Kubernetes. I'm not crazy.)
- didip 3y agoMy group and a bunch of my peer groups. And we are running them at the scale that most people can’t even imagine.