5 ms·
From a sideline viewers perspective, persistence in production for Docker containers is a big problem that I've yet to see a good solution for. You end up havin
by Oculus 12y ago
From a sideline viewers perspective, persistence in production for Docker containers is a big problem that I've yet to see a good solution for. You end up having to keep DB servers and such outside of your container pool.
- dugmartin 12y agoI haven't used Docker in production but isn't the standard practice for persistence to mount a host directory as a data volume?
- vidarh 12y agoYou don't. That's the point of volumes. You do need to be careful to ensure you mount volumes or everything that needs to persist, but in practice that's not a very onerous limitation. Since volumes are bind-mounted from the outside, you can put the volumes on whatever storage pool you want that you can bring up on the host (as long as it meets your apps requirements, e.g. POSIX semantics, locking etc). E.g. at work we have a private Docker repository that runs on top of GlusterFS volumes that are mounted on thehost and then bind mounted in as a volume in the container. I also run my postgres instances inside Docker, with the data on persistent volumes. We could certainly use better tools to manage it, though. Another pattern I ought to have mentioned (that I've used myself) is to set up "empty" containers whose only purpose is to act as a storage volume for another container. I don't like that as much, mostly since I've not had as much long-term experience with how the layering Docker uses would impact it.
- Rafert 12y agohttps://github.com/clusterhq/flocker https://github.com/clusterhq/flocker seems to be doing some interesting stuff in that direction.
- cpuguy83 12y agoAnd we're working with the folks @ clusterhq to bring some of that into Docker.
- eropple 12y agoHaving persistent containers strikes me as the infrastructural equivalent of a code smell. I don't use in-container storage for anything that needs to be persisted at all, and I'm not sure why you would in a modern environment. Everything can fail and fail hard, and writing meaningful (i.e., volume'd) data to disk seems like asking for trouble. Ephemeral containers just seems to fit the model of Docker's capabilities much better than maintaining state. The exception to this would be, I guess, you could wedge in Docker containers for database server isolation or something, but my databases don't run on multi-tenant instances so there isn't a huge win to it. (I use RDS most of the time, let somebody else manage that problem.)
- vidarh 12y ago> Having persistent containers strikes me as the infrastructural equivalent of a code smell. Which is why pretty much all advice regarding Docker is to use volumes, so that the persistent data is managed from outside the container, and the container itself can be discarded at will without affecting the data volume. My preferred method is bind-mounted volumes from the host. They are not part of the containers, and the purpose is exactly to remove the need or persistent containers. A lot of the examples I gave in my article relies heavily on this. This leaves the form of persistence up to the administrator. And of course that means you can do stupid things, or not. On my home server that means a mirrored pair of drives combined with regular snapshots to a third disk + nightly offsite backups. At work, we're increasingly using GlusterFS on top of RAID6, so we lose multiple drives per server, or whole servers before the cluster is in jeopardy (and even then we have regular offsite snapshot throughout the day + nightly backups). If you are referring to the pattern of creating "empty" containers to act as storage volumes, then I sort-of agree with you, but mostly because of the maturity of Docker. After all nothing stops you from putting the docker storage itself on equally safeguarded storage. It's not really the risk of losing storage that makes me prefer stateless containers, but that separating state and data substantially reduces the data volume that needs to be secured (since we can spin up new stateless containers in seconds, we only really care about preventing loss of the persistent data volumes).
- eropple 12y ago
- peteridah 12y agoWe decided not to use docker containers for our postgres DB in production; Volume mounts just don't make me sleep easy. We use an S3-backed private docker registry to store our images/repositories.
- cpuguy83 12y agoI'd be interested to know what your concerns are with volumes.
- vidarh 12y agoWhat is your issue with volumes? Bind mounts have been battle tested over many years. I e.g. have production Gluster volumes bind-mounted into LXC containers that have been running uninterrupted for 5+ years.
- klochner 12y agoWith docker specifically, I've had frustrations with the user permissions on the volume between {pg container, data container, host OS}, with extra trouble when osx/boot2docker is added to the stack. Also, docker doesn't add as much value for something like postgres that likely lives on it's own machine.
- vidarh 12y ago> I've had frustrations with the user permissions on the volume between {pg container, data container, host OS} Then don't use data containers. I don't see much benefit from that either. The stuff we put on data volumes is stuff we want to manage the availability of very carefully, so I prefer more direct control. And so when I use volumes it's always bind mounts from the host. Some of them are local disk, some of them are network filesystems. We have some Gluster volumes that are exported from Docker containers that imports the raw storage via bind mounts from their respective hosts, for example, and then mounted on other hosts and bind-mounted into our other containers, just to make things convoluted - works great for high availability (I'm not recommending Gluster for Postgres, btw.; it "should" work with the right config, but I'd not dare without very, very extensive testing; nothing specific with Gluster, just generally terrified of databases on distributed filesystesms). > for something like postgres that likely lives on it's own machine. We usually colocate all our postgres instances with other stuff. There's generally a huge discrepancy between the most cost-effective amount of storage/iops, RAM and processing power if you're aiming for high density colocation, so it's far cheaper for us that way.
- deleted 12y ago[deleted]
- cpuguy83 12y agoSo, we're making volumes better! Please see: https://github.com/docker/docker/pull/8484 https://github.com/docker/docker/pull/8484 And that is only the beginning. It may be presumptuous to say the referenced PR will be in 1.4, but that's certainly what I'm pushing for.
- vidarh 12y agoAny plans on making it possible to control what's backing those volumes more directly? For me, when I'm using data volumes, it's generally because I have specific requirements (e.g. want my database on my expensive SSD array's; want my high availability file storage on a Gluster volume or similar) that doesn't make it interesting to have them co-located with the container storage. I don't want to care if the container storage is totally destroyed - I want to just re-create it from our registry on a different host. I keep threatening the devs I work with that I'll wipe the containers regularly, for a reason, and they're specifically and intentionally not backed up. From what I hear, the idea of keeping the containers totally disposable is a key appeal of Docker for many (me included), so I would just not use any volume related functionality that co-mingles the data volumes other than in very special circumstances (e.g. lets say we had static datasets that we'd like to mix with an application in different configurations; but I don't have any actual, real-life use cases for that at the moment; basically I'd only do that if I then also could push those volumes into our registry)
- cpuguy83 12y agoFirst, volumes are 100% separated from containers. If you remove a container, the volumes are not removed unless you explicitly told docker to (docker rm -v <container>). And even then, it won't remove the volume if other containers are using it. That said, it is hard to re-use a volume right now if you removed the last container referencing that volume. The linked PR solves this issue. You can mount these SSD's on the host and then use bind-mounts (docker run -v /host/path:/container/path) to get those into the container. Or you can add the devices to the containers directly ("docker run --device /dev/sdb", for example), but you'd need privilaged access to actually mount the device. The referenced PR itself doesn't really make using things like specialized disks any easier, except you can register them with the volumes subsystem and it will track them form you, ie `docker volumes create --path /path/to/data --name my_speedy_disks`. There was some discussion around being able to deal with devices directly with Docker instead of expecting the admin to handle mounting those devices onto the host so they can be used as volumes. Nothing finalized here. Here is where that discussion happened: https://botbot.me/freenode/docker-dev/2014-10-21/?msg=23911528&page=5 https://botbot.me/freenode/docker-dev/2014-10-21/?msg=239115...