4 ms·
Why GlusterFS should not be integrated with OpenStack
- j_s 13y agoApparently there are nearly 20 supported storage backends for OpenStack, this article is discussing the shortcomings of one of them. Not sure why GlusterFS is singled out. https://wiki.openstack.org/wiki/CinderSupportMatrix https://wiki.openstack.org/wiki/CinderSupportMatrix
- epistasis 13y agoIf I understand this correctly, the complaints are: - Terminology -- Seriously? It's not a very strong complaint. - Snapshotting -- have to use qcow2 for this rather than native file system support for snapshotting an individual file - Have to use Layer2 separation for security -- but this should be done any way, shouldn't it? There's no reason to trust this to application level security, and I there's any need at all for this type of security, L2 is the only way to go. Personally, I think Ceph is the future, and I also have personal reasons for wanting Ceph to succeed. Having dealt a bit with both communities, I think it's clear that Ceph is going to be the standard go-to destributed file system soon, and I hope to switch our gluster filesystems to it soon (come on POSIX FS layer!). So I kind of have it in the bag for Ceph. However, I don't see these complaints as very strong. I'm only a dabbler with OpenStack, but fairly experienced with Gluster and its warts.
- mgalkiewicz 13y agoTerminology is not a problem:) It is just a little bit misleading when you start implementing Cinder with Glusterfs. Complaints are mostly about integration of both tools. I dont intend to discredit Openstack/Glusterfs in particular.
- notacoward 13y agoGlusterFS developer here. The OP is extremely misleading, so I'll try to set the record straight. (1) Granted, snapshots (volume or file level) aren't implemented yet. OTOH, there are two projects for file-level snapshots that are far enough along to have patches in either the main review queue or the community forge. Volume-level snapshots are a little further behind. Unsurprisingly, snapshots in a distributed filesystem are hard, and we're determined to get them right before we foist some half-baked result on users and risk losing their data. (2) The author seems very confused about the relationship between bricks (storage units) and servers used for mounting. The mount server is used once to fetch a configuration, then the client connects directly to the bricks. There is no need to specify all of the bricks on the mount command; one need only specify enough servers - two or three - to handle one being down at mount time. RRDNS can also help here. (3) Lack of support for login/password authentication. This has not been true in the I/O path since forever; it only affects the CLI, which should only be run from the servers themselves (or similarly secure hosts) anyway. It should not be run from arbitrary hosts. Adding full SSL-based auth is already an accepted feature for GlusterFS 3.5 and some of the patches are already in progress. Other management interfaces already have stronger auth. (4) Volumes can be mounted R/W from many locations. This is actually a strength, since volumes are files. Unlike some alternatives, GlusterFS provides true multi-protocol access - not just different silos for different interfaces within the same infrastructure but the same data accessible via (deep breath) native protocol, NFS, SMB, Swift, Cinder, Hadoop FileSystem API, or raw C API. It's up to the cloud infrastructure (e.g. Nova) not to mount the same block-storage device from multiple locations, just as with every alternative. (5) What's even more damning than what the author says is what the author doesn't say. There are benefits to having full POSIX semantics so that hundreds of thousands of programs and scripts that don't speak other storage APIs can use the data. There are benefits to having the same data available through many protocols. There are benefits to having data that's shared at a granularity finer than whole-object GET and PUT, with familiar permissions and ACLs. There are benefits to having a system where any new feature - e.g. georeplication, erasure coding, deduplication - immediately becomes available across all access protocols. Every performance comparison I've seen vs. obvious alternatives has either favored GlusterFS or revealed cheating (e.g. buffering locally or throwing away O_SYNC) by the competitor. Or both. Of course, the OP has already made up his mind so he doesn't mention any of this. It's perfectly fine that the author prefers something else. He mentions Ceph. I love Ceph. I also love XtreemFS, which hardly anybody seems to know about and that's a shame. We're all on the same side, promoting open-source horizontally scalable filesystems vs. worse alternatives - proprietary storage, non-scalable storage, storage that can't be mounted and used in familiar ways by normal users. When we've won that battle we can fight over the spoils. ;) The point is that even for a Cinder use case the author's preferences might not apply to anyone else, and they certainly don't apply to many of the more general use cases that all of these systems are designed to support.
- mgalkiewicz 13y ago1) It is great that snapshots are on their way. I am looking forward to use them. All in all you cannot benefit from them in Cinder right now. 2) I dont claim that all bricks must be specified in mount command. I just point out that having let's say 4 bricks it is impossible to mount volume by specifing only 2 servers if both of them are down, yet still the rest 2 work. 3) Like I wrote. It only considers CLI. 4) I totally agree with you. Mounting volume from many locations is one of advantages. It is not supported by Openstack. I dont blame GlusterFS for that. 5) My intension was not to describe GlusterFS cool features but current state (and preview of Openstack Havana implementation) of integration with Openstack.
- notacoward 13y ago(1) You can in some configurations. If you use qemu there's a block-device driver in qemu and another on the GlusterFS back end (as of 3.4), which both allow snapshots via methods external to us. I meant what I said about it being a hard problem. We're determined to deliver a general snapshot function. That's much harder than delivering snapshots that rely on an uncommon and/or unstable base technology, so it's taking a while. (2) Yes, if you want to survive N concurrent failures you need N+1 mount servers, and currently released code only supports N=1. However, http://review.gluster.org/#/c/5400/ http://review.gluster.org/#/c/5400/ has already been merged and will be available in the next release. (3) IMO you should also have mentioned that the problem only manifests in a specific deprecated use of the CLI (from machines other than the servers). Nonetheless, this is a known deficiency which I've been personally pushing to fix. (5) The current state includes many of these "cool features" (thanks!) without need for any specific OpenStack integration. That's kind of the point. Unlike some, we don't need to re-implement features for every access method or use case. 90% of that functionality would be available e.g. to CloudStack or OpenNebula today. IMO making a big deal of snapshots as a differentiator in one direction without mentioning myriad differentiators in the other doesn't leave people with the information they need to make progress toward their own decisions.
- mgalkiewicz 13y ago
- viraptor 13y ago> Compute node downloads such image, puts it on a local disk and boots a VM. This method makes it impossible to use the highly desired live migration That's not true. Live migration is possible both with glance images and cinder volumes.
- mgalkiewicz 13y agoCould you point out some docs about it?
- viraptor 13y agoIt works by default. I don't think you need to do anything crazy about it. Base images will be downloaded by Nova during migration and libvirt/qemu will copy the differences. (assuming you use libvirt, I don't know about other supervisors) See nova.virt.libvirt.driver:LibvirtDriver.pre_live_migrate - the code path for "if not is_shared_storage".
- mgalkiewicz 13y agoAre you sure that you are talking about true live migration? Is your instance available during the migration?
- viraptor 13y agoIt's the openstack's "live migration". There's going to be a pause of course, but it's up to libvirt how long it's going to take. Reconnecting volumes will be always faster than copying the data. You could also put the instances directory on a shared network drive - you don't have to wait for the copy then.
- mgalkiewicz 13y agoWell I am talking about true live migration where not a single icmp echo request is lost.