4 ms·
I worked with a GlusterFS deployment in production about 2 years ago, and it was such a nightmare that I both feel compelled to write about it and never touch a
by gamegod 11y ago
I worked with a GlusterFS deployment in production about 2 years ago, and it was such a nightmare that I both feel compelled to write about it and never touch anything made by that team ever again.
It was the whole shebang: Kernel panics, inconsistent views, data loss, very slow performance, split-brain problems all the time. Our set up IIRC was very simple: two bricks in a replicated volume. It worked so poorly that we had to take it out of production. Some of our experience can be explained by GlusterFS performing poorly under network partitions, but nothing could justify kernel panics. It blew my mind that Redhat acquired that company and product.
Edit: I hope there's been a big improvement to the reliability and performance of GlusterFS. Can anyone with more recent experience running it in production comment?
- illumin8 11y agoI'm not a GlusterFS expert, and haven't used it before, but you should know that most consensus algorithms (Paxos, Raft, etc) only function reliably with an odd number of nodes. I have to wonder if your problems were mostly self-inflicted from having 2 nodes. Of course, any network partition in a 2-node cluster has a huge potential for data corruption, as each node now thinks it is the master (split-brain). In a 3-node cluster, any system with a decent consensus algorithm (to be clear, I'm not sure if GlusterFS has one) would know that during a partition the cluster can only continue to operate if at least 2 nodes can communicate with each other to elect a new master.
- tveita 11y ago> Of course, any network partition in a 2-node cluster has a huge potential for data corruption, as each node now thinks it is the master (split-brain). This is not a given, the cluster can and should refuse to operate if it can not get a majority vote for the master, if that is required to prevent data corruption. With two nodes that means both nodes must be active and reachable, with three nodes one node can fail.
- gamegod 11y agoI didn't do the initial configuration myself but I wouldn't rule out this kind of gross, self-inflicted problem. One has to wonder why the software even lets you run it with an unsafe number of nodes though... (Lots of distributed software is guilty of this...)
- toredash 11y agoUh, well if you base your GlusterFS experience on that, no wonder you have a bad view on the technology. I can only assume that your setup only used two nodes, which is not really supported, and of course will cause split-brain problems. And two bricks, really? GlusterFS is best aimed for scale both on the server and client side, two bricks is not really scale.
- gamegod 11y agoThere's probably a good case study to be made here on why so many people have had bad experiences with GlusterFS. If the problem is that it's too easy to set it up in a fragile way, then they need to fix that, because otherwise you're going to end up with a lot unhappy users (as you can see in the other comments). Still, if I can't make it work well on a reasonably small scale, why would I expect it to work any better when I scale it up?
- toredash 11y agoI would suggest reading the documentation and understanding the architecture. It works well for many applications but some are less ideal. Maybe yours was a corner case? I don't know. But it's used by many, for instance Facebook (https://www.socallinuxexpo.org/scale/14x/presentations/scaling-glusterfs-facebook https://www.socallinuxexpo.org/scale/14x/presentations/scali...).
- oh_sigh 11y agoBut then in a split with a 3-node cluster, there is now a 2-node cluster in charge, and what happens if another partition happens?
- craigyk 11y agoMy experience was not as bad as yours, but the problems I saw at a similarly small scale (6 servers, 3 2X replica sets) make me glad I'm not having to scale up a GlusterFS infrastructure. The biggest problem I saw was with absolutely atrocious performance when healing after a temporary node loss. Even if it was just a loss of a minute or two (server reboot), it wouldn't lose availability or corrupt files, but performance would get so bad for the next 2 hours that it might as well have been down. If I had to make a comparison, I'd say GlusterFS reminds me a lot of MongoDB in the beginning. It wins a lot of kudos at the outset based on ease of setup, management and CLI UI, plus it has a good "story" on ability to scale up that gradually begins to fray when pushed. Hopefully there have been big improvements.
- craigyk 11y agobased on another comment. I forgot to add that I always had an odd number of servers in the GlusterFS cluster to help with consensus, even if that server did not actually hold any data bricks.
- codefreakxff 11y agoI feel compelled to write a disagreement. I've been running GlusterFS for 6 months with 15TB of data on a 43TB cluster using 5 servers with zero issues. I have no idea what your particular combination of bad luck was, but I don't think your experience is truly reflective of the product, the team, or Red Hat's sensibilities.
- gamegod 11y agoYou're living the dream then, and I'm glad to hear this is possible. :) Do you mind sharing which GlusterFS version you're on, which kernel, and roughly what your load profile is? (eg. lots of small files, read-heavy, write-heavy, etc.) Also, are you using Red Hat Gluster Storage or the open source version?
- codefreakxff 11y agoopensource glusterfs 3.7.6 Ubuntu 14.04.3 LTS 3.13.0-24-generic g read heavy with a mix of 100MB or 15GB files depending on datasets
- richardfontana 11y agoRed Hat Gluster Storage is open source.
- toredash 11y agoI'm not surprised that you have good experience with this setup.
- notacoward 11y agoYou're right: nothing justifies kernel panics. There is nothing that GlusterFS or any other user-space program should be able to cause one. We (yes, I'm a GlusterFS developer) don't really do anything that any application shouldn't able to do as far as the kernel is concerned. If what we do on behalf of our callers causes a crash, it's the kernel developers' fault and you should engage with them instead of blaming your peers out in user-land. As far as performing poorly under network partitions, I'd love to hear more. That is our responsibility, and sounds like something we can/should fix.