5 ms·
>MariaDB Galera is really easy to set up and maintain. I take exception to that. When you log into all of your nodes after a network snafu and discover that ev
by clon 9y ago
>MariaDB Galera is really easy to set up and maintain.
I take exception to that. When you log into all of your nodes after a network snafu and discover that every single node has attempted to perform a full state transfer, involving the destruction of the data directory...
1. rm /var/mysql
2. Attempt to do FST
3. Crap happens
After a while, all nodes were left with no data. Off to backups :)
I am sure there is a way to prevent that, but Galera is still not easy ops wise.
- _rami_ 9y agoOuch. FTR, patroni does a similar thing, but instead of rm /var/lib/postgres it does a mv /var/lib/postgres /var/lib/postgres.bak.date ;)
- clon 9y agoA much nicer approach, if you have the luxury of data disk usage below 50%.
- tetha 9y agoMaybe it's my small scale experience, but so far, multi-master relational databases with automated fail-over seems like a lot of risk for just a little payoff. If you can handle the risk, and the operational/development cost, and you need the payoff, go for it. I'm not in that spot. If I have a master01 with master02 replicating as a standby, I can switch between these two masters within 5 - 15 minutes depending on my setup and at what infrastructural level I do the switch - I could reconfigure the application, switch a dns entry, use a load balancer like maxscale. It's downtime, but it's a low-risk recovery with well-known impacts. With a multi-master setup, I have to do at least two things: First, I must ensure my applications transaction-safety. Read-Write splits with an application with bad transaction management is fun, and write-splitting will end up with even more of a mess. And then I need to setup and operate the multi-master setup, which is a non-trivial decision and selection imo. Just look at the number of possible solutions for postgres. This in turn allows the system to automatically failover in case of trouble - which my current infrastructure would have had to do 3 times over two years, and it would have helped 2 times at most. Except if the failover itself fails and ends up harder to handle than the database failover, like in your case or in other really scary postgres failover horror stories. And interestingly enough, in our b2b context, our customers actually prefer a well-known, low-risk failure plan, even if it is 30 minutes of outage.
- clon 9y agoThis is very much our experience as well. As complexity increases, the marginal utility goes down, as operating concerns skyrocket. You really need pretty special operational needs to run multi master relational databases. It is am engineering tradeoff between complexity and fragility. In our case Galera / PXC ended up with a significantly worse availability record than running our previous simple master slave failover system. See my other comment for more details.
- zzzeek 9y ago> When you log into all of your nodes after a network snafu and discover that every single node has attempted to perform a full state transfer, involving the destruction of the data directory... I support customers that are using Galera at Red Hat and while we have seen lots of network snafus and situations bringing individual nodes back online, I've never seen all nodes attempt to state transfer each other like that nor have I ever seen a data directory "destroyed". The situations with Galera in 100% of cases involve operators acting too hastily and not understanding what's going on as they do it.
- clon 9y agoBy destroyed I mean the data directory is wiped as a first step towards a full state transfer. By the time we restored stable shell access to the nodes, after a lot of intermittent networking, all 3 nodes had indeed wiped themselves. The cluster was no more. There may have been of course also something to do with SeveralNines CC that we were trialling at the time. Perhaps it was trying to do something hasty as you put it. In any case, a lot of moving parts to understand fully.
- zzzeek 9y agoGaleras SST is going to rsync the files over but I can't imagine how you got an SST to initiate from a wiped data directory that would somehow wipe another one. Would love to see a reproducer for the condition you describe.