5 ms·
No, and I wouldn't hold my breath for that to happen. The trouble with multi-master is that there are way too many possibilities what it might mean - different
by pgaddict 7y ago
No, and I wouldn't hold my breath for that to happen. The trouble with multi-master is that there are way too many possibilities what it might mean - different purposes require different trade-offs, and PostgreSQL is unlikely to commit to one of those options (at the expense of others). So I'd expect the current situation to continue, i.e. different multi-master projects (open-source and proprietary) built on top of PostgreSQL, catering to different use cases.
For HA, the best option at the moment is probably a physical standby with an external monitoring/management tool (repmgr, patroni, pacemaker, ...) handling the failovers. So not really a multi-master. There are ways to do something similar with logical replication (that's what BDR does), but I don't think there's a widely available tool to manage that.
- MuffinFlavored 7y agoHow long would an HA switchover like that take? Is your “slave” master off to the side, getting data streamed from your “hot” master the whole time to stay up?
- ants_a 7y agoI have measured a planned switchover on a lightly loaded system (100tx/s) at 500ms from commit to commit. On larger systems it might be a couple of seconds. Unplanned failover depends mainly on the chosen timeout. Typical used value is 30s, which gives tolerance for small network problems without causing failovers while not being excessively long. Yes, the standby is constantly streaming and applying transaction logs from the master.
- pgaddict 7y agoRight, that's about the right ballpark - tens of second for unplanned events (failover), a couple of seconds for planned events (switchover). We can have a lenghty discussion about all the caveats and options, but in my experience trying to reduce the times below these (somewhat vague) thresholds is mostly pointless. For switchovers, a couple of seconds should not be a big deal - you can pick when it happens, and there are ways to make it non-disruptive for the application (i.e. you can design the app to tolerate this, or you can use PAUSE in pgbouncer, or whatever). For failovers, the "tens of seconds" may seem a bit too high, but most of the time will be spent determining whether to do the failover or not. Make it too aggressive and you'll be sad. Ultimately, it's a matter of money. People sometime say things like "It has to be 24/7, absolutely not outages, it's a non-negotiable critical business requirement." A good response that is "So you're telling me a 60-second outage of this system will put you out of business?"
- MuffinFlavored 7y ago> I have measured a planned switchover on a lightly loaded system (100tx/s) at 500ms from commit to commit. What does that switchover look like? You have services that resolve the database DNS on startup and create a connection pool. While they are doing 100tx/s, how do you get them pointed to another instance without forcing them to restart? Also, how exactly is a Postgres commit streamed to a DR instance? I'm not interested in all of the possible configurations of master/slaver as much as I am just trying to understand... you've got a hot master receiving 100tx/s... what is mirroring the commit log / write ahead log to the DR master?