4 ms·
The upgrade path using logical replication is critical for minimizing downtime. The others should not be considered for production workloads. As far as RDS goe
by tbrock 4y ago
The upgrade path using logical replication is critical for minimizing downtime. The others should not be considered for production workloads.
As far as RDS goes it amazes me that, even with replicas, there are NO options for doing a zero downtime upgrade which is a major gotcha.
By this point they should buffer the writes during the critical section when the switchover happens and make this seamless for users and operators.
- shadowgovt 4y agoIt's basically impossible to do a zero-downtime upgrade with an RDS. Systems that have zero-downtime upgrade "solve" this problem by not being relational (because it pushes the fault-tolerance to the query developer when they have to build their own relations across systems that are defined by API contract to not always be available).
- CWuestefeld 4y agoRight now as we speak, our DBAs are applying patches to our SQL Servers. SQL Always On seems to solve the problem pretty well. Granted, this is a minor upgrade, but even so, we've not had any database downtime for any of the updates we've applied in the last year+ since we got this set up.
- dub 4y agoHypothetically it should be possible to make an entirely new RDS cluster as a replica at a new version and fail over to it, with a similar error rate to a normal replica failover. Setting up the infrastructure to manually manage your own cluster failover would kinda go against the spirit of using RDS and letting AWS manage infrastructure for you, though.
- guiriduro 4y agoYes, its an obvious gap in their offering. I assume they wrap pg_upgrade in their RDS upgrade process, but of course it is in-place and requires downtime. Their multi-AZ replication is a high-availability solution but the primary and secondary must both be the same version of PostgreSQL, useless for major upgrades. For our upgrade 9.6 -> 13.4, we had the added complication that we used PostGIS which added a kink to the upgrade path. A reliable, caveat-free, simple AWS-provided logical replication solution for zero DT cluster upgrade was sorely missed and the downtime on the recommended path was painful.
- rkeene2 4y agoAs a counter example there's COMDB2 which has zero downtime upgrades, because the client (library) maintains the transaction state and can replay it when the server comes back online.
- manigandham 4y ago> "solve this problem by not being relational" That has nothing to do with this. All databases are fundamentally the same regardless of data model. The standard process is to start a new version, replicate the existing database, then cutover once the replica is caught up. Advanced systems can seamlessly switch the primary so it'll redirect new queries to the new primary upon that cutoff. This has been done for decades through external tooling across many systems, including Postgres. However it would be easier if the database included this functionality internally.
- mmontagna9 4y agoYeah there isn't much publically available. We wrote a tool for this at Instacart, it relies on some internal tooling but much of it is generalizable (assuming you run pgbouncers). This is a good reminder for us to write a blog post or something.