4 ms·
I posted a comment on the blog, but will leave one here as well. In regards to the claim of: "quick instance size scaling with no downtime" I am currently eva
by mathrawka 16y ago
I posted a comment on the blog, but will leave one here as well.
In regards to the claim of: "quick instance size scaling with no downtime"
I am currently evaluating moving our MySQL server to Amazon's RDS. I liked the sounds of the Multi-AZ RDS as it will greatly decrease downtime. However upon examination I found the following issues:
- Changing an instance size results in downtime. Amazon's docs and support say up to 3 minutes.
- Any failover can take up to 3 minutes to promote the backup to becoming live.
- You will experience downtime during that time
- Browsing the support forum, some people complained of it taking more than 3 minutes to failover and in some cases the failover got stuck and they experience a longer downtime until they wrote on the forum and had AWS support manually fix it.
So in other words, this is not a silver bullet to making MySQL have 100% uptime. And in fact you will experience "up to 3 minutes" of downtime each week during the maintenance window when a failover will occur (unless they do the failover before the maintenance, which I have not found information on anywhere).
- pmpkiran 16y agoThe weekly maintenance window you define for Amazon RDS is just a place holder.It does not mean that maintenance happens every week. It only happens when there are specific security or MyQL patches that need to be applied and is pre-announced in the forums.The maintenance window could also be used to schedule instance scaling events, during which, like you pointed out, there would be downtime. However, the downtime would be limited to the instance fail over time (typically less than 3 minutes) if you use a Multi-AZ deployment.
- mathrawka 16y agoAre you aware if they do the failover, then after it is successful perform maintenance? Or do they take the just stop the master server to do maintenance and let failover happen automatically? The first method would eliminate downtime. The second method would have that up to 3 minute downtime. Also, in my basic testing, I only had 1 out of 10 failovers be completed within 15 seconds. The rest took between 2 and 3 minutes. The longest one was slightly over 3 minutes.