3 ms·
It is indeed. Proper procedure would be to pivot to a slave. We have reconfigured all our Redis hosts to prohibit a restart on masters to prevent exactly that
by RobSpectre 13y ago
It is indeed. Proper procedure would be to pivot to a slave. We have reconfigured all our Redis hosts to prohibit a restart on masters to prevent exactly that temptation.
- pilif 13y agoYeah. I totally agree with your conclusions in your article. I was making a general remark because this doesn't just apply to you guys but to everyone else too :-) Btw: When going to a slave, be careful because the state the slave is in might be out of date compared to overloaded master at the point of moving over.
- aphyr 13y agoI should mention that promoting a secondary to a primary in an asynchronously replicated system can (and almost certainly will, under load) result in the loss of committed transactions. Redis has two resynchronization modes. One just recovers a small delta from the primary, if it hasn't diverged that far. If the primary has accepted too many writes, the secondary has to initiate a full resync, which is much more expensive. Since Twilio's postmortem says their nodes initiate full resyncs, I suspect that had Twilio promoted a secondary, they would have lost a significant number of writes. Probably best, under these kinds of scenarios, to design the system such that lost writes don't cause overbilling. ;-)