6 ms·
If I read things correctly, they made a fairly... interesting... tradeoff: 954 still-as-yet-unreconciled DB writes in exchange for 24 hours of site downtime. I
by eric_b 8y ago
If I read things correctly, they made a fairly... interesting... tradeoff: 954 still-as-yet-unreconciled DB writes in exchange for 24 hours of site downtime.
I think I'd have made a different choice, but cool that they were upfront about it.
- TheDong 8y ago> we captured the ... writes ... that were not replicated ... For example, one of our busiest clusters had 954 writes Their wording ("one of") makes it sound like they had up to perhaps one order of magnitude more (1k-10k), but they do not actually give us a useful number, merely say "it wasn't too much" and "one of an unknown number of total clusters had 1k". > I think I'd have made a different choice I think you might misunderstand how these things go and the tradeoffs. At the beginning of the incident, they find themselves with lots of writes in the west-coast master which aren't in the east-coast one, and some in the east-coast one that aren't in the west-coast one. Orchestrator cannot promote a working master because they have diverged. Your choices are: 1. Lose data (large unknown amount) by dumping an hour of west-coast data and going back to east-coast data, can be done in maybe 3 hours total by just deleting west-coast and serving all traffic from east-coast for as you rebuild the west-coast cluster. 2. Lose data (small amount) by rebuilding east-coast from a backup so it can be promoted to, replicating west coast data to it, promoting to it (what they did, 24 hours of time), and then try to manually fix up the small amount of lost data after-the-fact (ongoing) 3. Develop tooling to automate the reconciliation of data while the site is down such that the east-coast side can be merge-promoted without a rebuild, probably takes at least 3 days to build and might break everything, but if it works it probably merge-promotes in under an hour. 4. Keep the site down until east-coast data is manually reconciled, still requires rebuilding east to promote west to it, but then requires manually handling some lost writes... probably about 30-40 hours total downtime. 5. Update the application servers to work fine when the us-west DC is the master, probably about 2 months of development with the site down for the duration. Which choice would you have made instead? I'm willing to bet they went with either 2 or 4 (and if 4, changed to 2 when it took longer than expected). They probably assumed it would take about 4 hours, and then it simply took much longer than expected because computers are complicated, it turns out.
- eric_b 8y agoI'd have opted for 1 personally (losing 30 minutes of data) assuming I knew the alternative was 24 hours. I'll grant you they probably figured it would be much faster than that. Alternatively, why not just let the West-Coast replicate to an East-Coast slave, and switch em. Surely the peering bandwidth between West-Coast/East-Coast is higher than whatever they were doing pulling full backups from the cloud?
- TheDong 8y ago> I'd have opted for 1 personally (losing 30 minutes of data) assuming I knew the alternative was 24 hours [of read-only access to prod] With no offence meant, I'm glad you don't get to make decisions at github then. I'd lose faith in github for that since their primary purpose is to store other people's data (git blobs+metadata, comments, etc). If it was their data, it could be fine, but since it's user data, I don't think losing it is acceptable, even if that means being read-only for an entire week. > why not just let the West-Coast replicate to an East-Coast slave, and switch em? At the time of the incident, east-coast has already diverged, so it's not possible to replay logs from west-coast without moving east-coast back in time to some point prior to west-coast (but a point recent enough that west-coast still has full replica logs to replay). It's not possible to simply "rewind" a mysql database typically without restoring a backup. It's not possible in many mysql clustering solutions to catch-up to a peer unless you're already within a recent time window where the replay logs are still available. As such, the only option is to restore from backup and then catch up, which is what they did. I assume if taking another backup of west-coast, transferring it to east-coast over the DC link, and restoring it was faster, they would have done it, but I would be totally unsurprised if that was about the same total time. I'd also say that if you have a practiced procedure (restoring from a scheduled backup in the cloud) vs an ad-hoc procedure (custom backup, custom storage and transfer), the former is probably much safer to do when you're under fire and want to minimize risky moves stressed engineers have to make.
- eric_b 8y agoThe source code is paramount, no argument there. But this wasn't about that. I find the metadata less important, but as you say, I don't make the decisions. The factor you aren't considering is opportunity cost given productivity loss. In your hyperbolic example where you claim you'd rather they were read-only for a week than lose 30 minutes of data? I'd rather get things done during that week personally, and if the price I have to pay is resubmitting a comment or re-approving a PR, I guess I'd pay it. Also, MSSQL has had quick delta snapshot/restores forever (assuming you have them turned on and going every hour or so), mysql really does not have a similar feature?
- itsdrewmiller 8y agoThe incident lasted 24 hours, but the vast majority of that time was spent in a degraded state, not fully down. It's likely they didn't know the scope of the unreconciled writes until later in the incident, too.