27 ms·
I'd have opted for 1 personally (losing 30 minutes of data) assuming I knew the alternative was 24 hours. I'll grant you they probably figured it would be much
by eric_b 8y ago
I'd have opted for 1 personally (losing 30 minutes of data) assuming I knew the alternative was 24 hours. I'll grant you they probably figured it would be much faster than that.
Alternatively, why not just let the West-Coast replicate to an East-Coast slave, and switch em. Surely the peering bandwidth between West-Coast/East-Coast is higher than whatever they were doing pulling full backups from the cloud?
- TheDong 8y ago> I'd have opted for 1 personally (losing 30 minutes of data) assuming I knew the alternative was 24 hours [of read-only access to prod] With no offence meant, I'm glad you don't get to make decisions at github then. I'd lose faith in github for that since their primary purpose is to store other people's data (git blobs+metadata, comments, etc). If it was their data, it could be fine, but since it's user data, I don't think losing it is acceptable, even if that means being read-only for an entire week. > why not just let the West-Coast replicate to an East-Coast slave, and switch em? At the time of the incident, east-coast has already diverged, so it's not possible to replay logs from west-coast without moving east-coast back in time to some point prior to west-coast (but a point recent enough that west-coast still has full replica logs to replay). It's not possible to simply "rewind" a mysql database typically without restoring a backup. It's not possible in many mysql clustering solutions to catch-up to a peer unless you're already within a recent time window where the replay logs are still available. As such, the only option is to restore from backup and then catch up, which is what they did. I assume if taking another backup of west-coast, transferring it to east-coast over the DC link, and restoring it was faster, they would have done it, but I would be totally unsurprised if that was about the same total time. I'd also say that if you have a practiced procedure (restoring from a scheduled backup in the cloud) vs an ad-hoc procedure (custom backup, custom storage and transfer), the former is probably much safer to do when you're under fire and want to minimize risky moves stressed engineers have to make.
- eric_b 8y agoThe source code is paramount, no argument there. But this wasn't about that. I find the metadata less important, but as you say, I don't make the decisions. The factor you aren't considering is opportunity cost given productivity loss. In your hyperbolic example where you claim you'd rather they were read-only for a week than lose 30 minutes of data? I'd rather get things done during that week personally, and if the price I have to pay is resubmitting a comment or re-approving a PR, I guess I'd pay it. Also, MSSQL has had quick delta snapshot/restores forever (assuming you have them turned on and going every hour or so), mysql really does not have a similar feature?
- sethhochberg 8y agoMySQL does not have any kind of native snapshotting capability, no. It is becoming somewhat common to run large MySQL deployments on filesystems which do (like ZFS) to get a similar effect, albeit without the ability to restore a snapshot on a host while that database host is online. There is even some experimental work going on in the community to use ZFS snapshots as a state-transfer mechanism for laggy nodes in a Galera MySQL cluster, which seems promising.
- 1stranger 8y agoWe're just speculating how much data would have been lost but one option would be to tell the owner of the data "hey you commented on pull request foo with comment 'baz' but we were unable to save it. Please comment again if necessary." It's not a great UX but neither is 24 hours of downtime.