4 ms·
> didn't do anything particularly more wrong than anyone else? This is almost completely on their admin staff, maybe other people aren't willing to say it, but
by problems 10y ago
> didn't do anything particularly more wrong than anyone else?
This is almost completely on their admin staff, maybe other people aren't willing to say it, but I will. Test your backups. Or at least make sure they're non-zero in size. It should really be Operations 101.
Whether you do this automatically or manually by setting a reminder on your calendar once a week or even month, doesn't matter. Something this simple would have solved their entire issue. I do this and we run a much smaller shop than GitLab. Heck, if we were larger I'd have hot spare database servers in another datacenter in case the primary got nuked by disk failure, network outages or mistakes.
- fsiefken 10y agoAs part of our disaster recovery plan when I was working as a sysadmin at a 150p company we had replicated server (database replicated and webfiles rsync) on hot standby, we just switched the front-facing servers manually. GitLab has a very short blurb on a similar styled HA setup, I'm not sure if and how they have implemented such themselves and if it would have helped in preventing or shortening the recent downtime. They have probably documented their own setup somewhere. "Automated failover can be achieved with pacemaker alongside STONITH network management. Keep in mind that application servers need to be prepared for transitioning to the new network addresses. In this situation you can also opt to synchronize the database via a database specific protocol instead of DRBD. In the documentation for each database you can find out more about the options for MySQL and the options for PostgreSQL." https://about.gitlab.com/high-availability/#filesystem-storage-and-dbms-slave-servers https://about.gitlab.com/high-availability/#filesystem-stora...
- geofft 10y agoSynchronizing the database isn't quite what you want. It's true that in the case of an errant rm -rf it would almost certainly have helped, but it's approximately as easy to run a "DELETE FROM importantdata" and leave off the "WHERE" clause, which would get replicated. And certainly if you're using DRBD (replicate the volume, not the database), an rm -rf will get replicated. I'm just genuinely unsure what a better outcome would have been here. (It's certainly a process failure that no backups existed other than the manual 6-hour-old snapshot, but I'm not sure you can do much better than automating that.)
- problems 10y agoYou can easily add snapshotting using on-disk technologies like LVM or ZFS to that and achieve reliability against such an issue though, as well as being able to do a full text backup (ie: to SQL) from your replicated server at higher performance than you would on production.
- cookiecaper 10y agoThey had LVM snapshots every 24 hours. They lost 6 hours of data because someone had coincidentally triggered a snapshot 6 hours prior to the deletion event. Otherwise, they would've lost several more hours of data.
- problems 10y agoCorrect me if I'm wrong, but wasn't that from staging data, not from an LVM snapshot?
- sytse 10y agoWe had a secondary server. The secondary we were using was basically dead. That's what lead to the problem. We were trying to fix it, but ran the wrong procedure on the primary instead of the secondary Thus wiping the primary, instead of what might have been left on the secondary. The secondary was already removed at that point. Basically the procedure was the following: 1. Secondary falls behind too much, stops replicating. At this point you need to manually re-sync the whole thing 2. A re-sync requires an empty PostgreSQL data directory, so this data was removed 3. Re-sync doesn't work, leading to the other problems 4. At some point team-member-1 thought previous re-sync attempts left data behind, so team-member-1 wiped the data directory again to be sure; except team-member-1 ran this on db1.cluster.gitlab.com (the primary) instead of db2.cluster.gitlab.com (the secondary)