7 ms·
What really happened at Ma.gnolia and lessons learned [with video]
- lbrandy 18y agoThe short version: the file system got corrupted. The backup was just a file-sync over a firewire network to another machine. Meaning the bad data was backed up and presumably overwriting the older, good data. They had a RAID but the problem was a software filesystem so the errors just got stored. He seems to understand how terrible of a design decision he made in regards to the back-up system, and he appears physically affected when having to admit, publicly, the details of the infrastructure (or lack thereof) that caused this.
- moe 18y agoA one-liner to add insult to injury: sed 's/rsync/rdiff-backup/g' <bin/my-backup.sh >bin/my-real-backup.sh
- moe 18y agoAfter watching the vid I have to take that back. Apparently there was a SQL database involved... But rdiff is highly recommended nonetheless.
- hachiya 18y agoAnyone know how rdiff-backup compares to duplicity? I know duplicity has an option to turn off encryption, if one wants to remove that overhead...
- jrockway 18y agordiff-backup keeps a live version of the filesystem available, in addition to backups. This means a full restore is just a `cp -a` operation. FWIW, I've used both, and I like the opacity of the duplicity backups, since I store them on S3. If you are syncing to a nearby disk, though, then you might like rdiff-backup better.
- moe 18y agoSeconded, both have their place. I have tried pretty much all of them, incl. snapBack, dirvish and various homegrown scripts building on top of rsync, rdup, rcs and so on. rdiff and duplicity are the most mature of the pack which shows mostly in their handling of corner cases (connection loss during backup, resume of a partial/failed backup, disk full during backup, handling of really large trees) but also in overall convenience and robustness (legibility of on-disk format, configuration sanity, tools to find a specific revision of a file, flexibility in retention/purge intervals etc.). I generally recommend rdiff as the default tool for backups to a remote spinning disk. duplicity, as parent suggests, is good when you need your archive to be a single large file which helps with handling in some situations. There is also dar worth mentioning which is less useful for incremental stuff but can add redundancy to archives which is good for archiving to unreliable/decaying media (DVD, Tape). Be aware though that older versions had problems with large archives, use a recent version. And last no least, if you have a tape library then Bacula is a mighty tool. Easier to use and pretty much on par in terms of features compared to the commercial offerings and the residents like Amanda. We generally deploy a single backup server here with lots of disks that pulls snapshots from everywhere via rdiff and either mirrors the local repository to a remote location or feeds the precious data to tape via Bacula.
- Maro 18y agoI think there is an implicit assumption here that if you copy over a live Mysql DB file (I assume that's what he was doing), you get a consistent (aCid) view of your DB. This is a false assumption. If this is what he was doing, then possibly the older 'good' backups weren't any good either.
- lennysan 18y agoA great lesson for every saas service, and how a little bit of transparency goes a long way: http://www.transparentuptime.com/2009/02/magnolia-downtime-saas-cloud-trust.html http://www.transparentuptime.com/2009/02/magnolia-downtime-s...
- jim-greer 18y agoIf you have any kind of staging/testing server I'd highly recommend using your production backups to populate that on a regular basis. That way you test your new code releases with real data, and you know that your backups work.
- wmoxam 18y agoThis is what I do. Using Amazon's EC2+EBS makes it dead simple. Every time we test a release I setup a staging env (from scratch) and restore the latest EBS snapshot. That way I'm testing both the database backup and that I've kept all of the server images up to date as well. It doesn't take much time either (about 30 minutes) and is done weekly.
- tptacek 18y agoQuick cherry bomb to lob into this conversation: populating insecure test servers with sensitive production data is a classic web app company security failure. It probably doesn't matter for you, but be cognizant of it.
- Tangurena 18y agoI agree. One of our big financial clients has an automated tool to scrub such data, but then they have social security numbers as well as lots of other juicy financial data. So they're worried about all sorts of stuff that most of us never ponder as a business risk. One of the santizing steps is to replace all passwords with a set value, such as six/seven of a letter (like "A") or a number (eg, "111111"). Another sanitizing step is to scramble names and addresses. Usually the first letter gets preserved, and the rest gets replaced with a hash (say, MD5 it, and then base64 it and truncate it to length, that way it preserves max lengths and typical size of words). example: John Doe, 1313 Mockingbird Lane might get munged into Jiqw Dyh, 1313 Masdfasdfas Lfds We just have username/password/address/phone, so all we do is set all passwords to a default value (all emails, if any, get set to mine), and munge up telephone numbers. Later this year I'll cobble up a better sanitizer. Our parent company has to worry about GLBA compliance, but our little apps don't "collect" enough information to worry about GLBA at this time.
- markup 18y agoIf you were to start a web application from scratch, how would you deal with important database backups?
- hachiya 18y agohttp://forums.theplanet.com/index.php?showtopic=91115&view=findpost&p=599193 http://forums.theplanet.com/index.php?showtopic=91115&vi... This looks like a simple and safe way of handling repeated MySQL backups.
- markup 18y agoYeah, I use a similar approach. However, for ma.gnolia we are talking about a database reaching half a terabyte (unless I misunderstood), so my question was more like: what do you people consider the best way to approach database backup so that it's sustainable, scalable and the most disaster proof? Got any testimoniance (whether personal or of public companies), or case study?
- joshu 18y agoI wonder how they managed to get to half a terabyte. Delicious's was smaller even for millions of users.
- markup 18y agoYeah indeed... .5TB is huge for bookmarks (title, url, tags, description). I have never had the chance to build anything this big but if you imagine the amount of text you could fit in 500GB, it makes you wonder. I gave a look at your comments and found this one which replies to my question perfectly (I didn't follow that thread, discovered it right now): http://news.ycombinator.com/item?id=459000 http://news.ycombinator.com/item?id=459000 -- thank you for sharing your experience!
- sjh 18y agoAccording to the write-up on Wired (http://blog.wired.com/business/2009/01/magnolia-suffer.html http://blog.wired.com/business/2009/01/magnolia-suffer.html), Ma.gnolia also took a snapshot of the page being bookmarked. This may account for the size of the database.
- jonasvp 18y agoMy recommendation for basic backup needs: rsnapshot. I backup our public server to our internal network as well as my desktop machine to an encrypted portable drive using it: http://www.rsnapshot.org/ http://www.rsnapshot.org/ It's probably similar to rdiff-backup, which I haven't used. If you're fine with daily or hourly backups and don't have too much data (<100 GB), rsnapshot together with regular SQL dumps works fine.
- Maro 18y agoHe had RAID and was doing filesystem level backup, ie. copying over the entire Mysql DB file. When filesystem-level corruption occured, the backup script overwrote a good (perhaps 1 day old) backup file with a corrupted file, so he's backup was worthless. The first thing that comes to mind is that he could have used application-level backup, ie. Mysql. The script would have noticed that the DB is corrupted because reads (SELECT) would have failed, and the backup script would have stopped and sent him an email to restore the good backup file. If he used a cloud service like Amazon SimpleDB, he wouldn't have to worry about filesystem-level corruption, because that's abstracted away by Amazon. (And it's replicated.) This is still not enough though. What if the site gets hacked and the hacker issues DELETE statements. Then all your data is deleted, and even if you have application-level backup, it will succeed (it will read the empty DB), thus overwriting your old backup. I guess the conclusion is to keep around several copies of the data, and have sanity-checks in place to avoid overwriting good backups. In his case it was hard (given it's a homegrown application) to keep around many copies, because his DB was 500G in size.
- amix 18y agoA simple tip for those that run any kind of database: Be sure to replicate them in master-slave (or master-master). And base your backups on taking a slave down for backups. Hot backups only work for very small databases - even those that are based on LVM snapshots, tarsnap, innodb hotbackups etc. With big databases, you will be most likely IO bound and a backup will take your site down. If you have lots of load and lots of data then re-creating a slave will require lots of downtime. For Plurk.com we have had a 4 hour downtime due to re-creating a slave, so be sure to run a master-slave setup and have fresh slaves replicated at all times (we have learned this the hard way :)).