3 ms·
Points to take away from this: - Random corruption is a thing - Always make sure that your replication can't accidently screw you: - Replication doesn't
by roeme 13y ago
Points to take away from this:
- Random corruption is a thing
- Always make sure that your replication can't accidently screw you:
- Replication doesn't replace a separate backup system
- A version of your backup data must become immutable at some point
- Make sure you monitor the right metrics of your system.
And most importantly:
Avoid unnecessary noise in your monitoring channel(s).
I keep preaching this; people think they can keep on top of a noisy log, but the ugly truth is that your brain becomes numb and you will miss things. At the very least, you begin to tune out since "it's not that important".
edit:fmt
- dspillett 13y agoYou missed on from the list: Always make sure a recent backup has been tested in a meaningful way. At least have an automated restore scheduled, and have it report to you any error. Have that process run what-ever verification tools your DB supports to as this will catch some corruption that won't be seen in a simple restore. Better still (though this is very app dependent so difficult to create general rules for) do some actual data verification. If you app keeps an audit trail, keep a copy of the last backup around and compare anything that should not have changed between them (as there are no relevant audit entries) still identical. All this serves to increase the confidence that your backup will save you if/when disaster strikes.