5 ms·
What kind of database are they using, I wonder, to end up with such a spectacular failure?
by tigerBL00D 4y ago
What kind of database are they using, I wonder, to end up with such a spectacular failure?
- lopkeny12ko 4y agoWhy would this be an indictment of any specific database technology? If your disk fails and corrupts the filesystem, you're toast, regardless of what database you are using.
- mdavidn 4y agoThe technology to detect and recover from disk failures does exist. RAID and ZFS, for example. I would not expect a disk failure to replicate to the backup.
- tanelpoder 4y agoYep and if you ship WAL transaction logs to standby databases/replicas, corrupt blocks or lost writes in the primary database won't be propagated to the standbys (unlike with OS filesystem or storage-level replication). Edit: Should add "won't be silently propagated"
- ct520 4y agoWorking with critical infrastructure and lack of true in depth oversight it wouldn’t surprise me DR plans were not ever executed or exercised in a meaningful manner.
- ChrisMarshallNY 4y agoThis is quite common. Comprehensive DR testing is really difficult. Many orgs settle for “on paper,” or “in theory” substitutions for real testing. They do it right; no problem. Doing it right, though … there’s the rub …
- ilyt 4y agoNeither checks the checksum on every read as that would be performance-prohibitive. So "bad data on drive -> db does something with corrupted data and saves corrupted transformation back to disk" is very much possible, just extremely unlikely. But they said nothing about it being bad drive, just corrupted data file, which very well might be software bug or operator error
- sigotirandolas 4y agoThis is wrong, both ZFS and btrfs verify the checksum on every read. It's not typically a performance concern because computing checksums is fast on modern hardware. Besides, historically IO was much slower than CPU.
- guenthert 4y ago> Neither checks the checksum on every read as that would be performance-prohibitive. It is expensive. It might be prohibitive in a very competitive environment. This is hardly the case here. Safety first!
- ExoticPearTree 4y agoRAID does not really protect you from bit rot that tends to happen from time to time. ZFS might because it checksums the blocks. But if the corruption happens in memory and then it is transferred to disk and replicated, then from a disk perspective the data was valid.
- techie128 4y ago> If your disk fails and corrupts the filesystem, you're toast, regardless of what database you are using. There are databases that maintain redundant copies and can tolerate disk / replica failure. e.g. Cassandra.
- efficax 4y agojournal databases are specifically designed to avoid catastrophic corruption in the event of disk failure. the corrupt pages should be detected and reported by the database will function fine without them
- redox99 4y agoIf you mean journaling file systems, no. They prevent data corruption in the case of system crash or power outage. That's different from filesystems that do checksumming (zfs, btrfs). Those can detect corruption. In any case, if you use a database it handles these things by itself (see ACID). However I don't believe they can necessarily detect disk corruption in all cases (like checksumming file systems).
- birdyrooster 4y agoImagine you have one node which is running as a replica of another and it takes the backups. Well, let’s pretend it is backing up the corrupted data once in a while and it happened to overwrite their cold backup. They could have any number of databases and still had this failure. It’s more their methodology for taking backups. They should have many points in time to choose from to rebuild their database. They should be testing their databases before backing them up blindly.
- tenken 4y ago> They should be testing their databases before backing them up blindly. Oh you mean they should be testing/validating the generated backup db file before replicating it to long-term archive ...
- readthenotes1 4y agoWay back when use cases were a thing, I used to chide people for saying that Backup was a use case. No, Restore is a use case. (Replace "use case" with "requirement" or "user story"...)
- samman 4y agoA corollary to this would be: “Backups are worthless. Restores are priceless.”
- birdyrooster 4y agosemantics but yes
- eurasiantiger 4y agoWell, for example, MySQL/MariaDB using utf8 tables will instantly go down if someone inserts a single multibyte emoji character, and the only way out is to recreate all tables as utf8mb4 and reimport all data.
- colinjoy 4y agoSurely nobody would use that format and allow a commit message including emojis to cause an effective DOS for a large Sonarqube project.
- NavinF 4y agoIt doesn't block inserts with invalid data? I thought that was the whole point of telling the database what types you're using
- ilyt 4y agoIt does and poster above is incompetent
- eurasiantiger 4y agoI have had customer production sites go down due to this issue when emojis first arrived. It was a common issue in 2015. I would hope it is fixed by now!
- dpcx 4y agoMySQL historically isn't very good about blocking bad data. Sometimes it would silently truncate strings to fit the column type, for example. It's getting better as time goes on, though.
- dolmen 4y agoI need more info about this.
- lsaferite 4y ago
- rini17 4y agoWe had Oracle corrupt itself due to software bug. It similarly went undetected for some time and thus ended in backups.