3 ms·
OK, so I know nothing about NOTAMs and even less about the FAA infrastructure. But: data errors introduced by a bad disk? In 2022? Storing your data on a hardw
by PreInternet01 4y ago
OK, so I know nothing about NOTAMs and even less about the FAA infrastructure.
But: data errors introduced by a bad disk? In 2022? Storing your data on a hardware-backed (or, ZFS-backed, if you must) RAID6 (or equivalent) partition has been bog-standard for the past decade or so, and a partition is either online (and OK) or degraded (but still OK), or offline, at which point 'restoring the latest backup' (or just failing over to a replica) should bring you back to an OK state.
Then: apparently, restoring the offline partition caused bad data to be ingested, over and over, crashing the system again and again. Even if we pretend that solutions like SQLite have not been available for the past decade or so, 'skip past the head of the corrupted data file' has been best-practice since I last saw abominations requiring such hacks (say: Novell MHS), and sorry to repeat myself, over a decade ago?
TL;DR: please pay me US$ 200M, and I'll make sure all your problems are at least not of the kind solved a long time ago?
- mmastrac 4y agoI'm going to disagree that "skip past the head of the corrupted data file" is something that should be done in a safety-critical environment. What happens if the head of the data file contains the location of a hazard that 100 planes need to be cautious of?
- PreInternet01 4y agoSure, "skipped past ## corrupted octets at offset ####, offending contents saved as `/mnt/log/data/yyyyMMdd/seq.json` and discarded" should absolutely be logged and trigger the highest alert level. It should not, however, impact the availability of the system.