3 ms·
> 3. Had the higher-priority request arrived even one millisecond earlier or later, the update would have completed normally. Well, that's comforting to know.
by dtf 14d ago
> 3. Had the higher-priority request arrived even one millisecond earlier or later, the update would have completed normally.
Well, that's comforting to know.
BBC: "Flight chaos caused by software defect in space of a millisecond, report says"
Sky: "'Millisecond' software error caused air traffic outage that grounded thousands of flights"
The Guardian: "Flight chaos for hundreds of thousands was caused in ‘millisecond’ by software error"
Sounds like pure bad luck.
- philipwhiuk 14d agoOr insufficient testing?
- nradov 14d agoTesting isn't an effective way to catch most race conditions. Code reviews, static analysis, and rigorous enforcement of concurrent coding standards is usually a better approach.
- jiehong 14d agoPerhaps something like what TigerBeetle does: deterministic simulation testing.
- dtf 14d agoOr maybe even just reviewing logic that is subject to pre-emption. Maybe I'm being too harsh.. on the plus side the system has at least failed hard every time there's been a fault. Nobody has died. But it's been 3 times now in the past couple of years, and two of those times resulted in over 2000 flights cancelled and days of backlog, and misery for hundreds of thousands. It's really not acceptable.
- chrisjj 14d ago> the system has at least failed hard every time there's been a fault You don't know that.
- cmpxchg8b 14d agoSafety critical systems demand formal verification. This wasn't bad luck, this was poor craft.