6 ms·
Many years ago I wrote navigation software for ocean going vessels/ships. We used double and triple redundancy on many of our sensor types. We generally used th
by Randor 7y ago
Many years ago I wrote navigation software for ocean going vessels/ships. We used double and triple redundancy on many of our sensor types. We generally used three control computers that would 'vote' before deciding to make vessel navigation changes. We also always included a "Dead man's switch" that allowed the bridge crew to take control at any time.
I can't even imagine designing aerial vehicle autopilot without redundency. The stakes are too high...
I would be interested in having a look at the statistical model they used to prove 'the system was reliable with zero-redundancy'. While designing these systems for ships the only way we were able to get an error probability near zero was when we used triple redundancy.
- adreamingsoul 7y agoI recently watched a documentary about a diving ship that had a tripple redundancy positioning system that failed during a dive. It was fixed with a hard reboot.
- inamberclad 7y agoIt wasn't an autopilot that failed. This was just supposed to be an accessory system.
- rkagerer 7y agoHow did you deal with redundancy "choke points"? eg. What if the component that tallies computer votes and actuates things based on the results fails (especially in a way that's hard to detect)? Were you able in some cases to maintain isolation all the way through from sensors to actuators and design such that a single failed one (in a worst case failure mode) could be overcome by the rest?
- fit2rule 7y ago(Disclaimer: SIL-4 programmer for safety critical rail transportation applications) The way this is done is that 2 of 3 computers need to 'agree' on the final decision in order for it to be considered the correct one - there isn't a single point of assessment, but rather a consensus that must be formed from the results of all 3 computers. Ideal case, all 3 produce the same results. This has worked successfully for decades. What's changing now, is that those 3 computers are now no longer the same architecture - you'll have a PPC and an x86 and an ARM-based CPU all attempting to agree to the same data, in order to prevent systemic failure throughout.
- WWLink 7y agoWhat would be cool is if the software and hardware design is done by 3 completely different teams. Hehe.
- crocal 7y agoTried already! It’s a disaster. Nancy Leveson & friends wrote articles about this. Yet, there are still people today convinced diverse programming is a good idea... Sigh. https://dblp.org/rec/conf/icse/JaffeL01 https://dblp.org/rec/conf/icse/JaffeL01
- jimktrains2 7y agoDo you have a link to a readable copy?
- crocal 7y agoWhoops. Wrong article. Apologies. Blame my big fingers on this phone. The article is: « Analysis of Faults in an N-Version Software Experiment." Susan S. Brilliant, John C. Knight, Nancy G. Leveson (1990) And I found a PDF version here: http://sunnyday.mit.edu/papers/nver2.pdf http://sunnyday.mit.edu/papers/nver2.pdf
- yborg 7y ago>1990 study >student programmers >toy problem To me, this proves very little about how independent development teams would perform with experienced engineers working with present day design tools. It's possible that the results would be the same, worse, or better, but I hope this isn't the most recent research that's been done on this. The main issue with this approach is of course - cost. So the question is whether having two systems with uncorrelated fault paths is worth doubling the development expense.
- crocal 7y ago(Same disclaimer) > What's changing now, is that those 3 computers are now no longer the same architecture I think / hope this idea will die a well deserved death. Rail systems must face 25 years lifetime and with such design obsolescence headaches are multiplied by 3. In addition this creates bugs and integration nightmares. And on top of all that, we have known for decades that CPUs can be protected against systemic failures using the vital coded processor technique [1]. Note to the curious: this is one of the most fascinating piece of software I have ever encountered. Make the software resilient to any hardware fault through the power of arithmetic... [1] https://www.semanticscholar.org/paper/Vital-software%3A-Formal-method-and-coded-processor-Dollé/3a4a1645e672353c49d2c41718fe010c5fa2405b https://www.semanticscholar.org/paper/Vital-software%3A-Form...
- crocal 7y agoVoters are redundant too. Actuators receive safety release authorizations from all of them and will activate as long as one valid authorization is received.