17 ms·
Resources about programming practices for writing safety-critical software
- Jtsummers 10y agoI don't have all my resources on hand right now, but off the top of my head this book should be added: https://mitpress.mit.edu/books/engineering-safer-world https://mitpress.mit.edu/books/engineering-safer-world This list is barely scratching the surface of safety-critical system engineering, but it's a start.
- stanislaw 10y agoThanks for the link. The book has been added.
- Jtsummers 10y agoNP. I'll try to find more when I get home, no promises though, pretty busy these days with life crap and can't do as much there as I'd like on the technical side (day job doesn't give me time to focus on these issues either).
- jonahx 10y agoI'm halfway through this, and not only is the theory insightful and often unexpected, but it's incredibly engaging, incredibly so for such an academic work.
- kqr2 10y agoThere is a draft version that is freely available: http://sunnyday.mit.edu/safer-world.pdf http://sunnyday.mit.edu/safer-world.pdf
- kevinr 10y agoParent's MIT Press link has a link to the final PDF, down the page on the left.
- watwut 10y agoThat is awesome, thank you.
- danaliv 10y agoDO-178B has been replaced by DO-178C.
- RaiO 10y agoIs there anything like this that specifically addresses reliability in a critical (but not "safety-critical") system?
- jacquesm 10y agoYes, Armstrong's thesis is a very good starting point: http://erlang.org/download/armstrong_thesis_2003.pdf http://erlang.org/download/armstrong_thesis_2003.pdf
- macintux 10y agoAn interesting paper that effectively describes the hardware equivalent to Erlang is Jim Gray's "Why Do Computers Stop and What Can Be Done About It?" http://www.hpl.hp.com/techreports/tandem/TR-85.7.pdf http://www.hpl.hp.com/techreports/tandem/TR-85.7.pdf
- swah 10y agoOther than the latest MISRA, I really enjoyed "Better Embedded System Software" by Phil Koopman. Ideally you should read it before starting your project, since it deal with the product specification/gathering requirements phase, which is your starting point in safety critical systems. [1] https://betterembsw.blogspot.com.br/2010/05/test-post.html https://betterembsw.blogspot.com.br/2010/05/test-post.html
- partycoder 10y agoI have read the JSF standard. I learned a lot from reading it. However, the JSF project has been reported to have lots of software defects.
- vonmoltke 10y agoDon't blame the tools, blame the carpenters (and the customers, and the customer's bosses).
- hackuser 10y ago> the JSF project has been reported to have lots of software defects I haven't read anything that differentiates between these two possible scenarios: 1) Poor engineering, execution, etc. 2) The bugs expected in this software project. When I think of it this way, I'm amazed it ever was completed (but maybe I'm thinking about it the wrong way): * Meet the specifications of not only three U.S. military services but also militaries and other entities in multiple national governments (with all the politics, compromise and complexity that involves). * Invent and implement technologies to provide capabilities so bleeding edge that few people will imagine some of them for years, if not decades. There are no prior designs; nothing like it has ever been done. Part of the point is to exceed competitors' engineering capabilities by as much as possible. * Integrate these technologies into a massive system of systems, arguably the most complex system in the history of humankind. * The system is human-rated. * Performance is the highest priority; there is no making easy compromises of performance for safety: Human lives, the outcomes of battles, the fates of nations, and the course of history may depend on performance. * Accomplish this in secret, greatly restricting your access to outside resources. Will this work? You can't publish a paper and get feedback, or make a presentation at a conference. * Accomplish this in coordination with thousands of suppliers in many countries. * Because it's hardware and very expensive, your ability to iterate is limited. My completely amateur guess based on the above is that it's a massive, decades-long waterfall-style project.
- nickpsecurity 10y ago"Integrate these technologies into a massive system of systems, arguably the most complex system in the history of humankind." You went a bit overboard there. There's plenty of systems probably more complex that work fine on a daily basis. They were usually designed centrally, though.
- ctz 10y agoThe obsession with C/C++ here is really weird. Like, take the MCO failure. That's a classic, textbook problem that can be structurally guaranteed not to happen with use of even a basic type system. It should be literally impossible to confuse values of different types/units/dimensions like this in something described as "safety-critical". It seems like all the resources here are concerned with trying to whittle C/C++ into an appropriate choice of tool, rather than choosing a different tool. It seems like a 1980s-1990s mindset.
- banachtarski 10y agoMCO?
- FigmentEngine 10y agoMars Climate Orbiter http://sunnyday.mit.edu/accidents/MCO_report.pdf http://sunnyday.mit.edu/accidents/MCO_report.pdf
- endorphone 10y agoHow would a basic type system protect against incorrectly interpreting an imperial floating point value as a metric floating point value? That seems like an especially weak example, and fundamentally falls under the realm of logical fault endemic of every possible programming language. There are legitimate gripes about C/C++, especially in a space with hostile actors an unknown inputs, but that example was particularly weak.
- joshmarlow 10y ago> How would a basic type system protect against incorrectly interpreting an imperial floating point value as a metric floating point value? You could wrap your value in a typed data structure: enum LengthUnit { Feet(f64), Meters(f64), } And provide conversion functions between them, then only operate on one type, `LengthUnit::Meters` and throw errors if `LengthUnit::Feet` is passed in. I'm using Rust syntax here because it's fresher on my mind, but you could do the same with Haskell, OCaml, F#, etc. IIRC, OCaml would optimize away the outer structure so you wouldn't have much/if any performance hit. Presumably compilers for the other languages could/would do the same. EDIT: for formatting and clarity.
- throwme_1980 10y agoc++ is not considered safe for any RTOS system, in fact you won't find it used in Aviation embedded devices (referring to the big 3 ) Tools yes, you can higher level languages to your heart's content.
- vonmoltke 10y agoHuh? The F-22A, F-35, P-8, and P-3 are all flying C++ code. Those are just the programs I have personally touched (not necessarily the code, though). Where did you get the idea that it "is not considered safe for any [real-time] system"?
- jordanb 10y agoThe F-22 avionics are mostly Ada. F-35 is mostly C++ though. If you want a good face-palm go read up on the "decision-makers" advocating making the switch. I saw a quote by a general blaming the F-22's cost overruns on the "Ada Operating System". But don't worry, C++ is a "COTS Industry Standard" so you can bet there were no overruns on the F-35. /s
- pjmlp 10y agoThe cost of the F-35 program shows how good that decision has been.
- vonmoltke 10y ago> The F-22 avionics are mostly Ada. Yes, including the piece I directly touched. It was deemed to minor to be worth the cost of rewriting.. Hell, when I left the program we still had an arthritic VAX to build on, should we ever need to rebuild the code. As for your last line, see my comment elsewhere about blaming the carpenters instead of the tools. :-)
- unboxed_type 10y agoAs far as I understand all vehicles you mentioned are not subject to DO-178 regulation. If they were then C++ code would have been much less likely be used. It is because C++ code is much more difficult to prove correct either using 100% test case coverege (DO-178B) or thru formal methods (DO-178C). Please correct me if I am wrong, I am doing a research on a similar topic.
- phelmig 10y agoDoes anyone know how software quality is handled in complex supply chains, e.g. automotive? From my point of view software is a 2nd grade citizen in areas dominated by manufacturing and classical engineering. I guess testing an over-the-update for a car that was build by ann OEM and thousands of suppliers must be quite a task.
- Jtsummers 10y agoIt's getting better, but hardware companies tend to view software as second-class. They think it's "easy", though they're finally accepting that it's not. It's taken decades of fatalities, cost overruns, and missed deadlines for them to realize this, but they're realizing it.
- nerdponx 10y ago> fatalities If someone dies because of a preventable bug in your software, shouldn't that be considered manslaughter? Obviously you formed a corporation in order to shield yourself from legal action (among other things). Fine, so you personally don't get charged with manslaughter. But in that case the corporation should be charged, and if convicted should be sentenced to the corporate equivalent of 25 years in jail. That would be a strong enough incentive to care about software. Of course, it never works like that in real life. Is this as ridiculous as it sounds to me or is my outrage misplaced somehow?
- Jtsummers 10y agoIt's as ridiculous as it sounds, but they go through a lot of effort on the corporate side to make sure they're in the clear. It's a lot of CYA paperwork and stuff demonstrating they've done what they could have. And then out-of-court civil suit settlements that are sealed so no one knows the details and can't form class action suits or coordinate well enough to initiate a criminal investigation (their family member's accident seems like a one-off to them, they don't know the extent of the problems).
- pjmlp 10y agoYou are completely right. As Hoare so elegantly described at his Turing award speech, regarding Algol compilers, back in 1981. "Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to--they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law. " Full speech here: http://www.labouseur.com/projects/codeReckon/papers/The-Emperors-Old-Clothes.pdf http://www.labouseur.com/projects/codeReckon/papers/The-Empe...
- yeslibertarian 10y agohopefully in a future not so far away, most safety-critical code will be formally verified, like http://sel4.systems/ http://sel4.systems/ for example
- kevinr 10y agoCode like the Boeing 787's avionics package gets one better: the spec specifies what the register values should be after each step of execution, and there's a company which takes the code, puts the processor in single-step mode, and checks.
- mrlyc 10y agoIn addition to MISRA, I've found the safety checklist in Lutz's "Targeting Safety-Related Errors During Software Requirements Analysis" at https://trs.jpl.nasa.gov/bitstream/handle/2014/35179/93-0749.pdf https://trs.jpl.nasa.gov/bitstream/handle/2014/35179/93-0749... to be very useful.