24 ms·
The Therac-25 Incident (2021)
- voxadam 1y ago(2021)
- deleted 1y ago[deleted]
- siva7 1y ago> With AECL's continued failure to explain how to test their device They can't. There was a single developer, he left, no tests existed, no one understood the mess to confidently make changes. At this point you can either lie your way through the regulators or scrap the product altogether. I've seen this kind of devs and companies running their software in regulated industries like in the therac incident, just now we are in the year 2025. I left because i understood that it's a criminal charge waiting to happen.
- snkline 1y agoI was kinda shocked by the results of his informal survey, because this was a big focus of my ethics course in college. I guess a lot of developers either didn't get a CS degree, or their degree program didn't involve an ethics course.
- ChrisMarshallNY 1y agoI worked for hardware manufacturers for most of my career, as a software guy. In my experience, hardware people really dis software. It's hard to get them to take it seriously. When something like this happens, they tend to double down on shading software. I have found it very, very difficult to get hardware people to understand that software has a different ruleset and workflow, from hardware. They interpret this as "cowboy software," and think we're trying to weasel out of structure.
- scottLobster 1y ago<deleted - wrong thread>
- ChrisMarshallNY 1y agoI think that was the F-35 incident (in Alaska). The Therac-25 incident was a radiation overdose in Texas.
- scottLobster 1y agoWhoops, wrong thread :P
- ChrisMarshallNY 1y agoBeen there, done that. Cheers
- kccqzy 1y agoHardware designers benefit from having multiple separate teams to test their product. A chip designer can rely on at least two other teams to test the designed chip, and one of them will be using formal verification. If software also has long release cycles and high cost to remedy mistakes, you bet we would also have multiple testers. In fact that was what happened in the 90s with shrink wrapped software and without easy updates.
- ChrisMarshallNY 1y agoThat was the case for the company that I worked at. The official QA organization was very powerful, and had no compunctions about stopping an entire product line, for one bug. When that happened, the department responsible for the bug would find themselves against the wall. As a result, all the software departments had pretty big teams of testers, who would validate the software, before it was released to the purview of the QA organization. It could be pretty restricting, but we always felt confident that what we shipped, worked.
- rendaw 1y agoSo reading about this my current company sounds exactly the same. And the one before it, and the one before that. Critical issues happen with customers, blame gets shifted, a useless fix is proposed in the post mortem and implemented (add another alert to the waterfall of useless alerts we get on call), and we continue to do ineffective testing. Procedural improvements are rejected by the original authors who were then promoted and want to keep feeling like they made something good and are now in a position to enforce that fiction. So IMO the lesson here isn't that everyone should focus on culture and process, it's that you won't have the right culture and process and (apparently) laws and regulation can overcome the lack of culture and process.
- MerrimanInd 1y agoEvery mechanical engineer educated in the USA knows the name of two famous collapses: the Tacoma Narrows Bridge and the Hyatt Regency balcony in Kansas City, MO. With an engineering ethics class being part of nearly every undergrad curriculum, these are two of the classic examples for us. I'm curious; do software engineers learn stories like the Therac-25 in their degrees?
- scottLobster 1y agoI was a Computer Engineer, so not quite the same, but we got taught about Therac-25 in our Engineering Ethics class when I took it over a decade ago. Unfortunately Computer Science is still in its too-cool-for-school phase, see OpenAI being sued over recently encouraging a suicidal teenager to kill themself. You'd think it would be common sense for that to be a hard stop outside of the LLM processing the moment a conversation turns to subjects like that, but nope.
- NoSalt 1y agoIs there a way to get the "gist" of the article, the lesson to be learned without reading the full article? I got to the screaming part and couldn't read any more.
- SirMaster 1y agoThe question I have is why was the hardware capable of delivering a fatal dose like this. Is that actually ever even a usable output for some legitimate reason? If not, why not hardware limit the power input to the machine, so even if the software completely failed, it would not be physically capable of delivering a fatal dose like this?
- PokestarFan 1y agoI believe that for X-ray mode, the radiation was indirect, so it needed a lot more power. Furthermore, older revisions had hardware locks, and the intent of the Therac-25 was to make it cheaper.
- salynchnew 1y agoGreat WTYP episode on this: https://www.youtube.com/watch?v=7EQT1gVsE6I https://www.youtube.com/watch?v=7EQT1gVsE6I
- koverstreet 1y agoA lot of people draw the wrong conclusions from Therac-25 today; becoming overly process driven can become a huge problem for software quality, because the processes have to be the right processes, and once processes are in place people have a natural tendency to defer to them and suspend their own judgement. That gets actively dangerous; a lot of more recent safety mishaps are more of the variety of "processes were followed, but things went hilariously off the rails and no one noticed and spoke up". Culture and expertise matter just as much if not more, especially today now that we all (in theory) should understand source control, testing, safer languages, etc. I think Admiral Rickover's methods apply just as much today, and applying that kind of thinking would fill major gaps in a lot of organizations - he emphasized good communication, a sense of responsibility, and thinking on your feet, and his safety record is unmatched. I think aviation also approaches process a bit better - by having much of it be more informal, less rigid checklists, it doesn't encourage people to suspend judgement so much. There's also the Tankship Tromedy, which really emphasizes the engineering legwork of just chasing down, understanding and fixing every last failure mode you can find. https://www.dieselduck.info/library/08%20policies/2006%20The%20Tankership%20Tromedy.pdf https://www.dieselduck.info/library/08%20policies/2006%20The...
- cantrevealname 1y ago> All of this software, from the individual processes to the OS itself, were the work of a single software developer. They left AECL in 1986, and no one has ever revealed their identity. I bet some readers are thinking that the developer that caused this tragedy retired with the millions he earned, maybe sailed his yacht to his Caribbean mansion. But the $300K FAANG salaries and multi-million stock options for senior developers represents the last decade or two. In the 1980's, developers were paid poorly and commanded little respect. The heroes in tech companies that sold expensive devices were the salesmen back then. The commission on the sale of a single Therac-25 probably exceeded the developer's salary. All of the following would indicate that this developer, no matter how senior or capable, was still a low-paid schlub: - It's Canada, so automatically 20% lower salaries than in the U.S. (AECL is in Canada, so it's a good bet that the developer was Canadian.) - It's the 1980s, so pre-web, pre-smartphones, pre-Google/Amazon, and developers had little recognition and low demand. - It's government, known to pay poorly for developers. (AECL is a government-owned corporation.) - It's mostly embedded software. Even though embedded software can be incredibly complex and life-critical, it's the least visible, so it's among the lower paid areas of software engineering (even today). For 1986, I would put his salary at $30-50K Canadian, or converted to U.S. dollars at that time would be $26-43K U.S., and inflation adjusted would be $78-129K U.S. today. And no stock options.
- Duanemclemore 1y agoThere's an excellent episode of Well There's Your Problem about Therac-25. https://youtu.be/7EQT1gVsE6I https://youtu.be/7EQT1gVsE6I
- MarkusWandel 1y agoIn a quick skim of the comments so far, I don't see the real smoking gun. The previous devices had hardware interlocks. So if the software glitched, it was just an annoying glitch - nobody got zapped. But mature software gets trusted, so they removed the hardware interlock as redundant. And then the annoying glitches became fatal. Total miscommunication. The people cost-reducing the hardware interlock only saw mature, trustworthy software. The people living with the glitches only saw them as annoying, but harmless. And then, disaster.
- zackmorris 1y agoOur power went off a couple off weeks ago due to wind probably knocking a branch into a power line. Now our Frigidaire microwave runs with the door open. Supposedly there are mechanical switches that prevent that, but evidently "modern" microwaves can control the gun through the logic board. The engineering failures that led to this, from conceptual to design to internal control, boggle my mind. I'm not even sure where to send a complaint or if it would result in any kind of compensation. Because billion dollar corporations know that they'll never have to face any kind of corporate death penalty because they're protected by limited liability. So we'll just buy another $150 microwave instead. Are smaller companies better at engineering safety? Evidently not.
- bongodongobob 1y agoI have a microwave from the early 80s. If I stick a pencil in the door latch I can get it to run with the door open as well. It's not the demon core. Just don't stick your head in it when it's running.
- simoncion 1y ago> If I stick a pencil in the door latch I can get it to run with the door open as well. "The safety interlocks don't work when the operator intentionally goes out of his way to defeat them." isn't a concern. There's only so much you can do to prevent someone who's dedicated to disabling them. "The safety interlocks fail dangerous because of an unexpected power cut." is a huge concern. What else did the manufacturer skimp on, or -worse- simply fail to understand was important to do for the safety of the operator of the device?
- bongodongobob 1y agoIt's crazy to me that you are ruling out a power surge frying a board. Same thing could happen to the 80s model as well. You have not root caused it and are making up a failure mode that fits your point. Hell, the hall sensor could be fried and that's pretty damn mechanical. Again, your microwave isn't a demon core. Inverse square law applies. Don't put any limbs inside when it's on, and it really isn't that dangerous, so I'm not shocked they didn't apply aircraft safety design rules.
- w10-1 1y agoThis is not the example readers need to understand, because the failures were so rudimentary and systemic that it seems "good process" is the answer. Having written and validated both FDA and CLIA software, I'd suggest that process is never sufficient. Plenty of well-meaning people will create and follow incomplete plans and hand-wave away issues when they sign off -- particularly people who gravitate towards rule-based, formulaic work in a hierarchy. You need people both capable of and willing to seriously question whether proof is really proof, and who will stand up for some random patient in the distant future over their boss and colleagues on a deadline -- and yet they cannot be oppositional or egotistical, and must have deep insight into the subject matter. It's really, really hard to find those people.
- csours 1y agoTo me, the Therac incident is the poster child for a category I call 'context change error'. Some of the controls were 'born' in a world of hardware interlocks, and so the engineers used the frame of mind where hardware interlocks exist. Some time later, the interlocks were replaced with software controls. Since everything had worked before, all the software had to do was what worked before. But it is VERY difficult to challenge all of your assumptions about what "working" means. --- This is also a good reminder that work is done by people and teams, not corporations. That is - just because somebody knows the fine details, that does not mean that the corporation knows the fine details.
- fogzen 1y agoAlmost. It’s a process problem. But the process is a step above the organization. It’s a socio-economic process that incentivizes these problems. It’s capitalism that’s the process problem. That’s the process that introduces the problem into the organization. Without the government regulators making them test nothing would have even been done at all. Because the problem is the organization exists within a framework that pits it against safety. Safety is at odds with what the organization is tasked to do within the process that it exists in.
- darepublic 1y agosad story. gotta blame canada for this crap. The elements of this story.. hospitals.. a janky attempt at innovation.. passive aggressive denials from otherwise timid demure canadians. Cold grey bureacracy. It all reminds me of the not so great north. The technicians were sipping on their tim hortons slop at the time to make it perfect.
- onewheeltom 1y agoThe manufacturer of the Therac-25, AECL, did not share customer incident reports with other customers when patients were injured. So, the hospitals believed that their incidents were isolated. This may have been legal, but was highly unethical.
- smarks 1y agoI believe the definitive analysis of the Therac-25 incident was written by Nancy Leveson, first in IEEE Computer,[1] and later as an appendix of her book.[2] The appendix is freely available as a PDF on the web [3][4] and probably other places. Many people here are asking questions about what happened and how it came about. The answers to many of these questions can be found there. I strongly recommend that anyone who is serious about safety and wants to learn more about this incident read Leveson’s analysis. [1]: N. G. Leveson and C. S. Turner, "An investigation of the Therac-25 accidents," in Computer, vol. 26, no. 7, pp. 18-41, July 1993. [2]: Nancy Leveson. Safeware: System Safety and Computers. Addison-Wesley, 1995. [3]: http://sunnyday.mit.edu/papers/therac.pdf http://sunnyday.mit.edu/papers/therac.pdf [4]: https://web.mit.edu/6.033/2014/wwwdocs/papers/therac.pdf https://web.mit.edu/6.033/2014/wwwdocs/papers/therac.pdf
- MilyMason2 1y ago[flagged]
- rokkamokka 1y agoI was taught this incident in university many years ago. It's undeniably an important lesson that shouldn't be forgotten
- napolux 1y agoThe most deadly bug in history. If you know any other deadly bug, please share! I love these stories!
- NitpickLawyer 1y agoThe MCAS related bugs @ Boeing led to 300+ deaths, so it's probably a contender.
- solids 1y agoWas that a bug or a failure to inform pilots about a new system?
- AdamN 1y agoBoth - and really MCAS was fine but the issue was the metering systems (Pitot tubes) and the handling of conflicting data. That part of the puzzle was definitely a bug in the logic/software.
- kijin 1y agoRemember the Airbus that crashed in the middle of the Atlantic because one of the pilots kept pulling on his yoke, and the computer decided to average his input with normal input from the other pilot? Conflict resolution in redundant systems seems to be one of the weakest spots in modern aircraft software.
- sgerenser 1y agoAir France 447: https://en.m.wikipedia.org/wiki/Air_France_Flight_447 https://en.m.wikipedia.org/wiki/Air_France_Flight_447 Inputs were averaged, but supposedly there’s at least a warning: Confused, Bonin exclaimed, "I don't have control of the airplane any more now", and two seconds later, "I don't have control of the airplane at all!"[42] Robert responded to this by saying, "controls to the left", and took over control of the aircraft.[84][44] He pushed his side-stick forward to lower the nose and recover from the stall; however, Bonin was still pulling his side-stick back. The inputs cancelled each other out and triggered an audible "dual input" warning.
- deleted 1y ago[deleted]
- benrutter 1y ago> software quality doesn't appear because you have good developers. It's the end result of a process, and that process informs both your software development practices, but also your testing. Your management. Even your sales and servicing. If you only take one thing away from this article, it should be this one! The Therac-25 incident is a horrifying and important part of software history, it's really easy to think type-systems, unit-testing and defensive-coding can solve all software problems. They definitely can help a lot, but the real failure in the story of the Therac-25 from my understanding, is that it took far too long for incidents to be reported, investigated and fixed. There was a great Cautionary Tales podcast about the device recently[0], one thing mentioned was that, even aside from the catasrophic accidents, Therac-25 machines were routinely seen by users to show unexplained errors, but these issues never made it to the desk of someone who might fix it. [0] https://timharford.com/2025/07/cautionary-tales-captain-kirk-forgot-to-put-the-machine-on-stun-2/ https://timharford.com/2025/07/cautionary-tales-captain-kirk...
- ChrisMarshallNY 1y agoI worked for a company that manufactured some of the highest-Quality photographic and scientific equipment that you can buy. It was expensive as hell, but our customers seemed to think it was worth it. > It's the end result of a process In my experience, it's even more than that. It's a culture.
- franktankbank 1y agoA culture of high-quality engineering, no doubt. Made up of: high quality engineers!
- ChrisMarshallNY 1y agoYes, but some of them were the most stubborn bastards I've ever worked with.
- franktankbank 1y ago
- michaelt 1y agoI'd be interested in knowing how many of y'all are being taught about this sort of thing in college ethics/safety/reliability classes. I was taught about this in engineering school, as part of a general engineering course also covering things like bathtub reliability curves and how to calculate the number of redundant cooling pumps a nuclear power plant needs. But it's a long time since I was in college. Is this sort of thing still taught to engineers and developers in college these days?
- FuriouslyAdrift 1y agoA big thing that was emphasized in my computer engineering courses at Purdue in the early 90s with regards to machine interfaces was hysteresis. A machine has a RANGE of behaviors throughout it's operating area that might not be accounted fro in your programing and you must take that into consideration (i.e. a robotic arm or electric motor doesn't just 'stop' instantly). Analog systems do not behave like computers.
- ramses0 1y agoThe "IBM Black Team Debugs a Tape Drive" story comes to mind: https://www.penzba.co.uk/GreybeardStories/TheBlackTeam.html https://www.penzba.co.uk/GreybeardStories/TheBlackTeam.html
- mlnhd 1y agoThis and Tacoma Narrows are literally the only topics covered in engineering ethics, which itself is literally only a one hour presentation.
- InvisibleUp 1y agoDon’t forget the Hyatt Regency walkway, too.
- mrguyorama 1y agoThe therac-25 was just one of the many incidents we covered in my Software Ethics course for my Computer Science degree. The problem is not "we have to teach it", the problem is that at least half the talented people in the room with me in that class considered the entire thing "a joke" bullshit class that just wasted their time. You can't teach people to care.
- rvz 1y agoWe're more likely to get a similar incident like this very quickly if we continue with the cult of 'vibe-coding' and throwing away basic software engineering principles out of the window as I said before. [0] Take this post-mortem here [1] as a great warning and which also highlights exactly what could go horribly wrong if the LLM misreads comments. What's even more scarier is each time I stumble across a freshly minted project on GitHub with a considerable amount of attention, not only it is 99% vibe-coded (very easy to detect) but it completely lacks any tests written for it. Makes me question the ability of the user prompting the code in the first place if they even understand how to write robust and battle-tested software. [0] https://news.ycombinator.com/item?id=44764689 https://news.ycombinator.com/item?id=44764689 [1] https://sketch.dev/blog/our-first-outage-from-llm-written-code https://sketch.dev/blog/our-first-outage-from-llm-written-co...
- mrguyorama 1y agoGod that "post mortem" is such a portent of things to come. I've seen this exact problem path happen locally nearly any time I use claude. It very obviously just picks what it should put where based on weighted random chance, and that random chance is going to not go in your favor at some point, in a way that no amount of training or job experience can help with, because no, a human would not have made this mistake. This is the kind of mistake that fails people out of CS101; It's obvious that the student is just manipulating symbols they don't really "get" rather than modifying code. Throwing the chinese room thought experiment at your code base is bad engineering.
- voxadam 1y agoThe idea of 'vibe-coding' safety critical software is beyond terrifying. Timing and safety critical software is hard enough to talk about intelligently, even harder to code, harder yet to audit, and damn near impossible to debug, and all that's without neophyte code monkeys introducing massive black boxes full of poorly understood voodoo to the process.
- isopede 1y agoI strongly believe that we will see an incident akin to Therac-25 in the near future. With as many people running YOLO mode on their agents as there are, Claude or Gemini is going to be hooked up to some real hardware that will end up killing someone. Personally, I've found even the latest batch of agents fairly poor at embedded systems, and I shudder at the thought of giving them the keys to the kingdom to say... a radiation machine.
- throwawayoldie 1y agoThey killed "only" about 350 people combined, but the two fatal crashes of the Boeing 737 MAX in 2018 and 2019 were due to poor quality software: https://en.wikipedia.org/wiki/Maneuvering_Characteristics_Augmentation_System https://en.wikipedia.org/wiki/Maneuvering_Characteristics_Au...
- the-grump 1y agoThe 737 MAX MCAS debacle was one such failure, albeit involving a wider system failure and not purely software. Agreed on the future but I think we were headed there regardless.
- jonplackett 1y agoYeah reading this reminded me a lot of MCAS. Though MCAS was intentionally implemented and intentionally kept secret.
- Maxion 1y ago> Personally, I've found even the latest batch of agents fairly poor at embedded systems I mean even simple crud web apps where the data models are more complex, and where the same data has multiple structures, the LLMs get confused after the second data transformation (at the most). E.g. You take in data with field created_at, store it as created_on, and send it out to another system as last_modified.
- SCdF 1y agoThe Horizon (UK Royal Mail accounting software) incident killed multiple postmasters through suicide, and bankrupted and destroyed the lives of dozens or hundreds more. The core takeaway developers should have from Therac-25 is not that this happens just on "really important" software, but that all software is important, and all software can kill, and you need to always care.
- deleted 1y ago[deleted]
- autonomousErwin 1y agoThis reminds me of the Belgium 2003 election that was impossibly skewered by a supernova light years away sending charged particles which manage to get through our atmosphere (allegedly) and flipping a bit. Not the only case it's happened.
- jve 1y agoOn the bright side, wow, those computers are really sturdy: takes a whole supernova to just flip a bit :)
- kijin 1y agoWell the thing is, millions of stars go supernova in the observable universe every single day. Throw in the daily gamma ray burst as well, and you've got bit flips all over the place.
- haddonist 1y agoWell There's Your Problem podcast, Episode 121: Therac-25 https://www.youtube.com/watch?v=7EQT1gVsE6I https://www.youtube.com/watch?v=7EQT1gVsE6I
- dpacmittal 1y agoThere's also this video from Kyle Hill which is pretty good (I think it's a different incident though, not sure) - https://www.youtube.com/watch?v=Ap0orGCiou8 https://www.youtube.com/watch?v=Ap0orGCiou8
- voidUpdate 1y agoMy go-tos are usually Fascinating Horror https://www.youtube.com/watch?v=nU5HbUOtyqk https://www.youtube.com/watch?v=nU5HbUOtyqk and Plainly Difficult https://www.youtube.com/watch?v=-7gVqBY52MY https://www.youtube.com/watch?v=-7gVqBY52MY. I've gone off Kyle Hill after a lot of people pointed out that he was promoting a scam (BetterHelp) on his video about fraud and his response was just to tell people to deal with it
- auggierose 1y agoWondering if that "one developer" is here on HN.
- Forgret 1y agoHahaha, it would be interesting, maybe he just commented on the post here?
- mellosouls 1y agoTIL TheDailyWTF is still active. I'd thought it had settled to greatest hits only some years ago.
- greatgib 1y agoThis story is kind of old. But also I'm suspicious that this was an AI generated content due to this weird paragraph (one becoming "they"): It's worth noting that there was one developer who wrote all of this code. They left AECL in 1986, and thankfully for them, no one has ever revealed their identity. And while it may be tempting to lay the blame at their feet—they made every technical choice, they coded every bug—it would be wildly unfair to do that.
- remyporter 1y agoI’ve been writing on the Internet since very early days, and have put almost twenty years into The Daily Wtf specifically. Which means I’m actually over represented in the training set. I don’t write like AI. AI writes like me.
- lopespm 1y agoI am really amazed by the frequency and the quality of your output. Would you have an article on your routines, how you structure your day / work? Essentially, what enables your consistency, and quality articles?
- HankStallone 1y agoIt writes that way because almost everyone writes that way these days. It's annoying if you learned English grammar from textbooks and other materials written over 50 years ago, but it's extremely common now anyway. So a large chunk of its training data will be that way. It's interesting, because all the older works in its training data will default to the masculine singular, and that has to be a massive number of books too. But maybe the modern writing, including lots of online sources, simply overwhelms that. Or it's one of the guardrails written into the AIs to avoid offending people.
- vemv 1y agoMy (tragically) favorite part is, from wikipedia: > A commission attributed the primary cause to generally poor software design and development practices, rather than singling out specific coding errors. Which to me reads as "this entire codebase was so awful that it was bound to fail in some or other way".
- ycombobreaker 1y agoSibling reply notes the "process" is the problem, amd I would second that. I would also like to add, it's perfectly possible to produce a high quality code base with poor practices. This can happen with very small, expert teams. However, certain qualities become high-variance, which becomes a hazard over time.
- rgoulter 1y agoHmm. "poor software design" suggests a high risk that something might go wrong; "poor development practice" suggests that mistakes won't get caught/remedied. By focusing on particular errors, there's the possibility you'll think "problem solved". By focusing on process, you hope to catch mistakes as early as possible.
- DamonHD 1y agoWhen trying to make better systems in moderately-critical roles (investment banking, not medicine though) my approach was both try to understand and fix the immediate fault, but also find out if (and fix if so) any systemic issue that would make other related errors likely.
- rossant 1y agoThe first commenter on this site introduces himself as "a physician who did a computer science degree before medical school." He is now president of the Ray Helfer Society [1], "an honorary society of physicians seeking to provide medical leadership regarding the prevention, diagnosis, treatment and research concerning child abuse and neglect." While the cause is noble, the medical detection of child abuse faces serious issues with undetected and unacknowledged false positives [2], since ground truth is almost never knowable. The prevailing idea is that certain medical findings are considered proof beyond reasonable doubt of violent abuse, even without witnesses or confessions (denials are extremely common). These beliefs rest on decades of medical literature regarded by many as low quality because of methodological flaws, especially circular reasoning (patients are classified as abuse victims because they show certain medical findings, and then the same findings are found in nearly all those patients—which hardly proves anything [3]). I raise this point because, while not exactly software bugs, we are now seeing black-box AIs claiming to detect child abuse with supposedly very high accuracy, trained on decades of this flawed data [4, 5]. Flawed data can only produce flawed predictions (garbage in, garbage out). I am deeply concerned that misplaced confidence in medical software will reinforce wrongful determinations of child abuse, including both false positives (unjust allegations potentially leading to termination of parental rights, foster care placements, imprisonment of parents and caretakers) and false negatives (children who remain unprotected from ongoing abuse). [1] https://hs.memberclicks.net/executive-committee https://hs.memberclicks.net/executive-committee [2] https://news.ycombinator.com/item?id=37650402 https://news.ycombinator.com/item?id=37650402 [3] https://pubmed.ncbi.nlm.nih.gov/30146789/ https://pubmed.ncbi.nlm.nih.gov/30146789/ [4] https://rdcu.be/eCE3l https://rdcu.be/eCE3l [5] https://www.sciencedirect.com/science/article/pii/S0022346821002104 https://www.sciencedirect.com/science/article/pii/S002234682...
- elric 1y agoOne of the commenters on the article wrote this: > Throughout the 80s and 90s there was just a feeling in medicine that computers were dangerous <snip> This is why, when I was a resident in 2002-2006 we still were writing all of our orders and notes on paper. I was briefly part of an experiment with electronic patient records in an ICU in the early 2000s. My job was to basically babysit the server processing the records in the ICU. The entire staff hated the system. They hated having to switch to computers (this was many years pre-ipad and similarly sleek tablets) to check and update records. They were very much used to writing medications (what, when, which dose, etc) onto bedside charts, which were very easy to consult and very easy to update. Any kind of dataloss in those records could have fatal consequences. Any delay in getting to the information could be bad. This was *not* just a case of doctors having unfounded "feelings" that computers were dangerous. Computers were very much more dangerous than pen and paper. I haven't been involved in that industry since then, and I imagine things have gotten better since, but still worth keeping in mind.
- superjan 1y agoIt”s worthwhile to mention that in the US and EU EMRs are generally not considered Medical Devices and are therefore not subject to a lot of regulations. https://www.medicaleconomics.com/view/what-if-emrs-were-classified-as-medical-devices https://www.medicaleconomics.com/view/what-if-emrs-were-clas...
- jacquesm 1y agoNow we have Chipsoft, arguably one of the worst players in the entire IT space that has a near monopoly (around me, anyway) on IT for hospitals. They charge a fortune, produce crap software and the larger they get the less choice there is for the remainder. It is baffling to me that we should be enabling such hostile players.
- amelius 1y ago> The Therac-25 was the first entirely software-controlled radiotherapy device. This says it all.
- theglocksaint 1y agoDon't know how to break this to you but they all are nowadays.
- mdavid626 1y agoSome sanity checks are always a good idea before running such destructive action (IF beam_strength > REASONABLY_HIGH_NUMBER THEN error). Of course the UI bug is hard to catch, but the sanity check would have prevented this completely and the machine would just end up in an error, rather than killing patients.
- b_e_n_t_o_n 1y agoinvariants are so useful to enforce even for toy projects. they should never be triggered outside of dev, but if they do sometimes it's better to just let it crash.
- bzzzt 1y agoMaking sure the beam is off before crashing would be better though.
- b_e_n_t_o_n 1y agoFor sure :P
- linohh 1y agoIn my university this case was (and probably still is) subject of the first lecture in the first semester. A lot to learn here and one of the prime examples how the DEPOSE model [Perrow 1984] works for software engineering.
- Forgret 1y agoWhat surprised me most was that only one developer was working on such an unpredictable technology, whereas I think I need at least 5 developers to be able to discuss options.
- throwaway0261 1y agoOne of the benefits of regulations in these areas, is that they require proper tests and documentation. This often requires more than one person to handle the load. We don't want to go back to the 80s YOLO mode just because we need to "move faster". BTW: Relevant XKCD: https://xkcd.com/2347/ https://xkcd.com/2347/
- mnw21cam 1y agoThough this XKCD might be even more relevant: https://xkcd.com/2030/ https://xkcd.com/2030/
- OskarS 1y agoIt's interesting to compare this with the Post Office Scandal in the UK. Very different incidents, but reading this, there is arguably a root assumption in both cases that people made, which is that "the software can't be wrong". For developers, this is a hilariously silly thing, but for non-developers looking at it from the outside, they don't have the capability or training to understand that software can be this fragile. And they look at a situation like the post office scandal and think "Either this piece of software we paid millions for and was developed by a bunch of highly trained engineers is wrong, or these people are just ripping us off". Same thing with Therac-25, this software had worked on previous models and the rest of the company just had this unspoken assumption that it simply wasn't possible that there was anything wrong with it, so testing it specifically wasn't needed.
- ndsipa_pomu 1y agoI'd consider the Post Office Scandal to be far more malicious. The higher ups in the post office were getting bonuses IIRC according to how much money was "recovered" (defrauded) from the subpostmasters. Also there was a lot of lying to the courts and ministers about the reliability of the software. As far as I know, the Therac-25 incidents were reasonably honest mistakes.
- OskarS 1y agoI agree, that is very true, Therac-25 was incompetence, Post Office was incompetence with a heavy dose of malice. This aspect just steuck me as similar, the unquestioning belief in the infallibility of software.
- jwr 1y agoNo, this is not a "hilariously silly thing" for developers. In fact, I'd say that most developers place way too much trust in software. I am a developer and whatever software system I touch breaks horribly. When my family wants to use an ATM, they tell me to stand at a distance, so that my aura doesn't break things. This is why I will not get into a self-driving car in the foreseeable future — I think we place far too much confidence in these complex software systems. And yet I see that the overwhelming majority of HN readers are not only happy to be beta-testers for this software as participants in road traffic, but also are happy to get in those cars. They are OK with trusting their life to new, complex, poorly understood and poorly tested software systems, in spite of every other software system breaking and falling apart around them. [anticipating immediate common responses: 1) yes, I know that self-driving car companies claim that their cars are statistically safer than human drivers, this is beyond the point here. One, they are "safer" largely because they drive so badly that other road participants pay extra attention and accommodate their weirdness, and two, they are still new, complex and poorly understood systems. 2) "you already trust your life to software systems" — again, beyond the point, not quite true as many software systems are built to have human supervision and override capability (think airplanes), and others are built to strict engineering requirements (think brakes in cars) while self-driving cars are not built that way.]
- haunter 1y agoMy "favorite" part: >One failure occurred when a particular sequence of keystrokes was entered on the VT100 terminal that controlled the PDP-11 computer: If the operator were to press "X" to (erroneously) select 25 MeV photon mode, then use "cursor up" to edit the input to "E" to (correctly) select 25 MeV Electron mode, then "Enter", all within eight seconds of the first keypress and well within the capability of an experienced user of the machine, the edit would not be processed and an overdose could be administered. These edits were not noticed as it would take 8 seconds for startup, so it would go with the default setup Kinda reminds me how everything is touchscreen nowadays from car interfaces to industry critical software
- hiccuphippo 1y agoAnd we have a concept, optimistic updates, for making the ui look responsive while the updates happen in the background and reconcile later. I can only hope they know when not to use it.
- kevincox 1y agoOptimistic updates should almost always be paired with some sort of indicator showing if/when a value has actually been persisted. In practice this is rarely implement. Even failures are often not shown and value rolled back. (Making it a very optimistic update indeed)
- ramses0 1y agoTry quickly typing 1+ 2 + 3 into the iOS 11 Calculator (reddit.com) 886 points by danso on Oct 24, 2017 | hide | past | favorite | 480 comments https://news.ycombinator.com/item?id=15538666 https://news.ycombinator.com/item?id=15538666 ...this _exact_ same failure mode in a "less" critical domain (eg: literally your most frequently used "pocket calculator"), unless you're using the calculator for Important Things(tm).
- throwaway0261 1y agoOne of the comments said this: > That standard [IEC 62304] is surrounded by other technical reports and guidances recognized by the FDA, on software risk management, safety cases, software validation. And I can tell you that the FDA is very picky, when they review your software design and testing documentation. For the first version and for every design change. > That’s good news for all of us. An adverse event like the Therac 25 is very unlikely today. This is a case where regulation is a good thing. Unfortunately I see a trend lately where almost any regulation is seen as something stopping innovation and business growth. There are room for improvements and some areas are over regulated, but we don't want a "DOGE" chainsaw to regulations without knowing what the consequences are.
- softwaredoug 1y agoSafety problems are almost never about one evil / dumb person and frequently involve confusing lines of responsibility. Which makes me very nervous about AI generated code and people who don’t clam human authorship. The bug that creeps in where we scapegoat the AI isn’t gonna cut it in a safety situation.
- 0xDEAFBEAD 1y ago>any bugs we see would have to be transient bugs caused by radiation or hardware errors. Can't imagine that radiation might be a factor here...
- tedggh 1y agoTL;DR The Therac-25 was a radiation therapy machine built by Atomic Energy Canada Limited in the 1980s. It was the first to rely entirely on software for safety controls, with no hardware interlocks. Between 1985 and 1987, at least six patients received massive overdoses of radiation, some fatally, due to software flaws. One major case in March 1986 at the East Texas Cancer Center involved a technician who mistyped the treatment type, corrected it quickly, and started the beam. Because of a race condition, the correction didn’t fully register. Instead of the prescribed 180 rads, the patient was hit with up to 25,000 rads. The machine reported an underdose, so staff didn’t realize the harm until later. Other hospitals reported similar incidents, but AECL denied overdoses were possible. Their safety analysis assumed software could not fail. When the FDA investigated, AECL couldn’t produce proper test plans and issued crude fixes like telling hospitals to disable the “up arrow” key. The root problem was not a single bug but the absence of a rigorous process for safety-critical software. AECL relied on old code written by one developer and never built proper testing practices. The scandal eventually pushed regulators to tighten standards. The Therac-25 remains a case study of how poor software processes and organizational blind spots can kill—a warning echoed decades later by failures like the Boeing 737 MAX.
- thijson 1y agoI remember my computer science professor talking about this, how critical safety can be in software. Another example he gave was the refueling machine at a nuclear power plant, it had fell off the tracks and broke the pipe that goes through the reactor due to a software bug. Also he mentioned the software in his pacemaker. Engineers in other fields need to sign off on designs, and can be held liable if something goes wrong. Software hasn't caught up to that yet.
- Tenemo 1y agoThe full 1993 report linked in the article has an intetesting statement regarding software developer certfication in the "Lessons learned" chapter: > Taking a couple of programming courses or programming a home computer does not qualify anyone to produce safety-critical software. Although certification of software engineers is not yet required, more events like those associated with the Therac-25 will make such certification inevitable. There is activity in Britain to specify required courses for those working on critical software. Any engineer is not automatically qualified to be a software engineer — an extensive program of study and experience is required. Safety-critical software engineering requires training and experience in addition to that required for noncritical software. After 32 years, this didn't go the way the report's authors expected, right?
- slavik81 1y agoI am a licensed professional software engineer in Canada. It's been fifteen years since I first registered with my professional association, but I will probably not be renewing my license this year as it's not providing any real benefit to my career. Two decades ago there was a lot of talk about turning software development into a structured engineering discipline, but that plan seems to have largely been abandoned.
- mitthrowaway2 1y agoI've had some discussions with my engineering regulator in Canada. It's clear they have no idea what software engineering even is or who should be regulated or why. I tried to get them to provide some examples of what would and would not count as software engineering, but they couldn't.
- firesteelrain 1y agoTo add. Safety-critical software is not something you pick up in a classroom, it is something built over years of disciplined practice. There are standards like DO-178 for avionics and IEC 61508 for industrial systems, but how rigorously they are applied often depends on cost and project constraints. That said, when failures happen, the audit trail will not matter to the people harmed. The history of safety engineering shows that almost every rule exists because someone was hurt first.
- armcat 1y agoTherac-25 was part of the mandatory "computer ethics" course at my uni, as part of the Computer Science programme, circa early 2000s.