13 ms·
A single line of code brought down a half-billion euro rocket launch
- jakeinspace 4y agoI’m very happy that the flight software codebase I’m currently working on doesn’t use any floating point. We don’t even have FPUs enabled. Then again, it’s not GN&C do the stakes are not as high.
- elashri 4y agoI think this story about code is to software engineering what the Nokia story is to business
- thedg 4y agodid you know that one time an integer overflow caused 184 billion bitcoins to be minted
- pugworthy 4y agoSeptember 12-13. An oceanographic research ship is doing gravity and magnetic surveys off the coast of Brazil. Suddenly, data acquisition software crash! byte day_of_year;
- squarefoot 4y agoWas by any chance the programmer an alien from either Mercury or Venus?:)
- Kukumber 4y agoThe world is not ready for the "Epochalypse", let's see what will collapse first, our civilization or our computers https://en.wikipedia.org/wiki/Year_2038_problem https://en.wikipedia.org/wiki/Year_2038_problem
- thedg 4y agoTIL
- nl 4y agoThe world is mostly ready for the 2038 problem because we have mostly transitioned to 64bit machines where it isn't a problem.
- ozarker 4y agoThere’s a lot of 32bit machines still around. There’s going to be a lot of problems in a lot of unexpected places
- serf 4y agoY2K bug affected a lot of machines, yet the fear of the impact beforehand had a larger effect on society-at-large than any of the software consequences during the event.
- actionfromafar 4y agoIndeed, but so very few things were automated back then compared to now.
- throwaway-blaze 4y agoSure, but we got thru y2k pretty unscathed. You don't think we can patch all the critical things in the next 15 years?
- actionfromafar 4y agoTell that to all IoT things and PLCs out there. The *D* in IoT stands for Date.
- phkahler 4y ago>> However, the reading is larger than the biggest possible 16 bit integer, a conversion is tried and fails. Usually, a well-designed system would have a procedure built-in to handle an overflow error and send a sensible message to the main computer. This, however, wasn’t one of those cases. This is so unbelievably untrue. I've never seen code anywhere that waits to fail before doing the right thing. This is exactly why I think exceptions are mostly useless, someone has to anticipate the problem, so why not write something that works right the first time. There are cases where exceptions can happen, but I don't think floating point arithmetic should be considered one of those cases.
- testrun 4y ago>>so why not write something that works right the first time Nobody writes 100% correct code. Ever.
- rvnx 4y ago"Remote computer is not answering" What is the correct code then ? In case there is no exception.
- WalterBright 4y agoSpeak for yourself.
- M4R5H4LL 4y agoThat's an untested assumption :)
- RantyDave 4y agoThey might do, we should test for that. These stories of gnat-brings-down-empire are, to me, always missing the point. There are going to be bugs. The hard part is in creating management around the actual software engineering such that this is not a problem.
- phkahler 4y agoThat's right. But nobody puts their best code in a generic exception handler either.
- dang 4y agoRelated: Ariane 5 – Flight 501 Failure (1996) - https://news.ycombinator.com/item?id=30556601 https://news.ycombinator.com/item?id=30556601 - March 2022 (1 comment) An overflow error costing 500M dollars (1996) - https://news.ycombinator.com/item?id=18939625 https://news.ycombinator.com/item?id=18939625 - Jan 2019 (20 comments) The Explosion of the Ariane 5 - https://news.ycombinator.com/item?id=5331474 https://news.ycombinator.com/item?id=5331474 - March 2013 (58 comments) There must be others. Anyone?
- raldi 4y agoThis is like saying a single little match blew up a building, neglecting to mention the garage full of oily rags and gasoline cans. The one line of code was the spark, yes, but the catastrophic consequences were due to a series of poorly-designed failsafes and insufficient testing.
- probablypower 4y ago> "A Single Line of Code Brought Down a Half-Billion Euro Rocket Launch" Blaming a system failure on a single point like this dooms that system to repeat similar failures (albeit in another element) in the future. There are numerous testing, quality and risk controls that could've been in place. There are probably even a few people who didn't do their job (besides the one person a decade ago who wrote the 'single line'). The point isn't to pin blame on any one point, but to look at the system (people, processes, technology) and try to understand why the system is fragile enough that a single person's error is able to escalate into a half-billion euro error. By focusing in on the point of failure, you end up falling victim to survivorship bias [0]. It is how you end up with developer teams swamped with unit-testing requirements and test coverage metrics, but still somehow end up with errors that impact the end-user anyway. It is how you get company surveys that always seem to miss the point, saying that the measures they implemented to improve company culture worked, yet everyone is burning out and miserable. [0] - https://en.wikipedia.org/wiki/Survivorship_bias https://en.wikipedia.org/wiki/Survivorship_bias
- naasking 4y agoThey're not focusing on one line of code, they cover the failover systems that also failed as well. It's also a mistake to try and fix bad tools, languages and programming practices with higher level processes. Just use better tools that do bounds checking (and unit checking which has also caused failures), preferably checked at compile-time and the problem is fixed without all the rigmarole you describe.
- atonse 4y agoAlso remember this happened in the late 90s which means the code was written in the 80s.
- fsckboy 4y agosounds like you're saying that being unable to write in-bounds code is the single-point-of-logic failure for a coder, and if you correct that part of the bad tool, all their other algos will be great... I think people who can write in-bounds code and type correct code with no safety rails have a leg up to write really good code.
- deleted 4y ago[deleted]
- photochemsyn 4y ago> "To achieve this, the guidance system converts the velocity readings, from 64 bit floating point to 16 bit signed integer". Oh, excellent possible interview question? "Write some code that reliably converts the full range of possible 64 bit floating point values to a 16 bit signed integer. What are the issues you'll have to deal with and what edge cases might arise?"
- l33t233372 4y agoIf I had this question and I couldn’t look up IEEE 754 I would be cooked. Maybe I have no business writing floating point code.
- yellow_postit 4y agoOr it’s a bad interview question as posed because in practice you could look up the spec and a better question would be to probe how the candidate would approach the problem and what clarify questions they would ask.
- touisteur 4y agoAnd most likely any serious Ada codebase worked by people worrying about such a question, has this in a generics or done (correctly) 7000 times... 'I'd just instantiate the Quantizer or Quantization_Manager generics, y'all have one, right?'
- gowld 4y agoIt's a bad question because it's completely impossible, so the interviewee has to guess if the interviewer knew that and looking for pushback, or didn't know that and you have to figure out what they actually meant and what they incorrectly think the answer is supposed to be.
- vba616 4y agoI'd be like: Step 1: Enumerate all possible inputs. <--- this is the important part Step 2: Map each input to something in the output domain. Step 3: Are you using Ada??? C++ is *right out*. ...it looks like the proper name for what is needed is a "non-injective surjective total function".
- enraged_camel 4y ago>> The cause? A simple, and very much avoidable coding bug, from a piece of dead code, left over from the previous Ariane 4 mission, which started nearly a decade before. >> The worst part? The code wasn’t necessary after takeoff, it was only part of the launch pad alignment process. But sometimes a trivial glitch might delay a launch by a few seconds and, in trying to save having to reset the whole system, the original software engineers decided that the sequence of code should run for an extra… 40 seconds after the scheduled liftoff. The author appears to be using a different definition of "dead code" than I'm used to. To me, dead code is code that is no longer called by anything else, and has no chance of running. Maybe a more accurate term is "legacy code"?
- WalterBright 4y ago> The system is designed to have a backup, standby system, which unfortunately, runs the exact same code. At Boeing, the backup system runs on a different CPU architecture, with a different program design, a different programming language, and a different team that isn't allowed to talk with the team on the other path.
- mrtksn 4y agoVery interesting, considering what happened with MCAS.
- peterhunt 4y agoThis is super interesting. Is there anywhere I can read more about this?
- sumthinprofound 4y agohttps://militaryembedded.com/avionics/computers/optimizing-reliability-dissimilar-redundant-architectures https://militaryembedded.com/avionics/computers/optimizing-r...
- WalterBright 4y agohttps://www.digitalmars.com/articles/b39.html https://www.digitalmars.com/articles/b39.html https://www.digitalmars.com/articles/b40.html https://www.digitalmars.com/articles/b40.html
- namaria 4y agoThen they go and make a single failure prone sensor capable of overriding pilots' control of flight surfaces.
- WalterBright 4y agoActually, the electric thumb trim switches overrode MCAS. This is why the Lion Air crew managed to restore normal trim 25 times. The Egypt Air did twice.
- xwdv 4y agoI hate these “single line of code did X” type headlines. It will always be a single line of code. The nature of most programs is to execute commands in a sequence. Eventually you hit one that fails. Hell, you could reduce it to even be less than a line of code. It could be a single variable. A single instruction. It could be a couple bits. A couple bad 1’s and 0’s in memory blew up a multibillion dollar rocket launch.
- opportune 4y agoI have found that among software engineers, it is surprisingly not common knowledge that floating point operations have all these sharp edges and gotchas. The most common situation in which it crops up is when dealing with quantities that require fractional units/arithmetic of some commonly discrete unit of measure. For example, you implement some complex logic to do request sampling, and in your binary you convert the total number of active requests to a float, add some stuff, divide some stuff, add some more stuff, multiply it again, then convert back to an int something like “number of requests that should be sampled.” Because floating point operations are non-associative, non-distributive, and commonly introduce remainder artifacts, you can end up with results like sampling 1 more request than there are total requests active, even when the arithmetic itself seems like that should be impossible. This is also common when dealing with time, although typically the outcome is not that bad. Despite time having a simple workaround of just changing the unit of measure (eg using milliseconds instead of seconds) and using int operations on that, because people don’t know why they shouldn’t use floating point operations in this case, they don’t always reach for it. The worst is when some complicated operation is done to report a float (or int converted from a float) as a metric. In the request sampling example, that would likely be noticed quickly and fixed. But when the float value looks reasonable enough and doesn’t violate some kind of system invariant, it can feed you bad data for a very long time before someone catches it.
- mooooooooooooo 4y agoWould you happen to have any resources on how to treat floating point values? I noticed some odd behaviour recently when using ruby to save to Postgres where the handoff between the two systems introduced imprecision in the saved value. Didn’t get to dig into it because it wasn’t a priority but it’s definitely an annoying unanswered question.
- cbhl 4y agoIf you want a thorough understanding: you'll want to look up "numerical computation", "numerical methods" or "computational methods"; in particular computing error bounds and error propagation. Typically covered in bachelor's level university-level Math, CS, or Engineering departments. If you just want to fix the odd behavior: adjust the schema so that you only work with whole numbers. For example, in the database schema, you can use DECIMAL instead of REAL/DOUBLE columns, or use two columns to specify the ratio of two integers (for example num/denom in frame rates in various video containers/codecs). In the application code: work only in whole numbers (e.g. cents, satoshis) instead of fractionals, using bigint or string types as applicable instead of e.g. double.
- jasontedor 4y agoI think James Gleick had a much better write up about this: http://www.maths.mic.ul.ie/posullivan/A%20Bug%20and%20a%20Crash%20by%20James%20Gleick.htm http://www.maths.mic.ul.ie/posullivan/A%20Bug%20and%20a%20Cr...
- twawaaay 4y agoWhy would the program react like that to a SINGLE wrong signal that disagrees with everything else and produce a signal that cannot do anything good in any circumstances? This just smells like a truly naive piece of implementation. There should be layers upon layers of safeties to prevent this dumb thing from happening. The computer should know the position, orientation and velocity of the rocket at any point in time and new signals should be interpreted in the context of what the computer already knows and in context of what other sensors are telling. It is not like the rocket can turn itself around in 1ms and if it does there probably isn't much it can do anyway. This suggests to me the problem is not the bug, it is the overall quality of development.
- amelius 4y agoI'm also wondering if they had a simulation environment, and why it didn't catch the problem.
- twawaaay 4y agoPretty much. The trajectory is preplanned anyway. I would expect nothing less than every launch to be ran through simulated environment multiple times if only to catch wrong launch information. Also this FP to integer does not smell any better. Anybody who's been interested in programming for any length of time will learn this is just a bad idea asking for trouble. If I was tech lead for the project I would definitely make sure there is static analysis that prevents these kinds of conversions from happening.
- allenrb 4y agoAnd you'd think catching incorrect launch information would be an obvious focal point, and yet... https://spaceflightnow.com/2018/02/23/investigators-say-erroneous-navigation-input-led-ariane-5-rocket-off-course/ https://spaceflightnow.com/2018/02/23/investigators-say-erro...
- twawaaay 4y ago
- ekanna 4y ago[dead]
- jaynate 4y agoWow. It makes you wonder what other types of bugs could exist?
- simondotau 4y agoIt'd probably be more accurate to say that a technology environment which allowed any single line of code to cause catastrophic failure is what brought down the launch. Or a failure of sufficiently accurate testing brought down the launch.
- joshAg 4y ago> With 16-bit unsigned integers, you can store anything from 0 to 65,535. If you use the first bit to store a sign (positive/negative) and your 16-bit signed integer now covers everything from -32,768 to +32,767 (only 15 bits left for the actual number). Anything bigger than these values and you’ve run out of bits. That's, oh man, that's not how they're stored or how you should think of it. Don't think of it that way because if you think "oh 1 bit for sign" that implies the number representation has both a +0 and a -0 (which is the case for ieee 754 floats) that are bitwise different in at least the sign bit, which isn't the case for signed ints. Plus, if you have that double zero that comes from dedicating a bit to sign, then you can't represent 2^15 or -2^15, because you are instead representing -0 and +0. Except, you can represent -2^15, or -32,768, by their own prose. So there's either more than just 15 bits for negative numbers or there's not actually a "sign bit." Like, ok, sure, you don't want to explain the intricacies of 2's complement for this, but don't say there's a sign bit. Explain signed ints as a shifting the range of possible values to include negative and positive values. Something like > With 16-bit unsigned integers, you can store anything from 0 to 65,535. If you shift that range down so that 0 is in the middle of the range of values instead of the minimum and your 16-bit signed integer now covers everything from -32,768 to +32,767. Anything outside the range of these values and you’ve run out of bits.
- mturmon 4y ago> ...If you shift that range down so that 0 is in the middle of the range of values instead of the minimum... Not a downvoter, but: your concept of "shifting the range" is also misleading. In the source domain of 16-bit numbers, [0...65535] can be split into two sets: [0...32767] [32768...65535] The first set of numbers maps to [0...32767] in 2's complement. But the second interval maps to [-32768...-1]. So it's not just a "shift" of [0...65535] onto another range. There's a discontinuous jump going from 32767 to 32768 (or -1 to 0 if converting the other direction). And actually, we don't know if the processor used 2's complement or 1's complement -- if it was 1's complement, they would have a signed 0! I think they'd have to say "remapping" the range? On the whole, I think OP did about as well as you're going to do, given the audience.
- 4y ago
- motohagiography 4y agoMillions of lines of code was what brought down that rocket launch. Saying it was one line just absolves everyone else.
- aaron695 4y ago[dead]
- albert_e 4y agoHow many such errors may have happened on successful projects but never got post-mortemed because there was luckily no catastrophic failure
- monksy 4y agoWhat was their unit test coverage?
- martyvis 4y agoEven very earth tied machines suffer similar issues. I worked on what was known as a "hot leveller" computer, a PDP-11/73 at the steelworks I was employed at. It had something like 9 rolls (maybe 200mm in diameter) that would be applied to a very hot steel plate (maybe 10mm to 150mm thick) after it had been rolled from a maybe 300mm thick slab. The levellers job was to smooth out any waves that might have acquired during the rolling process - almost like a clothes iron. The gap between the rolls needed to be adjusted by hydraulically positioning backup rolls that are even able to bend those work rolls across their width (maybe 3000mm). As you are always intending apply a huge amount of force anyways, to achieve the desired results, it was a mix of metallurgical driven algorithms and hard limits to doing the "setup". While there was always an operator that had to accept the setup before the run, there was always the risk of hitting the machines surfaces too hard, and straining components, and maybe causing a prolonged as expensive outage. Obviously the biggest risks were when there were changes or even experiments by both engineers and metallurgists. It was fun times as a quite junior engineer, and think there were a few times when over zealous setups resulted in some big noises. But I don't think I broke anything fortunately.
- tpoacher 4y agoHow can you write a whole article about "a single line of code" and not have that line appear anywhere in the article? Even worse, why was I completely unsurprised, nay expecting this to be the case when I clicked? (to the article's credit, it didn't quite start in the typical "George was walking his dog home when he noticed something wrong" fashion...)
- aaron695 4y ago[dead]