6 ms·
>This all has to happen without any human intervention, so the central computer software has been programmed and extensively tested to make sure all corrections
by JonoW 8y ago
>This all has to happen without any human intervention, so the central computer software has been programmed and extensively tested to make sure all corrections can be made on the fly.
I'd love to get some deeper insight into how NASA writes and tests software, I can only guess it's a million miles from how most of us work. Anyone know of any good talks, articles from engineers there?
- Aaargh20318 8y agoHere's an article about it I read a while back, interesting read: https://www.fastcompany.com/28121/they-write-right-stuff https://www.fastcompany.com/28121/they-write-right-stuff
- marsRoverDev 8y agoThis redundant software and hardware setup typically isn't necessary when humans aren't involved. The space shuttle system is similar to what you will find on a Boeing or Airbus aircraft. Redundant software, written by different people in different countries with completely different cultures in different languages (on purpose), running on multiple machines with different hardware and voting on the decisions to be made. It is complete overkill when "all" you're going to lose is a robot and some pride, as with a space probe you want to have lots of features and this level of safety is very restrictive on development effort. More than likely, the spacecraft in question is written in C or C++ with the help of RTEMS or VxWorks. It is probably running a radiation hardened, very slow processor.
- deleted 8y ago[deleted]
- rtkwe 8y agoThey don't do 3x calculations and voting but they do often have redundant computers they can switch over to in case of failure. Curiosity had to switch to it's 'B-side' computer back in 2013 when A-side had a memory issue. Even when not carrying humans it's still a billion/million dollar mission that probably wouldn't be replicated for a while if ever (within the researchers life times at least) that could be scuttled by a softwer bug. If anyone is interested JPL publishes their code standards doc for C: https://lars-lab.jpl.nasa.gov/JPL_Coding_Standard_C.pdf https://lars-lab.jpl.nasa.gov/JPL_Coding_Standard_C.pdf
- kevin_thibedeau 8y agoMost spacecraft have some form of redundancy to guard against single point failures. It's a waste of money to send up failure prone hardware. Amateurs building cubesats, probably not, but the big players aren't going to take that sort of risk.
- marsRoverDev 8y agoYou are right, they have redudnancy in all cases - but it isn't usually software written by multiple teams with different hardware.
- mi3law 8y agoThere was a discussion on HN a while back RE NASA's software safety guidlines. Here is the link to the discussion: https://news.ycombinator.com/item?id=12014271 https://news.ycombinator.com/item?id=12014271 The PDF linked to in the discussion is no longer there, but I found it on standards.nasa.gov here: https://standards.nasa.gov/standard/nasa/nasa-gb-871913 https://standards.nasa.gov/standard/nasa/nasa-gb-871913 There are also some interesting product management related guidelines from NASA, like this from 2014: https://snebulos.mit.edu/projects/reference/NASA-Generic/NPR-7150-002A.pdf https://snebulos.mit.edu/projects/reference/NASA-Generic/NPR...
- xtian 8y agoTalk on JPL's software for the Curiosity rover: https://www.usenix.org/conference/hotdep12/workshop-program/presentation/Holzmann https://www.usenix.org/conference/hotdep12/workshop-program/...
- firebacon 8y agoThe part that I find the most intriguing is "corrections can be made on the fly". I can see how you would ensure reliability through proper requirements specification, a good software development process, separate independent implementations and extensive verification. However, every time I read a popsci article about space flight software, they talk about this capability to push new code to the spacecraft while it is in flight. I'm really curious to learn what this looks like in practice (technical details). Do they really have the ability to do an "ad-hoc" upload and execution of arbitrary code on these systems? If so, how are the ad-hoc programs tested and verified?
- SomeHacker44 8y agoFrom previous articles, remote updates seem to be a core part of spacecraft software/operating systems. I even recall one situation where a spacecraft had a REPL built in that was used to fix a problem (slowly) remotely! They also have multiple levels of operation and watchdog functionality. I have no direct experience with that beyond following news about spacecraft.
- firebacon 8y agoRemote updates -- where you replace a full (sub)system -- are one thing, since you can always run the normal software validation procedure on the new version of the software. So an OTA update of a system (even in flight) does not sound like rocket science (yet)... But: Once you include a REPL or another mechanism to push and execute arbitrary code "ad-hoc", I wonder how that could possibly be tested an validated? Surely as soon as you add the ability to run arbitrary code, there is no way of testing for all possible states of the system as part of the validation process? In other words, how do you allow the user to push arbitrary code, but prevent them from putting the spacecraft into a condition from which it can not be recovered? The only way I could naively think of would be to only allow the user to push code to a completely isolated CPU that has a remote-reset functionality from the main/comms CPU. Still, the popsci articles I read made it sound like there might be more to it. It would be excellent to find some first-hand accounts/sources on how this looks like in reality.
- noselasd 8y agoThere are many good videos from the yearly Flight Software workshop http://flightsoftware.jhuapl.edu/ http://flightsoftware.jhuapl.edu/
- bitexploder 8y agoTheir cost per line of code is also, pardon me here, astronomical. That quality has a cost most shops cannot stomach.
- saalweachter 8y agoThere are a lot of really good links, but to be honest 99% of the secret to writing bulletproof code is “write the most simple, boringest program you can”. Which is not to say that what NASA and its contractors do isn’t cool or that they don’t spent ungodly amounts of time and money on testing and verification, but you also don’t load one line of code more than is absolutely necessary onto a machine that absolutely must work at all times. It’s an important lesson to learn and a good skill to exercise from time to time, but honestly it’s also something that doesn’t apply to most of our work as software engineers. For most software most people are willing to knock a couple of nines off the reliability of a piece of software in exchange for higher-quality output, lower costs, and more features. If my data analysis pipeline fails one time in ten because an edge case can use all the memory in the world or some unexpected malformed input crashes the thing but yields more useful output than if I kept it simple and hand-verified every possible input, well, that can be a fine trade off. If your machine learning model for when to retract the solar panel occasionally bricks and leaves the panel out to be destroyed, that’s less acceptable.
- reaperducer 8y agoyou also don’t load one line of code more than is absolutely necessary Coincidentally, I spent the weekend banging around with an old TRS-80 Model 100, and it's been very interesting to see what workarounds and compromises were made to conserve space. For example, the machine ships with no DOS at all, so if you're working with cassettes or modem only, you don't have that overhead. If you do add a floppy drive, when you first plug it in, you flip some DIP switches on the drive and it acts like an RS-232 modem, and you can download a BASIC program from the drive into the computer that, when run, generates a machine-language DOS program and loads it out of the way into high memory. I don't have one of those sewing machine drives, so I went with a third-party DOS, which weighs in at... wait for it... 747 BYTES.† An entire disk controller with command line interface in 2½ tweets. † http://bitchin100.com/wiki/index.php?title=TEENY.CO_MANUAL http://bitchin100.com/wiki/index.php?title=TEENY.CO_MANUAL
- corerius 8y agoalso, I would imagine that there would be a strong bias towards reuse...which leads you to long term standardization of not just language but also CPU architecture.
- GimbalLock 8y agoHello! FSW dev from NASA Langley here. We do try to do reuse as much as possible, but small satellites (CubeSats) are starting to change that. There are so many new pieces of hardware and so much experimentation going on to see what’s feasible in space. There are new RTOS frameworks being developed both by commercial and government (CFS, F-prime). If you’re interested in this in particular there is a conference called SmallSat which hosts the talks from previous years. https://smallsat.org https://smallsat.org
- GimbalLock 8y agoHello! I’m a FSW dev at NASA Langley. As others have said, the talks from the FSW workshop are a great start. If you want to see a well-used framework, check out CFS (https://cfs.gsfc.nasa.gov https://cfs.gsfc.nasa.gov)
- throwawaymath 8y agoOff topic, but I've always been interested by the way that government agencies almost exclusively choose acronyms for their software. Meanwhile private companies (especially in the last decade or two) almost always choose unrelated, single words. It initially seems kind of ridiculous to me that everything has an acronym, but I suppose it's no more ridiculous than choosing a name that sounds like a Pokemon. Maybe less so. In any case, thanks for sharing that.
- btrettel 8y agohttps://ntrs.nasa.gov/search.jsp?R=20100024508 https://ntrs.nasa.gov/search.jsp?R=20100024508 > The development and verification of the Charring Ablating Thermal Protection Implicit System Solver (CATPISS) is presented. [...] Not sure industry would try this one either, though it is very memorable.