5 ms·
Who said overflow was wrong? The sort of code I work on has to continue running, no matter what. Continuing on is always better than crashing, at least in my co
by nynx 3y ago
Who said overflow was wrong? The sort of code I work on has to continue running, no matter what. Continuing on is always better than crashing, at least in my context.
It’s also basically impossible to establish invariants about a system for code that is run in signal handlers
- lxgr 3y agoAnd I’ll take a crash over a silent wraparound any day. Out of curiosity, what domain is that? I can’t think of many use case where silent wrapping wouldn’t likely cause problems down the road.
- nynx 3y agoI write software for spacecraft. I’d much rather take my chances that the overflowed value is meant for humans or for some tertiary system than to end the mission regardless.
- bvrmn 3y agoOh shi, not very reassuring knowledge spacecraft SW engineers allow signed overflow in their code.
- Am4TIfIsER0ppos 3y agoI would guess they don't use a language where signed overflow is undefined. I use C all day every day and that's gotta be one of the worst legacy artefacts of it. As someone else said "it isn't 1970" the system is twos complement.
- bvrmn 3y agoDefined silent signed overflow (wrapping) is worse then UB. For UB there is at least runtime diagnostic. Parent clearly ok with `-fwrapv`.
- Dylan16807 3y ago> For UB there is at least runtime diagnostic. I have no idea what you mean by this. A couple types have diagnostics, but so do overflow and wrapping. Most don't. In general, it's much easier to have runtime diagnostics for arithmetic than for UB. I guess if I squint and "silent" implies you're not allowed to install diagnostics, then it's worse? But that's so artificial of a comparison as to be silly.
- Dylan16807 3y agoEven if the system is definitely twos complement, it's bad for i+1 to sometimes be less than i.
- lxgr 3y agoUnmanned spacecraft, hopefully? "Taking chances" is not a thing that I'd like the designers of a system that controls the vessel I'm on (or that overflies my house) routinely do during development. Undefined behavior can transitively expand to the entire system very quickly. For any subsystem, the proper error handling for an overflow/arithmetic error that's not properly handled right in the place it occurs is to declare the subsystem inoperative, not to keep going. If that approach routinely takes out too many and/or critical subsystems, you likely have much larger problems than integer overflows.
- torstenvl 3y agoThere are almost no circumstances where a crash in production code is preferable. Can you imagine working on a complicated Excel document when you run into a 2038 bug and you lose all your work instead of just having one of your future dates calculated as being in the 1970s? Or flying a plane and the system has to reboot midair instead of saying you're at -32768 feet of altitude? Or there are now 2.15 billion virus definitions, so your entire security suite just stops running? Or your Tesla stops keeping you in your lane without warning because your odometer turned over? Most things are not Therac-25. Much more software is protecting us and our data than is exposing us to danger. Loss of data, loss of navigation, loss of safety systems or life support... simply unacceptable. Turn on -ftrapv when testing and especially when fuzzing, sure, but production code should almost always be built with -fwrapv.
- nynx 3y agoThis is exactly right.
- astrange 3y agoCrash-only software works for Erlang, and I think people expect telephony switches to be reliable. (Swift also crashes on overflows and runs your phone.) Most of your examples would be just as bad if they started having unbounded incorrect behavior, and an overflow you didn't know about could lead to that. So don't get the math wrong!
- cesarb 3y ago> (Swift also crashes on overflows and runs your phone.) AFAIK, over half of the mobile phones in the world, and probably nearly all of the fixed phones, do not run Swift, but instead some older language which does not have this "crash on overflow" behavior.
- lxgr 3y agoMany landline and mobile phone switches do run on Erlang. Not getting a connection in case of e.g. an integer overflow for "pick the next available trunk line" is preferable to kicking out an existing connection on line 0. Note that an exception doesn't need to literally mean "the entire system comes to a screeching halt": You'd usually limit exception handling to the current operation.
- quickthrower2 3y agoNeural nets?
- Dylan16807 3y ago> Who said overflow was wrong? The word "overflow". > It’s also basically impossible to establish invariants about a system for code that is run in signal handlers Signals are a pain but it doesn't specifically have to be a signal.
- nynx 3y agoOverflow is perfectly well defined in most languages. What could it be other than a signal?
- eru 3y agoAlas, it's not well defined for signed integers in C (and C++) which are still popular languages.
- Dylan16807 3y ago> What could it be other than a signal? A normal exception without stack gremlins trying to break everything.
- kelnos 3y agoOverflow is fine if you're aware of it and have code that either doesn't need to care about it, or can work around it. Consider protocols with wrapping sequence numbers. Pretty common in the wild. If I increment my 32-bit sequence counter and it goes from 2^32-1 back to 0, that's likely just expected behavior.
- eru 3y ago> Continuing on is always better than crashing, at least in my context. In that case, you can just detect a crash (eg with an external watchdog), and replace your process with one that runs fizzbuzz. Fizzbuzz is wrong, but so would be continuing-no-matter-what-happened in general.
- MathMonkeyMan 3y agoThere are the kinds of defects that people notice, and then there are the kinds of defects that people don't notice. Crashes are usually noticed, while nonsense calculations and undefined behavior are sometimes not noticed. I worked at a company that sold a product that some customers used to make decisions involving a lot of money. A program misbehaved, or gave the wrong answer, or returned from a function too soon, or something. It didn't crash. A customer lost a lot of money, and we were liable. It turned out that the bug was a symptom of undefined behavior that would have been caught by one of the flavors of assertions built into standard releases of the core C++ libraries. The team in charge of the app in question had long ago disabled those assertions because "they keep crashing the code." Infrastructure within the company then went through a small internal crisis about whether to force assertions down people's throats or to let it be wrong on prod. The compromise was to have the assertions expand to loud warnings on prod. Ops then had a tool that would list all of the "your code is provably broken" instances, and that went to a dashboard used the shame the relevant engineering managers. Not sure if it worked, but when you have a _lot_ of C++ code, a great deal of it is broken and "it's fine." Until it isn't.