25 ms·
I totally agree with de-emphasizing the old "recoverable" vs. "unrecoverable" dichotomy (https://blog.burntsushi.net/unwrap/#what-about-recoverable-vs-unrecover
by dcsommer 4y ago
I totally agree with de-emphasizing the old "recoverable" vs. "unrecoverable" dichotomy (https://blog.burntsushi.net/unwrap/#what-about-recoverable-vs-unrecoverable-errors https://blog.burntsushi.net/unwrap/#what-about-recoverable-v...). Every time I've heard programmers (especially in the context of exceptions) try to define it, I've found it imprecise and open to debate.
When invariant violations or mistakes by programmers (aka bugs) are detected, the program should halt as it is in an inconsistent state and continuing could be very dangerous (think privacy/security/data corruption). Otherwise, don't halt (handle it or have the caller handle it).
- arcticbull 4y agoYep, there's not really any such thing (IME) as a 'recoverable' error - except with respect to I/O. There's either I/O errors - or there's logic errors. A failure with logic should nuke due to the app being in an inconsistent state; trust is lost. An I/O error should fail softly.
- ithkuil 4y agoA parser library can return an error when the input is wrong. It doesn't necessarily mean the app is in an inconsistent state when that happens; it all depends on what the application does. It follows that aborting is often a sensible decision in application code and rarely in library code.
- arcticbull 4y agoI would personally consider parsers as part of I/O, but point taken.
- fatherzine 4y agoPartially true. In practice people implement multiplexed servers, for many reasons, including performance / throughput. A logic failure should nuke _the offending request_, not the entire server with all the unrelated concurrent requests.
- AshamedCaptain 4y agoThis is an attitude that I often see -- library authors who believe they own the process. Aborting may be the only sensible thing to do in runtimes where you can end up corrupting the entire process memory or the like, making recovery "dubiously possible", but absolutely not for anything higher level, where recovery may be safe and possible.
- burntsushi 4y agoI gave examples covering this area in the post. How would you rewrite the code in this section[1], for example, to conform to your views? (Assume this code is in a library.) [1]: https://blog.burntsushi.net/unwrap/#what-about-when-invariants-cant-be-moved-to-compile-time https://blog.burntsushi.net/unwrap/#what-about-when-invarian...
- fatherzine 4y ago* Short term: add explicit runtime checks for every step that may panic. With the way jump prediction works in modern processors, the runtime cost may be smaller than one may naively assume it is. Some of us would take, say, a 10% perf. degradation instead of chasing prod panics in the middle of the night. * Long term: plug-in a richer type (logic) system, so one can safely prove the most costly runtime invariants at compile time.
- burntsushi 4y ago> add explicit runtime checks for every step that may panic And do... what when the check fails? Can you please write the code for it? Because I don't understand what the heck you mean here. If you don't know Rust, pseudo code is fine. > Long term: plug-in a richer type (logic) system, so one can safely prove the most costly runtime invariants at compile time. This is just copying what I already said in the blog post. Until someone can show me how to prove the correctness of arbitrary DFA construction and search from a user provided regular expression pattern and demonstrate its use in a practical programming language like Rust, I consider this a total non-answer, not a "long term" answer. Basically, your answer here confirms for me that your views on how code should be structured are incoherent.
- jcelerier 4y ago> There's either I/O errors - or there's logic errors. A failure with logic should nuke due to the app being in an inconsistent state; trust is lost. nah. in GUI apps for instance you want the failure in the logic of a sub-sub-function to just tell the error "wops" when the button that triggered the action was clicked, not nuke the app (unless you hate your users). e.g. imagine a 3D software which allows to do mesh operations - user clicks on the "Smooth the mesh" button somewhere. Programmer forgot to handle a division by zero in some degenerate case of the smoothing computation which ends up leading to an exception: a value becomes zero, someone used unsigned integers for n in an "n - 1" computation which ends up in a call to array_of_floats.resize(0xffffffffffffffff) (and a likely std::bad_alloc being thrown if you're in c++). The original mesh is unchanged as the operation waits until the computation is complete to replace the old mesh with the new. If you ever decide to crash in this situation I am sure you will have great reviews on 3D modeling software comparisons.
- arcticbull 4y agoI would actually want that to crash, yes. That would ensure it gets caught and resolved by developers during development. Or at least increase the odds thereof. Further, if it allocates a few gb, or if it's allocating a large amount of memory because the size parameter got smashed? If it crashes enough that you'd get terrible reviews it would definitely be caught during development. If it messes up your mesh silently then you'll definitely still get bad reviews. And actually having an std::bad_alloc thrown is even worse since you have no idea what state it left things after running some destructors when you weren't planning for it.
- jcelerier 4y ago> That would ensure it gets caught and resolved by developers during development i don't know in which reality you live but in mine there's not much relationship between the existence of crashes, and them being resolved during development > And actually having an std::bad_alloc thrown is even worse since you have no idea what state it left things after running some destructors when you weren't planning for it. i don't even know what to say. that's the whole point of destructors - you know that things will be unwound in the reverse order from which they were created on automatic storage. that's, like, why c++ exists
- hinkley 4y agoRecoverable vs unrecoverable comes down to requirements. Certain companies known for having software that 'just works' tend to have both very few unrecoverable errors and very conservative feature sets to help facilitate that short list. It is very clearly a choice, even if many people are deciding by default. By not tackling an issue, you've chosen to have that issue.
- dcsommer 4y agoI think we have compatible views. Each layer of the software must decide it's requirements and handle errors appropriately per requirements. You're right I didn't articulate when to handle an issue locally vs. pass it up. I think that's where requirements (and also explicit API guarantees) come into play. I do think that APIs that "overpromise" by not returning the errors they do not handle to the caller, and instead halt or throw an exception, do their users a disservice in the long-run. These just become undocumented cases that bite you later on. Better libraries have all these conditions baked into the API itself.
- tsimionescu 4y agoBut an exception is a part of the API, why are you putting it at the same level as halting?
- dcsommer 4y agoIn my experience, APIs that throw rarely define all the exceptions that can come from it, especially transitively. I see exceptions as a failed (because undocumented, but still important for correctness) attempt at compromising between halting and returning an error.
- hinkley 4y agoI wonder if there's a moral equivalent of borrow semantics where we more formally define error propagation.
- jcranmer 4y agoThe criteria I tend to prefer is "expected" versus "unexpected" errors. I/O errors, especially network errors, are things that are going to be expected under reasonable operation, therefore it make sense that code should handle them. Similarly, user input resulting in incorrectly formatted code should be reasonably expected and therefore handled. But the same kinds of failures might not be reasonably expected in other circumstances--I wouldn't expect that the internal configuration files of an application should occur in reasonable operation, and therefore it makes sense to panic if they're corrupted... even if the cause is an I/O operation on a local disk, or parsing some JSON or TOML or INI or whatnot file. One implication of this is that it needs to be easy for any error system to promote an "expected error" into an "unexpected error"--which is what unwrap/expect does. The recoverable/unrecoverable error suggests that there ought to be no reason to do this, but there is absolutely a reason to do so: what category an error falls into is ultimately decided by the context of the error, not the generation of the error itself.
- merb 4y agoNetwork errors might be retryable/routable differently, but often (especially when starting out) should probably returned to the User. I mean if s3 is down you can retry the call but often it is down then
- dymk 4y agoSort of. At my company, if we removed retries from our services, our reliability would drop precipitously. Something like 99.99% of retries succeed on the second try, if there's not a hard service outage. If there is a hard outage, well, not much to do about that.
- preseinger 4y agoOnly the root of the call stack, fn main, should be able to return anything to the user. Everything else should return errors to their callers through the normal return mechanisms. Anything else, anything that introduces the possibility of shadow control flow, makes it basically impossible to maintain a working mental model of nontrivial programs.
- 4y ago
- simion314 4y agoDon't exception also halt your program if you ignore them? Also if using a library I don't want a bug in it to bring my program down, then I am forced to use workarounds like create a child process to use the library, start the child process from the main process and check on it to see if it fails or succeeds, that would be bad for performance and ugly.
- alerighi 4y ago> When invariant violations or mistakes by programmers (aka bugs) are detected, the program should halt as it is in an inconsistent state and continuing could be very dangerous (think privacy/security/data corruption). Otherwise, don't halt (handle it or have the caller handle it). Well it's not always the case. There are situations in which if you detect errors you want the program to continue running, and have only that particular functionality to fail. I tend to write resilient code, since I work in embedded systems and what you never want is the system to crash. Halting a CPU on an invariant violation (i.e. and assert failing) is something useful for debugging (you trigger the debugger and you then analyze why it happened), but something you generally don't want in production. Bette to have a ton of checks more and in case of an invariant violation (that maybe is resulting from a programmer mistake, but there is always the possibility of hardware memory corruption errors) to return an error and handle it in some ways (for example restart the task that returned the error, trying to go back to the last working state).
- burntsushi 4y ago> There are situations in which if you detect errors you want the program to continue running, and have only that particular functionality to fail. Yes, like a web server. If a request handler fails by panicking, in a Rust program, you catch the panic, respond with a 500 error and log the panic somewhere. But you continue serving other requests. I talked about this in the blog post. The problem with your strategy is that it requires you to be aware of your own mistakes. That doesn't sound like a robust strategy, unless you're investing huge resources into sophisticated tooling and have drastically restricted the expressivity of your programming environment. That exists and is fine, and I even addressed that in the blog post too.
- marshray 4y agoGreat article BTW, loved it! Will certainly become a classic. The web server example scares me. Something happened during execution that the programmer didn't expect. There's a bug in the program. What if the panic is due to memory corruption (less likely in Rust) or internal data structure corruption? Without knowledge to the contrary, swallowing a panic and YOLO'ing execution is driving full speed down the road of very poorly defined behavior. If the programmer had sufficient knowledge to conclude it was safe, they could have just used a Result<> to report the error. So ... don't make panic part of your API, and don't recover from panics?
- stormbrew 4y agoTo me the real issue is this is an extremely forced binary and there's really at least three meaningful categories (especially in software with a UI of any sort): - unactionable invariant violation (poisoned mutex, hard memory errors): crash immediately, something that should ever happen happened and there's no way to either handle or present the error to the user in a meaningful way. - unactionable (at the call site) but normal errors (couldn't open a file, disconnected from the remote end of a connection, etc): these need to be propagated up to where they can be turned into actionable information for a user, ideally. This is rarely a thing the call site where it happened can usefully do. - immediately actionable and normal errors (user input didn't validate, file user wanted to open doesn't exist, connection failed but can be retried with a backoff, etc). These need to be handled at the call site or maybe one or two levels up. You need an exception-like mechanism (or at least a process for emulating one, a la go MRV or C errno) to handle the second case, you often want it for the third case, but it never really makes sense to use it for the first. That said, I think in non-test rust code you should use expect instead of unwrap, because sometimes invariants do trip and that little tiny extra bit of info can make a huge difference to resolving it.