4 ms·
Ad #1: And bring down the whole process? Seriously, no, this is about the worse advice ever. Unless you are writing a tiny binary that works all on its own and
by fnl 9y ago
Ad #1: And bring down the whole process? Seriously, no, this is about the worse advice ever. Unless you are writing a tiny binary that works all on its own and on a single job/item, this is about as close as you can get to a mortal sin in any medium- or large-sized program.
Figure out what the remaining viable state of the program is, report the error (yes, loudly!), and recover. This would be the far more correct advice for #1 (except for tiny binaries doing a single job only, as said).
- ensiferum 9y agoSo you're one of those people who try to weasel their way out and clutter program logic to try to deal with bugs instead of just fixing the bugs. Seriously, in some cases when you're going for example OOB, the only reasonable action is to dump core. Trust me. This will improve your program, not make it worse.
- fnl 9y agoWhilst your application is in shatters for minutes or hours, your customer hotline going crazy, and your CFO will hate you for having to return thousands of good $$$ to your company's clients for the outage, to make them happy again. Trust me, you don't want to bring down the whole house just to fix a pipe...
- ensiferum 9y agoTo me this is an orthogonal problem. Ultimately high availability (in this context) is about making your software correct. When your software is correct you won't have availability issues because of bugs. Meanwhile if your software does crash (or terminate in a controlled fashion) but you need high availability you'll need to restart it automatically.
- fnl 9y agoThere is no such thing as an "orthogonal problem" here, both things connect. That's why you (a) want to have your processes staying alive in about any software as long as possible (hence invalidating your argument for #1), and (b) as I clearly mentioned, you do need to report such unexpected errors to the developer of the program, to avoid that they get swept under the rug. But handling cases from #1 as you described is nonviable in (virtually) all cases (except as described: if the error prevents the binary from completing that single job it was run on). Just imagine any of the programs you are running on your machine right now would just crash because they are not resilient to an off-by-one fencing error or whatever other bug (and there will be issues, even in your three-liners). I bet you'd be pretty pissed off if your programs would crash on you every other day or so...
- ensiferum 9y agoYou have it so backwards... Whats the point of a program producing incorrect results? Whats the value of that? Let's have an example. Person& findPerson(key_t key) {...} Now this is a core function, the basic precondition of the function is that the key is valid and exists. The function is called with a key that doesn't meet the precondition. A programmer who wrote the client code has made a bug. How do you deal with this situation? Throw and exception? Change the signature and return an error code + person? Great, now you've turned a bug into an "error", masking the actual problem. Then what does the calling code do? How is it prepared to handle "bugs". What does the logic to deal with bugs look like? Then you write bugs in your code trying to deal with bugs? Alternatively (I know some big orgs do this..) they put in a log and then return some "default object" and pretend there was no problem. Now how's the calling codes computation going to go, it's given bogus data to deal with. Bottom line, I want my programs to be bug free and correct. I don't want to even try to clutter my code with trying to write logic to deal with bugs. What I want on every bug is clean termination, program state and stack trace. It makes it very easy to fix the bug thus actually improving my software quality. I'm happy to know that my software actually has very few bugs. I know that you might think that it's inconvenient for the user to have his software abort. But if the software produces corrupt results then whats the value of that? Even worse, the user can think that results are correct even though it's garbage. Seriously, going back to the example. iF the precondition is validated you abort. Your developers can easily fix the offending code, release a new version of the software and your system is just that much better. If you incorporate this methodology from the start of the project you can greatly improve your quality in general.
- fnl 9y agoJust to clarify: I also don't think you should bring down the whole house when no financial interests are associated with the process dying. But when money gets involved, people typically start to behave and act more professionally, so its probably a good example, even for "non-financial cases."