4 ms·
Sorry, I'm kind of confused here. Why is the approach I'm defending considered random and not based on data? Is the generation of random hypotheses in general
by ectoplasm 11y ago
Sorry, I'm kind of confused here. Why is the approach I'm defending considered random and not based on data? Is the generation of random hypotheses in general considered unscientific? What about fuzz testing or pharmaceutical R&D? What is the precise difference between hypotheses and assumptions in the context of the scientific method? What is the difference between the presence of a small portion of code being responsible and the code being more broadly responsible? Why the emphasis on presence?
- rjurney 11y agoBecause it seems you're making an assumption about the problem being reproducible and identifiable by running half the code. I've never found this to be the case. Ever. Its a strange idea. Where does this come from? This isn't a hypothesis based on data, it is an assumption. And a bizarre one.
- ectoplasm 11y agoThe following program segfaults: main() { a(); b(); } You don't have a debugger. You don't have the source for a or b. Strategy? Sure printf works, but so does //.
- rjurney 11y agoI get the feeling you don't actually program. Or your writing is disconnected from the programmer part of you.
- ectoplasm 11y agoThat's interesting, could you elaborate?
- actualprogram 11y agoI can. It's long and tedious, because a lot of it is very simple things you learn very early on writing and maintaining software. You seem to be arguing that your example demonstrates the fact that you can localize a bug in the code with binary search; specifically that the bug must exist in either a() or b(), and that running them independently is both possible and will determine the single location of the bug. This is not true. It is only true in the case that one assumes bugs must have single locations, and that those locations can be found with binary search over what amount to incomplete programs. In other words, you're assuming the prior, begging the question, etc. It's tempting to say "real code isn't like this contrived example" but in fact this contrived example is the best possible demonstration of why blind binary search is a poor strategy. Let us say the 'bug' exists in a(). You seem to be assuming this is the necessary consequence: main () { a(); b(); } -- segfaults main () { a(); /* b(); / } -- segfaults main () { / a(); / b(); } -- doesn't segfault But this isn't the only possibility. With the bug existing in a(), you could also see this result: main () { a(); b(); } -- segfaults main () { a(); / b(); / } -- segfaults main () { / a(); / b(); } -- segfaults You would expect this in situations where b() uses data structures created by a(). We may also see this result, still assuming the bug exists in a(): main () { a(); b(); } -- segfaults main () { a(); / b(); / } -- doesn't segfault main () { / a(); / b(); } -- segfaults If above we learned nothing, here we've actually got a falsehood - our binary search has localized the bug to b(), but it actually exists in a()! So, in practice, binary search fails to localize a bug in a(). All of these situations can be created by having a() write a global which is relied upon by b() - a() may write to a protected area, b() may have a default value, a() may write nonsense, or pass an integer where b() expects an address - none of this is particularly exotic, they're all the sort of things you get every day when debugging segfaults. We might now delve into a competing series of contrived examples of a() and b() and argue about their relative prevalence in the world (which none of us are capable of knowing), because if some case is particularly rare, it may make binary search very slightly better than flipping a coin in this case. Instead, I will point out that this is once again* assuming the prior, and that we have these simple (if shockingly annoying) facts: 1) side effects exist 2) in the presence of side-effects, binary search cannot predict the location of a bug. And it follows that in the sense that a "scientific" hypothesis has predictive power, then binary search over the codebase is basically the homeopathy of debugging.