4 ms·
The main distinction I'd draw is whether you'd doing formal or informal reasoning. In the formal reasoning world, undefined behavior is mostly a good thing. On
by raphlinus 4y ago
The main distinction I'd draw is whether you'd doing formal or informal reasoning. In the formal reasoning world, undefined behavior is mostly a good thing. On one side of the contract, it's a set of proof obligations for the code, and not especially onerous in the grand scheme of things - correct programs won't do UB. On the other side, it's a clear statement of what the compiler is and is not allowed to optimize. The more UB, the more opportunities to optimize.
When you're doing informal reasoning, the calculus changes. There's all kinds of stuff that can go wrong that is not motivated by what the machine is actually doing. In fact, it's something of a nightmare. Doing a memcpy of a struct that has padding in it? What are the exact semantics of restrict? And threading. Benign data races used to be a thing, but in an undefined behavior world, it's game over. C makes things worse than they need to be with its wonky integer rules (left shift of a negative integer and wrapping multiplication of two unsigned shorts are both UB), but a lot of that is potentially fixable and those mistakes won't be repeated in new languages.
In the context of Rust, more undefined behavior makes sense, and Ralf's work takes us much closer to a solid spec. But when you're doing mostly informal reasoning, I can see why people are so emotionally against it, and decisions such as turning off strict aliasing might be justified.
- jcranmer 4y agoWhile I do sympathize with some of the user complaints with UB, and the issues with things like signed integer overflow and strict aliasing seem entirely gratuitous, I think most users complaining about UB fail to comprehend that the issue with UB is that it's often really hard to constrain just what can possibly go wrong--and that's even without compiler optimizations kicking into play. It should be pretty clear that memory unsafety produces all sorts of crazy havoc--a write to errant memory could overwrite stack return locations and then basically do whatever it wants given the power of ROP gadgets. At first glance, it looks like uninitialized memory is "merely" an issue of reading more or less random data, but there are cases (e.g., MADV_FREE) where it turns out that the value of uninitialized memory can change underneath you. Traps cause lots of program state to become rather indeterminate, simply because of what may or may not live in a register or in memory, but on some architectures (e.g., Alpha), code may keep running for a while after an instruction traps, to the point that you're no longer even in the same function. Sanely describing what happens in data races are beyond the ken of formal semanticists (see the still-unsettled discussions over the semantics of relaxed atomics); what hope do programmers have of reasoning about these memory semantics? It also doesn't help that the distinction between undefined, unspecified, and implementation-defined behavior is poorly grasped by a large segment of the community.
- petergeoghegan 4y ago> While I do sympathize with some of the user complaints with UB, and the issues with things like signed integer overflow and strict aliasing seem entirely gratuitous, I think most users complaining about UB fail to comprehend that the issue with UB is that it's often really hard to constrain just what can possibly go wrong--and that's even without compiler optimizations kicking into play. That's probably true, but compiler people do themselves no favors by pretending that these things come from some higher echelon, that they couldn't possibly presume to question. It just doesn't pass the smell test. The fact that -wfrapv and -Wno-strict-aliasing are not the defaults in GCC is a choice made by GCC. A bad choice, in my opinion. MSVC made different choices, and lots of people still use it, so there is an existence proof that you can just not do these things on a mainstream compiler. (In fact, MSVC doesn't even offer type-based aliasing as an option that can be enabled, last I checked.)
- tialaramex 4y agoHow much performance gets left on the table as you disable ever more optimisations though? The justification for monstrously unsafe languages like C was that they're faster. If after removing optimisations which are too tricky to write for they're no longer faster then the languages don't pay their way any more and there's no reason to use them. I was expecting it would be easy to find benchmarks trying the same C or C++ code with GCC, Clang and MSVC and giving performance numbers, but I didn't find that. Maybe it exists and I can be directed to it ?
- petergeoghegan 4y ago> The justification for monstrously unsafe languages like C was that they're faster. I don't think that that's true. I find the explanation given by "Some Were Meant for C" [1] far more plausible. But leaving that aside: what does that have to do with anything that I said? And might I be permitted to make a point about GCC that is wholly unrelated to Rust, without getting a generic lecture about memory safety? > I was expecting it would be easy to find benchmarks trying the same C or C++ code with GCC, Clang and MSVC and giving performance numbers, but I didn't find that. Maybe it exists and I can be directed to it ? I don't doubt that there are silly compiler microbenchmarks somewhere. And I know for sure that strict aliasing could in principle make a huge difference. For example, an autovectorization optimization could take place once the compiler had leeway to applying an assumption about two pointers not aliasing, but not otherwise. However, in practice it doesn't seem to make all that much difference for most kinds of C programs, for all kinds of reasons that are very difficult to pin down. The big exceptions generally involve numerical code, which is why Fortran has always tended to be faster than C for numerical applications. At least it definitely was for most of the history of both languages. (Yes, C was openly understood to be slower than Fortran in cases that were important for Fortran 40+ years ago. I refer you to [1] once more.) [1] https://www.cs.kent.ac.uk/people/staff/srk21//research/papers/kell17some-preprint.pdf https://www.cs.kent.ac.uk/people/staff/srk21//research/paper...
- AlotOfReading 4y agoI don't think that's a useful distinction here. For context, I often work with high reliability software, including formal methods. What's needed for actual programs is the ability to say one of two things: A) There is no UB in program X or B) UB in program X cannot lead to a violation of constraint Y The current situation in the C family, whether you're using formal methods or not, is that you cannot generally prove (A) and the time traveling, no-holds-barred results of UB in the spec means that (B) is impossible. While rust doesn't entirely solve this, the fact that those statements are true everywhere except unsafe means that the scope of code you have to manually review is limited to something smaller than "everything".
- tialaramex 4y ago> the fact that those statements are true everywhere except unsafe This is not only a technical feature of Rust's standard library, but perhaps more importantly a cultural feature of Rust's ecosystem. The compiler has no technical problem with your "safe" implementation of Index for your type actually just doing unsafe pointer dereferences internally and trusting users to always pick valid indices, just like C. It's a bad idea, but the compiler is not a cop. However Rust's culture says if you're providing unsafe stuff that must be marked unsafe so that other people don't cut themselves on the sharp edges of your code by mistake. You can imagine with a different culture, you'd end up with popular code that's labelled "safe" but has UB all over the place because the interfaces lie everywhere as an "optimisation" and the community just puts up with it.