4 ms·
No, those are things that _usually_ happen in most C implementations. But reading from uninitialized memory provokes "undefined behavior" according to the stand
by peff 12y ago
No, those are things that _usually_ happen in most C implementations. But reading from uninitialized memory provokes "undefined behavior" according to the standard. Which might mean returning an unspecified value, might mean whole chunks of code being skipped, or might mean demons flying out of your nose.
The first one seems like an obvious choice of behavior, and one that some C programs doubtless depend on. But when you throw advanced optimizations into the mix, compilers may do unexpected things (e.g., assuming that a certain thing can't happen because it's not allowed by the standard, and then manipulating the code under that assumption). If you read other articles from Regehr, he describes some examples.
- angersock 12y agoThe linked flamewar on the GCC issue tracker ( https://gcc.gnu.org/bugzilla/show_bug.cgi?id=30475#c36 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=30475#c36 ) is an excellent example of this: citing "undefined behavior", compiler writers will gleefully do all kinds of goofy shit and, if confronted by developers who actually program for a living, stick their fingers in their ears and go "nyah nyah nyah undefined behavior we can do whatever we want go learn c nyah nyah nyah". "Undefined behavior" is a really great trick for promoting pet optimization projects and stonewalling practical feature requests by use of language lawyering.
- DSMan195276 12y agoHow was it goofy? Personally what the compiler did made perfect sense to me. If you assume integers can't overflow, then 'b' must be larger then 'a'. Thus why would the compiler bother performing the statement 'c=(b>a)' when it's 'obvious' that it's just going to be 'c=1'? That said, looking at that page the guy who made this bug was being more then extremely annoying, the person responding to the bug was fairly civil all things considered. You're complaining about UD but it's a necessary evil in C. Integer overflow was perhaps a bad choice by the standards makers, but the fact still stands that even if GCC did the 'right' thing, there's no guarantee that clang or any other compiler will do the same thing. The code would still be broken, it just might be harder to figure that out. If you want integer overflow and wrapping then use the compiler flag for it and write non-standard code. IMO, the bigger problem is that people write their code, compile it with gcc, and then assume it's standards compliant because it 'works' with gcc.
- angersock 12y agoStack overflow question illustrating the problem: http://stackoverflow.com/questions/7682477/why-does-integer-overflow-on-x86-with-gcc-cause-an-infinite-loop http://stackoverflow.com/questions/7682477/why-does-integer-... "Principle of Least Surprise" is that an integer will wrap, because that's what the hardware does in pretty much all cases. Any clever optimizations or undefined behavior should be happening due to explicit flags--which is exactly what the guy in that bug report wanted. The thing is that, for like the last half-century, we've expected integers to overflow and wraparound..that's just how they work. Ignoring that kind of expectation is asinine.
- DSMan195276 12y agoThere is an explicit flag, he's compiling with '-O2'. As he noted, without -O2 the output is correct. gcc does exactly what you're saying it should do in this instance, so I don't see what you're unhappy about.
- angersock 12y agoConsulting the original bug report, the optimization is hardly clear in performance benefits. Note also that, again, optimization somehow breaking 50 years of numerical reasoning is probably not a good 'default' behavior (even in O2! especially without clear benchmarks proving its utility!).
- wahern 12y ago50 years? Two's complement didn't predominate until the late 1970s, early 1980s. Before that time ones' complement predominated. And there are plenty of processors today which only use sign-magnitude. In particular, floating point-only CPUs. Compilers must emulate two's complement for unsigned arithmetic, and so signed arithmetic is significantly faster. The C standard is what it is for good reason. It's not anachronistic. Rather, now there are a million little tyrants who can't be bothered to read and understand the fscking standard (despite it being effectively free, and despite it being 1/10th the size of the C++ standard) and who are are convinced that the C standard is _obviously_ wrong. Which isn't a comment on this friendly-C proposal. But the vast majority of people have no idea what the differences are between well-defined, undefined, implementation defined, or unspecified behavior, and why those distinctions exist.
- ben0x539 12y ago> citing "undefined behavior", compiler writers will gleefully do all kinds of goofy shit and, if confronted by developers who actually program for a living, Do you imagine that compilers are created by lawyers or something?
- to3m 12y agoBut undefined behaviour isn't disallowed by the standard, and these things aren't not allowed to happen! If they were, the standard wouldn't even bother to mention any of it, and certainly wouldn't bother to suggest that one option is for things to behave "during translation or program execution in some documented manner characteristic of the platform". (See the C11 draft standard, 3.4.3.2; wording is basically the same in C99 I think.) It seemed obvious to me from the moment I first heard about this stuff that undefined behaviour is there to avoid binding implementations' hands too tightly. It's a way of allowing as wide a range of implementations as practical to be standard-compliant, by not forcing the compiler to patch over every last difference between systems or provide missing functionality. But is it a way to let gcc do whatever it likes, having proven your program invalid on a technicality? Well... I'm less sure about that one. (Obligatory links: http://blog.metaobject.com/2014/04/cc-osmartass.html http://blog.metaobject.com/2014/04/cc-osmartass.html, http://robertoconcerto.blogspot.co.uk/2010/10/strict-aliasing.html http://robertoconcerto.blogspot.co.uk/2010/10/strict-aliasin...)
- comex 12y agoThe suggestions that performance improvements brought by compiler optimizations are meaningless bother me, though. First, because hardware isn't getting faster that quickly anymore: Moore's Law hasn't meant for a long time that CPUs actually double their per-thread performance every 18 months, so that "1/10th as effective" from your first link, which refers to compilers hypothetically doubling performance every 18 years, starts to get more and more attractive. Second, because while in C/C++ code the programmer can often avoid useless machine code, newer languages such as Rust and Swift tend to do more stuff implicitly (safety checks, reference counting) which could often be eliminated by a Sufficiently Smart Compiler - increasing the need for good optimizations. (I think that this somewhat mirrors C++'s early history compared to C, but I was too young then to have any personal experience.) Of course those languages also tend to have no undefined behavior, so it's a bit different... Third, because I don't think undefined behavior is as evil and hard to avoid as people think it is. I think there are a few weird points (you can cast pointers into malloced buffers to any type as long as you're consistent, but there's no way to do that for a static buffer without fully static layout), and it would be nice to have more control over things like aliasing - both to loosen rules and to tighten them (i.e. more flexible restrict-like functionality). But despite being a fun topic, it doesn't seem to come up that often in practice from what I've seen, so the performance gain is close to free. And when you're, say, fighting for 60fps in a CPU limited scenario, it's hard to turn down even a small free performance gain. Fourth, because as someone who reads assembly frequently, idiotic looking assembly bothers me aesthetically even if it often doesn't have much performance impact. I speak in particular of reloading struct fields over and over when the data is already in a register and any person looking at the code would know it would be illogical for it to change in memory since it was loaded - but the compiler isn't smart enough to prove it can't alias... Sure, in individual cases it's easy to cache it in a local variable to stop this from happening, but in the large it's hard to avoid. Strict aliasing improves the situation somewhat, which is one reason I like it.