26 ms·
I Do Not Know C: Short quiz on undefined behavior (2015)
- eon1 10y agoC also: https://news.ycombinator.com/item?id=12902304 https://news.ycombinator.com/item?id=12902304
- userbinator 10y agoIMHO the problem is with compilers (and their developers) who think UB really means they can do anything, when what programmers usually expect is, and the standard even notes for one of the possible interpretations of UB, "behaving during translation or program execution in a documented manner characteristic of the environment". Related reading: http://blog.metaobject.com/2014/04/cc-osmartass.html http://blog.metaobject.com/2014/04/cc-osmartass.html http://blog.regehr.org/archives/1180 http://blog.regehr.org/archives/1180 and https://news.ycombinator.com/item?id=8233484 https://news.ycombinator.com/item?id=8233484
- monocasa 10y agoThere's no compiler writers throwing out if(undefined_behavior) { ruin_developers_day(); } It tends to be the effects of valid by the spec optimizations making assumptions that would only not be true during undefined behavior.
- nsajko 10y agoWhy restrict yourself to one compiler if you can write portable code? Clang and gcc provide flags that enable nonstandard behavior, and you can use static and dynamic (asan, ubsan) tools to detect errors in your code, it does not have to be hard to write correct code.
- jcranmer 10y agoStrict aliasing and ODR violations are extremely difficult to detect; these are the poster children for "undefined behavior that's hard to avoid and could seriously ruin your day if the compiler gets wind of it." There does appear to finally be a strict aliasing checker, but I have no experience with it.
- to3m 10y agoIn the main, people seem to be unfamiliar with what lies underneath C, so they never seem to really get this idea that you might be able to (or want to) expect any behaviour other than that imposed by its own definition.
- TillE 10y agoRight. Except for a few optimizer edge cases, you generally know what "undefined behavior" is going to spit out on a particular machine. Signed integer overflow, for example, almost always happens exactly the way you'd expect.
- maxlybbert 10y agoPeople have been a little sloppy with the terms, but there's a difference between implementation defined behavior and undefined behavior. Generally, the committee allows undefined behavior when it doesn't believe a compiler can detect a bug cheaply. Of course, many programmers complain about how the committee defines "cheaply." Trying to access an invalid array index is undefined because the way to prevent that kind of bug would be to add range checking to every array access. So, each extra check isn't expensive, but the committee decided that requiring a check on every array access would be too expensive overall. The same applies to automatically detecting NULL pointers. And the fact that the standard doesn't require a lot -- a C program might not have an operating system underneath it, or might be compiled for a CPU that doesn't offer memory protection -- means that the committee's idea of "expensive" isn't necessarily based on whatever platforms you're familiar with. But it is certainly true that a compiler can add the checks, or can declare that it will generate code that acts reliably even though the standard doesn't require it. And it's even true that compilers often have command line switches specifically for that purpose. But in general I believe those switches make things worse: your program isn't actually portable to other compilers, and when somebody tries to run your code through a different compiler, there's a very good chance they won't get any warnings that the binary won't act as expected.
- sjolsen 10y ago>the problem is with compilers (and their developers) who think UB really means they can do anything But that's exactly what undefined behavior means. The actual problem is that programmers are surprised-- that is, programmers' expectations are not aligned with the actual behavior of the system. More precisely, the misalignment is not between the actual behavior and the specified behavior (any actual behavior is valid when the specified behavior is undefined, by definition), but between the specified behavior and the programmers' expectations. In other words, the compiler is not at fault for doing surprising things in cases where the behavior is undefined; that's the entire point of undefined behavior. It's the language that's at fault for specifying the behavior as undefined. In other other words, if programmers need to be able to rely on certain behaviors, then those behaviors should be part of the specification.
- E6300 10y agoOn the other hand, the expected and desirable behavior in one platform might be different from that in another platform. It's possible to overspecify and end up requiring extra code when performing ordinary arithmetic operations, or lock yourself out of useful optimizations.
- sjolsen 10y agoWhich is exactly the motivation behind implementation-defined behavior. There's a broad range of "how much detail do you put in the specification" between the extremes of "this is exactly how the program should behave" and "this program fragment is ill-formed, therefore we make no guarantees about the behavior of the overall program whatsoever."
- E6300 10y agoImplementation-defined behavior at best just tells you that the behavior is guaranteed to be deterministic (or not). You still cannot reason about the behavior of the program by just looking at the source. And I'm not sure if optimizations such as those that require weak aliasing would be possible if the behavior was simply implementation-defined.
- E6300 10y ago1. Unless C's variable definition rules are completely different from C++'s, int i; is a full definition, not a declaration. If both definitions appear at the same scope (e.g. global), this will cause either a compiler error or a linker error. A variable declaration would be extern int i;
- khedoros1 10y agoC's variable definition rules are different from C++'s. gcc happily compiles those two lines, g++ exits with the "redefinition" error.
- E6300 10y agoThat was unexpected.
- shabbyrobe 10y agoSaid every C programmer ever!
- caf 10y agoYes, in C a plain int i; at file scope is a tentative definition - if, by the end of the compilation unit, no definition has been seen, one of them will become a definition, otherwise it is just a declaration. On the other hand, this: int i = 0; is a definition, and you can't have two of those.
- sparky_ 10y agoI suppose this sort of ambiguity is what drives the passion of Rust and Go programmers.
- barsonme 10y agoSorta. I write mostly Go (some JS, PHP) and I got 6/10, forgetting mostly stupid stuff like passing (-INT_MIN, -1) to #12. But some of those are prevalent in Go. For example, 1.0 / 1e-309 is +Inf in Go, just as it is in C—it's IEEE 754 rules. int might not always be able to hold the size of an object in Go, just like C. In Go #6 wraps around and is an infinite loop, just like C. The questions that don't, in some way, translate to Go are #2, #7, #8, and #10. But, to your credit, I do like how Go has very limited UB (basically race conditions + some uses of the unsafe package) and works pretty much how you'd expect it to work.
- hermitdev 10y agoIt's worth noting that for example #12, the assert will only fire for debug builds (i.e. the macro NDEBUG is not defined). So, depending on how the source is compiled, it may be able to invoke the div function with b == 0.
- DSMan195276 10y agoI'll be honest, I didn't find any of these to be particularly surprising. If you've been using C and are familiar with strict-aliasing and common UB issues I wouldn't expect any of these questions to seriously trip you up. Number 2 is probably the one most people are unlikely to guess, but that example has also been beaten to death so much since it started happening that I think lots of people (Or at least, the people likely to read this) have already seen it before. I'd also add that there are ways to 'get around' some of these issues if necessary - for example, gcc has a flag for disabling strict-aliasing, and a flag for 2's complement signed-integer wrapping.
- mjevans 10y agoI don't think #2 has been fully beaten to death yet. Assuming a platform where you don't segfault (say that 'page 0' variables are valid) and thus runtime does proceed; I still can't think of any /valid/ reason to eliminate the if that follows (focus line 2 in the comments). Under what set of logic does being able to de-reference a pointer confer that it's value is not 0 (which is what the test equates to)? In my opinion that is an, often working but, incorrect optimization.
- E6300 10y ago> Under what set of logic does being able to de-reference a pointer confer that it's value is not 0 (which is what the test equates to)? Simple: undefined behavior makes all physically possible behaviors permissible. In reality though, such an elimination would only be correct if the compiler was able to prove that the function is ever called with NULL, and if the compiler is smart enough to do that, hopefully the compiler writers are not A-holes and will warn instead of playing silly-buggers.
- mjevans 10y agoYou can't though. It's always possible to re-link the objects.
- 10y ago
- Tharre 10y agoI don't think this Q&A format makes for a good case of not knowing C. I mean I got all answers right without thinking about them too much, but would I too if I had to review hundreds of lines of someone else's code? What about if I'm tired? It's easy to spot mistakes in isolated code pieces, especially if the question already tells you more or less what's wrong with it. But that doesn't mean you'll spot those mistakes in a real codebase (or even when you write such code yourself).
- moosingin3space 10y agoThis is further compounded by how difficult it is to build useful abstractions in C, meaning that much real-world C consists of common patterns, and reviewers focus on recognizing common patterns, which increases the chances that small things slip through code review. Agreed that these little examples aren't too difficult, especially if you have experience, but I certainly do not envy Linus Torvalds' job.
- aidanhs 10y agoMy 'favourite' bit of surprising (not undefined) behaviour I've seen recently in the C11 spec is around infinite loops, where void foo() { while (1) {} } will loop forever, but void foo(int i) { while (i) {} } is permitted to terminate...even if i is 1: > An iteration statement whose controlling expression is not a constant expression, that performs no input/output operations, does not access volatile objects, and performs no synchronization or atomic operations in its body, controlling expression, or (in the case of a for statement) its expression-3, may be assumed by the implementation to terminate To make things a bit worse, llvm can incorrectly both of the above terminate - https://bugs.llvm.org//show_bug.cgi?id=965 https://bugs.llvm.org//show_bug.cgi?id=965.
- deleted 10y ago[deleted]
- adamnemecek 10y agoWhat's the point of this?
- bcoates 10y agoIt allows the optimizer to assume away the halting problem; all nontrivial loops are obligated to halt.
- Gibbon1 10y agoI read this as, will confuse the snot out of hapless newbie programmers trying to learn C via stepping though their code in an IDE. While providing no practical benefit to programmers writing production code.
- blackflame7000 10y agoIt is necessary in order for the compiler to do transformation optimizations which do impact production code. Newbies shouldn't be writing production code without guidance anyways IMO.
- junk_disposal 10y agoHonestly, Optimizing compilers will kill C. It killed the one thing C was good at - simplicity (you know exactly what happens where, note I'm not saying speed, as C++ can be quite a bit faster than C). Now, due to language lawyering, you can't just know C and your CPU, you have to know your compiler (and every iteration of it!). And if you slip somewhere, your security checks blow up (http://blog.regehr.org/archives/970 http://blog.regehr.org/archives/970 https://bugs.chromium.org/p/nativeclient/issues/detail?id=245 https://bugs.chromium.org/p/nativeclient/issues/detail?id=24...) .
- msbarnett 10y ago> Now, due to language lawyering, you can't just know C and your CPU, you have to know your compiler (and every iteration of it!). This mythical time never existed. You always had to know your compiler -- C simply isn't well specified enough that you can accurately predict the meaning of many constructs without reference to the implementation you're using. It used to, if anything, be much much worse, with different compilers on different platforms behaving drastically different.
- vyodaiken 10y agoThis is not really correct. The kinds of implementation dependencies usually encountered reflected processor architecture. The C standards committee and compiler community have created a situation in which different levels of "optimization" can change the logical behavior of the code! Truly a ridiculous state of affairs. The standards committee has some mysterious idea I suppose, but the compiler writers who want to do program transformation should work on mathematica or prolog, not C.
- prodigal_erik 10y agoCompiler writers have to use program transformation to do well on benchmarks. Developers who don't prioritize benchmarks probably don't use C, and if they do they really shouldn't, because sacrificing correctness for speed is the only thing C is good for these days.
- deleted 10y ago[deleted]
- rdc12 10y agoIsn't this line from #3, undefined behavior not mentioned in the article (sequence point violation) zp++ = xp + *yp;
- msbarnett 10y agoThat's not a sequence point violation. The C standard makes it clear that zp gets xp + *yp prior to the increment. Quoting 6.5.2.4 > The result of the postfix ++ operator is the value of the operand. After the result is obtained, the value of the operand is incremented. (That is, the value 1 of the appropriate type is added to it.) See the discussions of additive operators and compound assignment for information on constraints, types, and conversions and the effects of operations on pointers. The side effect of updating the stored value of the operand shall occur between the previous and the next sequence point. The last sentence is key.
- brianmurphy 10y agoAs a former C programmer, you know not to fool around at the max bounds of a type. That avoids all of the integer overflow/underflow conditions. When in doubt, you just throw a long or unsigned on there for insurance. :)
- kvakkefly 10y agoAnyone who enjoys this will also enjoy http://cppquiz.org http://cppquiz.org
- deleted 10y ago[deleted]
- AndyKelley 10y agoI made this post as a response. Disclaimer: yet another programming language trying to dethrone C. People seem to be less enthusiastic about the subject these days. http://andrewkelley.me/post/zig-already-more-knowable-than-c.html http://andrewkelley.me/post/zig-already-more-knowable-than-c...
- Hydraulix989 10y agoI feel bad because I'm smart enough to answer these questions correctly in a quiz format, but if I saw any of them in production code, I would not even think twice about it. (the quiz questions themselves lead you on, plus I read the MIT paper on undefined behavior that was posted on here back in 2013)
- federicoponzi 10y agoBefore: What? I know C. After 3 questions: Ok, I don't know C. Well played sir.
- nightcracker 10y agoI got every single one right. Does that mean I know C through and through? Perhaps. But all of these are the 'default' FAQ pitfalls of C, not the really tricky stuff.
- wmu 10y ago#4 is not really language issue, rather a floating point numbers feature.
- raarts 10y ago(2015)
- Kenji 10y agoI'm sorry, but the answer this website gives to 1. is wrong. See for yourself: int i; int i = 10; int main(int argc, char* argv[]){ return 0; } Try to compile it. It doesn't work (gcc.exe (GCC) 5.3.0), the error is: a.cc:2:5: error: redefinition of 'int i' int i = 10; ^ a.cc:1:5: note: 'int i' previously declared here int i; ^ Either I misunderstood the author and this example, or I do know C.
- mauricioc 10y agoJudging by the .cc extension, you are compiling this with a C++ compiler. Quoting from Annex C (which documents the incompatibilities between C++ and ISO C) of the C++ standard: Change: C++ does not have “tentative definitions” as in C E.g., at file scope, int i; int i; is valid in C, invalid in C++. This makes it impossible to define mutually referential file-local static objects, if initializers are restricted to the syntactic forms of C. For example, struct X { int i; struct X *next; }; static struct X a; static struct X b = { 0, &a }; static struct X a = { 1, &b }; Rationale: This avoids having different initialization rules for fundamental types and user-defined types. Effect on original feature: Deletion of semantically well-defined feature. Difficulty of converting: Semantic transformation. Rationale: In C++, the initializer for one of a set of mutually-referential file-local static objects must invoke a function call to achieve the initialization. How widely used: Seldom.
- Kenji 10y agofacepalm of course, even if I use gcc, if I compile a.cc it switches to the c++ compiler. Thanks.