10 ms·
Show me a C compiler that miscompiles the following code and I'll concede the point: uint32_t add_1(uint32_t a) { a += 1; return a; }
by jmillikin 2y ago
Show me a C compiler that miscompiles the following code and I'll concede the point:
uint32_t add_1(uint32_t a) {
a += 1;
return a;
}
- eru 2y agoI know that for 'int a' the statement 'a += 1' can give rather surprising results. And you made a universal statement that 'a += 1' can be trusted. Not just that it can sometimes be trusted. In C++ the code you gave above can also be trusted as far as I can tell. At least as much as the C version.
- jmillikin 2y agoI'll expand my point to be clearer. In C there is no operator overloading, so an expression like `a += 1` is easy to understand as incrementing a numeric value by 1, where that value's type is one of a small set of built-in types. You'd need to look further up in the function (and maybe chase down some typedefs) to see what that type is, but the set of possible types generally boils down to "signed int, unsigned int, float, pointer". Each of those types has well-defined rules for what `+= 1` means. That means if you see `int a = some_fn(); assert(a < 100); a += 1` in the C code, you can expect something like `ADD EAX,1` somewhere in the compiler output for that function. Or going the other direction, when you're in a GDB prompt and you disassemble the current EIP and you see `ADD EAX,1` then you can pretty much just look at the C code and figure out where you are. --- Neither of those is true in C++. The combination of completely ad-hoc operator overloading, function overloading, and implicit type conversion via constructors means that it can be really difficult to map between the original source and the machine code. You'll have a core dump where EIP is somewhere in the middle of a function like this: std::string some_fn() { some_ns::unsigned<int> a = 1; helper_fn(a, "hello"); a += 1; return true; } and the disassembly is just dozens of function calls for no reason you can discern, and you're staring at the return type of `std::string` and the returned value of `true`, and in that moment you'll long for the happy days when undefined behavior on signed integer overflow was the worst you had to worry about.
- eru 2y agoI heartily agree that C++ is a lot more annoying here than C, yes. I'm just saying that C is already plenty annoying enough by itself, thanks eg to undefined behaviour. > That means if you see `int a = some_fn(); assert(a < 100); a += 1` in the C code, you can expect something like `ADD EAX,1` somewhere in the compiler output for that function. Or going the other direction, when you're in a GDB prompt and you disassemble the current EIP and you see `ADD EAX,1` then you can pretty much just look at the C code and figure out where you are. No, there's no guarantee of that. C compilers are allowed to do all kinds of interesting things. However you are often right enough in practice, especially if you run with -O0, ie turn off the optimiser. See eg https://godbolt.org/z/YY69Ezxnv https://godbolt.org/z/YY69Ezxnv and tell me where the ADD instruction shows up in the compiler output.
- MaulingMonkey 2y agoerror: could not convert 'true' from 'bool' to 'std::string' {aka 'std::__cxx11::basic_string<char>'} I don't think anyone's claiming C nor C++'s dumpster fires have signed integer overflow at the top of the pile of problems, but when the optimizer starts deleting security or bounds checks and other fine things - because of signed integer overflow, or one of the million other causes of undefined behavior - I will pray for something as straightforward as a core dump, no matter where EIP has gone. Signed integer overflow UB is the kind of UB that has a nasty habit of causing subtle heisenbugfuckery when triggered. The kind you might, hopefully, make shallow with ubsan and good test suite coverage. In other words, the kind you won't make shallow.
- jmillikin 2y agoFor context, I did not pick that type signature at random. It was in actual code that was shipping to customers. If I remember correctly there was some sort of bool -> int -> char -> std::string path via `operator()` conversions and constructors that allowed it to compile, though I can't remember what the value was (probably "\x01"). --- My experience with the C/C++ optimizer is that it's fairly timid, and only misbehaves when the input code is really bad. Pretty much all of the (many, many) bugs I've encountered and/or written in C would have also existed if I'd written directly in assembly. I know there are libraries out there with build instructions like "compile with -O0 or the results will be wrong", but aside from the Linux kernel I've never encountered developers who put the blame on the compiler.
- TinkersW 2y agoa+=1 will not produce any surprising results, signed integer overflow is well defined on all platforms that matter. And we all know about the looping behavior, it isn't surprising. The only surprising part would be if the compiler decides to use inc vs add, not that it really matters to the result.
- eru 2y ago> a+=1 will not produce any surprising results, signed integer overflow is well defined on all platforms that matter. I'm not sure what you are talking about? There's a difference between how your processor behaves when given some specific instructions, and what shenanigans your C compiler gets up to. See eg https://godbolt.org/z/YY69Ezxnv https://godbolt.org/z/YY69Ezxnv and tell me where the ADD instruction shows up in the compiler output. Feel free to pick a different compiler target than Risc-V.
- jmillikin 2y agoI don't think "dead-code elimination removes dead code" adds much to the discussion. If you change the code so that the value of `a` is used, then the output is as expected: https://godbolt.org/z/78eYx37WG https://godbolt.org/z/78eYx37WG
- xmcqdpt2 2y agoThe parent example can be made clearer like this: https://godbolt.org/z/MKWbz9W16 https://godbolt.org/z/MKWbz9W16 Dead code elimination only works here because integer overflow is UB.
- jmillikin 2y agoTake a closer look at 'eru's example and my follow-up. He wrote an example where the result of `a+1` isn't necessary, so the compiler doesn't emit an ADDI even though the literal text of the C source contains the substring "a += 1". Your version has the same issue: unsigned int square2(unsigned int num) { unsigned int a = num; a += 1; if (num < a) return num * num; return num; } The return value doesn't depend on `a+1`, so the compiler can optimize it to just a comparison. If you change it to this: unsigned int square2(unsigned int num) { unsigned int a = num; a += 1; if (num < a) return num * a; return num; } then the result of `a+1` is required to compute the result in the first branch, and therefore the ADDI instruction is emitted. The (implied) disagreement is whether a language can be considered to be "portable assembly" if its compiler elides unnecessary operations from the output. I think that sort of optimization is allowed, but 'eru (presumably) thinks that it's diverging too far from the C source code.
- tux3 2y agoYou show one example where C doesn't have problems, but that's a much weaker claim than it sounds. "Here's one situation where this here gun won't blow your foot off!" For what it's worth, C++ also passes your test here. You picked an example so simple that it's not very interesting.
- eru 2y agoActually even here, C has some problems (and C++), too: I don't think the standard says much about how to handle stack overflows?
- jmillikin 2y ago'eru implied `a += 1` has undefined behavior; I provided a trivial counter-example. If you'd like longer examples of C code that performs unsigned integer addition then the internet has many on offer. I'm not claiming that C (or C++) is without problems. I wrote code in them for ~20 years and that was more than enough; there's a reason I use Rust for all my new low-level projects. In this case, writing C without undefined behavior requires lots of third-party static analysis tooling that is unnecessary for Rust (due to being built in to the compiler). But if you're going to be writing C as "portable assembly", then the competition isn't Rust (or Zig, or Fortran), it's actual assembly. And it's silly to object to C having undefined behavior for signed integer addition, when the alternative is to write your VM loop (or whatever) five or six times in platform-specific assembly.
- eru 2y agoForth might be a better competition for 'portable assembly', though.
- eru 2y agoYes, 'a += 1' can have undefined behaviour in C when you use signed integers. (And perhaps also with floats? I don't remember.) Your original comment didn't specify that you want to talk about unsigned integers only.
- WJW 2y agoC might be low level from the perspective of other systems languages, but that is like calling Apollo 11 simple from the perspective of modern spacecraft. C as written is not all that close to what actually gets executed. For a small example, there are many compilers who would absolutely skip incrementing 'a' in the following code: uint32_t add_and_subtract_1(uint32_t a) { a += 1; a -= 1; return a; } Even though that code contains `a += 1;` clear as day, the chances of any incrementing being done are quite small IMO. It gets even worse in bigger functions where out-of-order execution starts being a thing.
- eru 2y ago> It gets even worse in bigger functions where out-of-order execution starts being a thing. In addition, add that your processor isn't actually executing x86 (nor ARM etc) instructions, but interprets/compiles them to something more fundamental. So there's an additional layer of out-of-order instructions and general shenanigans happening. Especially with branch prediction in the mix.
- johnisgood 2y agoWhy would you want it to increment 1 if we decrement 1 from the same variable? That would be a waste of cycles and a good compiler knows how to optimize it out, or what am I misunderstanding here? What do you expect "it" to do and what does it really do? See: https://news.ycombinator.com/item?id=43320495 https://news.ycombinator.com/item?id=43320495
- lifthrasiir 2y agoIt is unlikely as is, but it frequently arises from macro expansions and inlining.
- acdha 2y agoThat’s a contrived example but in a serious program there would often be code in between or some level of indirection (e.g. one of those values is a lookup, a macro express, or the result of another function). Nothing about that is cheating, it just says that even C programmers cannot expect to look at the compiled code and see a direct mapping from their source code. Your ability to reason about what’s actually executing requires you to internalize how the compiler works in addition to your understanding of the underlying hardware and your application.
- gpderetta 2y agoIf my misocompile, you mean that it fails the test that a "C expression `a += 1` can be trusted to increment a numeric value", then it is trivial: https://godbolt.org/z/G5dP9dM5q https://godbolt.org/z/G5dP9dM5q
- Retr0id 2y agoHere's another one, just for fun https://godbolt.org/z/TM1Ke4d5E https://godbolt.org/z/TM1Ke4d5E
- jmillikin 2y agoI was being somewhat terse. The (implied) claim is that the C standard has enough sources of undefined behavior that even a simple integer addition can't be relied upon to actually perform integer addition. But the sources of undefined behavior for integer addition in C are well-known and very clear, and any instruction set that isn't an insane science project is going to have an instruction to add integers. Thus my comment. Show me a C compiler that takes that code and miscompiles it. I don't care if it returns a constant, spits out an infinite loop, jumps to 0x0000, calls malloc, whatever. Show me a C compiler that takes those four lines of C code and emits something other than an integer addition instruction.
- Retr0id 2y agoWhy are you talking about miscompilation? While the LLVM regression in the featured article makes the code slower, it is not a miscompilation. It is "correct" according to the contract of the C language.
- IshKebab 2y agoThat will be inline by any C compiler and then pretty much anything can happen to the 'a += 1'.
- deleted 2y ago[deleted]