10 ms·
C and C++ prioritize performance over correctness (2023)
- meling 2y agoI love the last sentence: “…, if you set yourself the goal of crossing an 8-lane freeway blindfolded, it does make sense to focus on doing it as fast as you possibly can.”
- ptsneves 2y agoMost comments like this, fall into a form of scarecrow fallacy. They assume performance in C/C++ can only come at the cost of correctness and then go on to show examples of such failures to prove the point, and there are so many. On the other hand the event space also has sets of cases where you can be correct and get faster. There are also sets of events where the failing risk is an acceptable tradeoff. Even the joke can come up empty, if you want to cross the 8bit lane the fastest possible and dont mind failing/dying some times it might be worth it. Also all the over emphasis on security is starting to be a pet peeve of mine. It sounds like all software should be secure by default and that is also false. When I develop a private project or a project with a threat scenario that is irrelevant i dont want to pay the security setup price, but it seems nowadays security became a tax. Cases in point: I cannot move my hard disk from one computer to another because secure boot was enabled by default without jumping hoops. I cannot install self signed certificates for my localhost without jumping hoops. I cannot access many browser APIs from an HTTP endpoint even if that endpoint is localhost. In that case i cannot do anything about it, the browser just knows better for my safety. I cannot have a localhost server serving mixed content. I mean come on why should i care about CORS locally for some google font. I cannot use docker build kit with a private registry with HTTP but to use a self signed certificat I need to rebuild the intermediate container. I must be nagged to use the latest compatibility breaking version library version for my local picture server because of a new DoS vulnerability. [...] On and on, and being a hacker/tinkerer is a nightmare of proselytizing tools and communities. I am a build engineer at heart and even I sometimes just want to develop and hack, not create the next secure thing that does not even start up This is like being in my home and the contractor forcing me to use keys to open every door to the kitchen, bedroom or toilet. The threat model is just not applicable, let me be.
- mananaysiempre 2y agoThat is more glib than insightful, I think: the programming equivalent of “as fast as you can” in this metaphor would likely be measured in lines of code, not CPU-seconds.
- gblargg 2y agoThe infinite loops example doesn't make sense. If count and count2 are volatile, I don't see how the compiler could legally merge the loops. If they aren't volatile, it can merge the loops because the program can't tell the difference (it doesn't even have to update count or count2 in memory during the loops). Only code executing after the loops could even see the values in those variables.
- jurschreuder 2y agoAlready all fixed in C++. And I don't know why now everything has to be beginner friendly. Then just use a high level language. C++ is never advertised as a high level language it's a mid-level language. Still with that even C++ has never shut down my computer or bricked my computer even once. All these young people are just too spoiled.
- jurschreuder 2y agoAlso in my experience, but you have to take my word for it, C++ "feels" more mathematically correct than Python. For me as an experienced developer Python has more undefined behaviour because it executes slightly different every time. The delays are different, threads work inconsistently. It's a bit like the metaphor that our brain is so big with so many small rules that it feels like you have free will. That's how Python feels. Somewhere in the millions of lines it generates there are bugs. I can sense it. Not enough to make it crash but they are there.
- dataflow 2y agoIt's not really that they prioritize performance over correctness (your code becomes no more correct if out-of-bounds write was well-defined to reboot the machine...), it's that they give unnecessary latitude to UB instead of constraining the valid behaviors to the minimal set that are plausibly useful for maximizing performance. E.g. it is just complete insanity to allow signed integer overflow to format your drive. Simply reducing it to "produces an undefined result" would seem plenty adequate for performance.
- hn-acct 2y agoThe author points out near the bottom that “performance” was not one of the original justifications for its UB decisions, afatct. Your example is a slippery slope but I get your point. I agree that there needs to be a “reasonable UB” But I’ve moved on from c++.
- tialaramex 2y ago> I agree that there needs to be a “reasonable UB” What could you possibly want "reasonable UB" for? If what you want is actually just Implementation Defined that already exists and is fine, no need to invent some notion of "reasonable UB".
- Brian_K_White 2y agoThey let the programmer be the ultimate definer of correctness. They don't prioritize performance over correctness, they prioritize programmer control over compiler/runtime control.
- hn-acct 2y agoBut they don’t really when some compilers silently remove code without mentioning it.
- Calavar 2y agoThe compiler removes code under the assumption that your code doesn't have UB. If your code has UB, that's a bug. "When my code is buggy the compiler outputs a buggy executable, but it's buggy in a different way than I want" has always struck me as somewhat of an odd complaint. Of course it can be difficult to know when you've unintentionally hit UB, which leaves room for footguns. This is probably an unpopular opinion, but to me that's not an argument for rolling back UB-based optimizations; it's an argument for better diagnostics (are you *sure* you meant to do this), rigorous testing, and for eliminating some particularly tricky instances of UB in future revisions of the standard.
- Maxatar 2y ago>... has always struck me as somewhat of an odd complaint. On the contrary I'd argue that the idea that any arbitrary bug can have any arbitrary consequence whatsoever is odd. There's nothing odd about expecting that the extent of an operation is bounded in space and time, it's a position that has a great body of research backing it.
- adgjlsfhk1 2y agothe problem is that any ub that is too difficult for a compiler to turn into a compile time error is also too difficult for humans to reliably prevent.
- 2y ago
- ajross 2y agoThis headline is badly misunderstanding things. C/C++ date from an era where "correctness" in the sense the author means wasn't a feasible feature. There weren't enough cycles at build time to do all the checking we demand from modern environments (e.g. building medium-scale Rust apps on a Sparcstation would be literally *weeks* of build time). And more: the problem faced by the ANSI committee wasn't something where they were tempted to "cheat" by defining undefined behavior at all. It's that there was live C code in the world that did this stuff, for real and valid reasons. And they knew if they published a language that wasn't compatible no one would use it. But there were also variant platforms and toolchains that didn't do things the same way. So instead of trying to enumerate them all individually (which probably wasn't possible anyway), they identified the areas where they knew they could define firm semantics and allowed the stuff outside that boundary to be "undefined", so existing environments could continue to implement them compatibly. Is that a good idea for a new language? No. But ANSI wasn't writing a new language. They were adding features to the language in which Unix was already written.
- bgirard 2y agoDid anything prevent them from transitioning undefined behavior towards defined behavior over time? > It's that there was live C code in the world that did this stuff, for real and valid reasons. If you allow undefined behavior, then you can move towards a more strictly defined behavior without any forward compatibility risk without breaking all live C code. For instance in the `EraseAll` example you can define the behavior in a more useful way rather than saying 'anything at all is allowed'.
- bluGill 2y agoNo, and that has been happening over time. C++26 for example looked at uninitialized variables and defined them. The default is intentionally unreasonable for all cases where this would happen just forcing everyone to initialize (and also because the value is unreasonable makes it easy for runtime tools to detect the issue when the compiler cannot)
- VWWHFSfQ 2y ago
- on_the_train 2y agoHoney it's time for your daily anti C++ post
- grandempire 2y agoI can only reply to so many UB comments in a day…
- 01HNNWZ0MV43FF 2y agoEvery day until my employer gives the juniors a blessing to learn something better
- pcwalton 2y agoI was disappointed that Russ didn't mention the strongest argument for making arithmetic overflow UB. It's a subtle thing that has to do with sign extension and loops. The best explanation is given by ryg here [1]. As a summary: The most common way given in C textbooks to iterate over an array is "for (int i = 0; i < n; i++) { ... array[i] ... }". The problem comes from these three facts: (1) i is a signed integer; (2) i is 32-bit; (3) pointers nowadays are usually 64-bit. That means that a compiler that can't prove that the increment on "i" won't overflow (perhaps because "n" was passed in as a function parameter) has to do a sign extend on every loop iteration, which adds extra instructions in what could be a hot loop, especially since you can't fold a sign extending index into an addressing mode on x86. Since this pattern is so common, compiler developers are loath to change the semantics here--even a 0.1% fleet-wide slowdown has a cost to FAANG measured in the millions. Note that the problem goes away if you use pointer-width indices for arrays, which many other languages do. It also goes away if you use C++ iterators. Sadly, the C-like pattern persists. [1]: https://gist.github.com/rygorous/e0f055bfb74e3d5f0af20690759de5a7 https://gist.github.com/rygorous/e0f055bfb74e3d5f0af20690759...
- AlotOfReading 2y agoThere's half a dozen better ways that could have been addressed anytime in the past decade. Anything from making it implementation defined to unspecified behavior to just throwing a diagnostic warning or having a clang-tidy performance rule. I'm also incredibly suspicious of the idea that FAANG in particular won't accept minor compiler slowdowns for useful safety. Google and Apple for example have both talked publicly about how they're pushing bounds checking by default internally and you can see that in the Apple Buffer hardening RFC and the Abseil hardening modes.
- pcwalton 2y ago> Anything from making it implementation defined to unspecified behavior to just throwing a diagnostic warning or having a clang-tidy performance rule. To be clear, you're proposing putting a warning on "for (int i = 0; i < n; i++)"? The most common textbook way to write a loop in C? > I'm also incredibly suspicious of the idea that FAANG in particular won't accept minor compiler slowdowns for useful safety. I worked on compilers at FAANG for quite a while and know quite well how these teams justify their existence. Telling executives "we cost the company $1M a quarter, but good news, we made the semantics of the language easier for programming language nerds to understand" instead of "we saved the company $10M last quarter" is an excellent strategy for getting the team axed next time downsizing comes around.
- netbioserror 2y agoThere's a way I like to phrase this: In C and C++, it's easy to write incorrect code, and difficult to write correct code. In Rust, it's also difficult to write correct code, but near-impossible to write incorrect code. The new crop of languages that assert the inclusion of useful correctness-assuring features such as iterators, fat-pointer collections, and GC/RC (Go, D, Nim, Crystal, etc.) make incorrect code hard, but correct code easy. And with a minimal performance penalty! In the best-case scenarios (for example, Nim with its RC and no manual heap allocations, which is very easy to achieve since it defaults to hidden unique pointers), we're talking about only paying a 20% penalty for bounds-checking compared to raw C performance. For the ease of development, maintenance, and readability, that's easy to pay.
- spacechild1 2y ago> but near-impossible to write incorrect code. Rust makes it near-impossible to make typos in strings or errors in math formulas? That's amazing! So excited to try this out!
- grandempire 2y ago> but near-impossible to write incorrect code. Except most bugs are about unforeseen states (solved by limiting code paths and states) or a disconnect between the real world and the program. So it’s very possible to write incorrect code in rust…
- deleted 2y ago[deleted]
- netbioserror 2y agoTrue, but I think errors in real-world modeling logic are part of our primary problem domain, while managing memory and resources are a secondary domain that obfuscates the primary one. Tools such as exceptions and contract programming go a long way towards handling the issues we run into while modeling our domains.
- 2y ago
- gavinhoward 2y agoAs a pure C programmer [1], let me post my full agreement: https://gavinhoward.com/2023/08/the-scourge-of-00ub/ https://gavinhoward.com/2023/08/the-scourge-of-00ub/ . [1]: https://gavinhoward.com/2023/02/why-i-use-c-when-i-believe-in-memory-safety/ https://gavinhoward.com/2023/02/why-i-use-c-when-i-believe-i...
- muldvarp 2y agoTo quote your article: > The question is: should compiler authors be able to do whatever they want? I argue that they should not. My question is: I see so many C programmers bemoaning the fact that modern compilers exploit undefined behavior to the fullest extent. I almost never see those programmers actually writing a "reasonable"/"friendly"/"boring" C compiler. Why is no one willing to put their ~money~ time where their mouth is?
- bsder 2y ago> I almost never see those programmers actually writing a "reasonable"/"friendly"/"boring" C compiler. Why is no one willing to put their ~money~ time where their mouth is? Because it is not much harder to simply write a new language and you can discard all the baggage? Lots of verbiage gets spilled about undefined behavior, but things like the preprocessor and lack of "slices" are way bigger faults of C. Proebsting's Law posits that compiler optimizations double performance every 20 years. That means that you can implement the smallest handful of compiler optimizations in your new language and still be within a factor of 2 of the best compilers. And people are doing precisely that (see: Zig, Jai, Odin, etc.).
- WalterGillman 2y agoI'm willing to write a C compiler that detects all undefined behavior but instead of doing something sane like reporting it or disallowing it just adds the code to open a telnet shell with root privileges. Can't wait to see the benchmarks.
- 2y ago
- uecker 2y agoYou can implement C in completely different ways. For example, I like that signed overflow is UB because it is trivial to catch it, while unsigned wraparound - while defined - leads to extremely difficult to find bugs.
- dehrmann 2y agoSome version of ints doing bad things plagues lots of other languages. Java, Kotlin, C#, etc. silently overflow, and Javascript numbers can look and act like ints until they don't. Python is the notable exception.
- AlotOfReading 2y agoThere's 3 reasonable choices for what to do with unsigned overflow: wraparound, saturation, and trapping. Of those, I find wrapping behavior by far the most intuitive and useful. Saturation breaks the successor relation S(x) != x. Sometimes you want that, but it's extremely situational and rarely do you want saturation precisely at the type max. Saturation is better served by functions in C. Trapping is fine conceptually, but it means all your arithmetic operations can now error. That's a severe ergonomic issue, isn't particularly well defined for many systems, and introduces a bunch of thorny issues with optimizations. Again, better as functions in C. On the other hand, wrapping is the mathematical basis for CRCs, Error correcting codes, cryptography, bitwise math, and more. There's no wasted bits, it's the natural implementation in hardware, it's familiar behavior to students from a young age as "clock arithmetic", compilers can easily insert debug mode checks for it (the way rust does when you forget to use Wrapping<T>), etc. It's obviously not perfect either, as it has the same problem of all fixed size representations in diverging from infinite math people are actually trying to do, but I don't think the alternatives would be better.
- jcranmer 2y ago> There's 3 reasonable choices for what to do with unsigned overflow: wraparound, saturation, and trapping. There's a 4th reasonable choice: pretend it doesn't happen. Now, before you crucify me for daring to suggest that undefined behavior can be a good thing, let me explain: When you start working on a lot of peephole optimizations, you quickly come to the discovery that there are quite a few cases where two pieces of code are almost equivalent, except that they end up giving different answers if someone overflowed (or some other edge case you don't really care about). Rather interestingly, even if you put a lot of effort into a compiler to make it aggressively infer that code can't overflow, you still run into problems because those assumptions don't really compose well (e.g., knowing that (A + (B + C)) can't overflow doesn't mean that ((A + B) + C) can't overflow--imagine B = INT_MAX and C = INT_MIN to see why). And sure, individual peephole optimizations don't make much of a performance effect. But they can sometimes have want-of-a-nail side effects, where a failure because of inability to assume nonoverflow in one place causes another optimization to fail to kick in and the domino effect results in measurable slowdowns. In one admittedly extreme example, I've seen a single this-might-overflow result in a 10× slowdown, since it alone was responsible for the autoparallelization framework to fail to kick in. This is happened enough to me that there are times I just want to shake the computer and scream "I DON'T FUCKING CARE ABOUT EDGE CASES, JUST GIVE ME THE DAMN FASTEST CODE." The problem with undefined behavior isn't that it risks destroying your code if you hit it (that's a good thing!); the problem is that it too frequently comes without a way to opt-out of it. And there is room to argue if it should be opt-in or opt-out, but completely absent is a step too far for me. (Slight apologies for the rant, I'm currently in the middle of tracking down a performance hit caused by... inability to infer non-overflow of an operation.)
- indigoabstract 2y agoAfter perusing the article, I'm thinking that maybe Ferraris should be more like Volvos, because crashing at high speed can be dangerous. But if one doesn't find that exciting, at least they'd better blaze through the critical sections as fast as possible. And double check that O2 is enabled (/LTCG too if on Windows).
- nine_k 2y ago/* I can't help but remember a joke on the topic. One guy says: "I can operate on big numbers with insane speed!" The other says: "Really? Compute me 97408743 times 738423". The first guy, immediately: "987958749583333". The second guy takes out a calculator, checks the answer, and says: "But it's incorrect!". The first guy objects: "Despite that, it was very fast!" */
- agentultra 2y agoIf you don't write a specification then any program would suffice. We're at C23 now and I don't think that section has changed? Anyone know why the committee won't revisit it? Is it purely, "pragmatism," or dogma? (Are they even distinguishable in our circles...)
- dang 2y agoDiscussed at the time: C and C++ prioritize performance over correctness - https://news.ycombinator.com/item?id=37178009 https://news.ycombinator.com/item?id=37178009 - Aug 2023 (543 comments)