5 ms·
Good to see this discussed - debuggability is not talked about enough but, done right, it could be a superpower. Setting the build for an old x64 machine (http
by mark_undoio 2y ago
Good to see this discussed - debuggability is not talked about enough but, done right, it could be a superpower.
Setting the build for an old x64 machine (https://dhashe.com/how-to-build-highly-debuggable-c-binaries.html#set-the-build-architecture-to-base-x86_64 https://dhashe.com/how-to-build-highly-debuggable-c-binaries...) for reversible / time travel debuggers seems unnecessarily restrictive to me. I'd expect a modern time travel debug tool (e.g. either rr or Undo - disclaimer, which I work on) to cope fine with most modern instructions (I believe GDB's built-in record / replay debugging tends to be further behind the curve on new CPU instructions - but if you're doing anything at scale it's not the right choice anyhow).
Regarding compilation (https://dhashe.com/how-to-build-highly-debuggable-c-binaries.html#partition-your-tus-into-debuggable-and-fast https://dhashe.com/how-to-build-highly-debuggable-c-binaries...) - we generally advise customers to use -Og rather than -O0. As the article states, this will still optimise out some code but should be a good trade-off without being too slow. (NB. last I checked, clang currently uses -Og as an alias for -O1, so it may behave less satisfactorily than under GCC).
It's also not said enough but: you don't need a special debug build to be able to debug. It's less user-friendly to debug a fully-optimised release build but it's totally possible. You just need to retain the DWARF debug info (instead of throwing it away). This is really important to know if you're debugging on a customer system or analysing a bug that's only in release builds.
- saagarjha 2y agoHaving debugged a lot of optimized code I would strongly recommend against it unless you are in a context where performance of your build is paramount (games?) or you cannot reproduce the bug when compiled without optimizations. Compilers really do a terrible job at preserving useful debug info when you turn them on. It’s a massive pain to have everything be marked as “optimized out” and reassemble the things you want from other variables or by using a disassembler to manually track which register the value is hiding in.
- SleepyMyroslav 2y agoIt's not only games. Anything sizeable that needs to run to repro will crumble under 20-100x times slowdown. Multithreaded behaviors will be just different. All those wonderful templated abstractions do not come for free in -O0. Ranges are especially egregious example. Debug build is truly dead outside of unit testing (imho). Realistic scenario that gamedev uses: deoptimize translation units you are interested in finding or reproducing bugs.
- senkora 2y ago> Realistic scenario that gamedev uses: deoptimize translation units you are interested in finding or reproducing bugs. Yep, looks like that’s this bullet point: https://dhashe.com/category/blog.html#partition-your-tus-into-debuggable-and-fast https://dhashe.com/category/blog.html#partition-your-tus-int...
- flohofwoe 2y ago> 20-100x This would be really unusual though right? For reference, in plain C code I see a 2x slowdown, in Zig a 4x slowdown (which I still need to investigate why exactly that's the case), and in C++ (even with heavy stdlib usage and on MSVC) at most 10x - which is the absolute worst case I've seen yet. My C++ info is a bit outdated though, have things gotten much worse in "modern" C++? Or in other words: if you see a slowdown of 100x in debug mode, I would be really concerned about why the performance is so heavily dependent on the optimizer doing it's thing and would start investigating what's the reason for such a massive slowdown.
- Cu3PO42 2y agoAt work, I have a large C++ codebase in which I can observe slowdowns of one order of magnitude or more with -O0 compared to -O3. And for many bugs it still takes 10 minutes of runtime to hit them, so not using optimizations just isn't tenable. As far as I can tell, this is worse than for the average C++ codebase. I have some ideas for why the optimizer is able to achieve such large improvements with this particular code and some of that could surely be done with better source, but it's a huge code base and a refactoring on that scale just isn't going to happen.
- mark_undoio 2y ago> It’s a massive pain to have everything be marked as “optimized out” and reassemble the things you want from other variables or by using a disassembler to manually track which register the value is hiding in. If you've got a time traveling / reversible debugger than you can (sometimes) go back to a point where the value was being written / used, at which point it'll often reappear in scope and be accessible. I believe DWARF's built-in virtual machine should be able to recompute missing values in many cases but I don't think compilers are great at putting the relevant info in, even where it should be possible to compute the right value fairly easily.
- mark_undoio 2y agoThe other trick I've found really helpful for "optimized out" values is to find places where they cross boundaries that block optimizations (e.g. procedure calls to another translation unit, so long as you're not doing some kind of link-time optimisation). e.g. if the value you're interested in is being passed to / returned from a function then inspecting it around the call / return site should have the value available.
- o11c 2y agoTwo particular notes around this (exact commands assuming gdb, but other debuggers should have equivalents): Pass various arguments to `backtrace` rather than just relying the default. Chances are it will have some non-optimized-out variables, which you can use to figure out what's going on. Use `info registers` and see what looks like a pointer, then cast it to a type you suspect it is. Note that this can be done for any stack frame.
- account42 2y agoEven with LTO or inside translation units, compilers are (sadly) extremely conservative about changing function boundaries.
- mark_undoio 2y agoI've seen gcc do an enthusiastic job on `static` functions within a C compilation unit - specifically where they only have one call site and so can be fully inlined into the caller. In that case, the code did become pretty hard to debug due to the extensive inlining and reordering it had allowed. Unfortunate because the only reason such functions exist is to make the structure of the code more apparent! Maybe that's an exception (and / or maybe it's easier with C than for C++). Maybe the less is that it's still always worth trying function call boundaries, in case the compiler has been conservative!
- DyslexicAtheist 2y agofeels strange still to see complaints about debugging in production being inconvenient when we should have caught these issues in test/staging. secondly I think not having debug tools and debug data in production is a security feature.
- saagarjha 2y agoYou’re trolling, right?
- o11c 2y ago> you don't need a special debug build to be able to debug Note that this is highly dependent on choice of compiler. Clang is utter crap for debugging even at -O1, but I've encountered basically no trouble ever using GCC at -O2 (you do have to learn a little about how the binary changes but that's easy enough to pick up). I really would not recommend -O3; historically it introduced bugs and regardless it makes the build process much slower, and the performance gain is fairly negligible (I can't say how much it destroys debuggability due to lack of experience). I can't speak for MSVC personally but it's a bad sign that its culture strongly promotes separate debug builds. That said, sanitizers are a place where a special debug build does help. Valgrind can do many of the things that sanitizers can but is around 10× slower which is a real pain if you can't isolate what you're targeting, so recompiling for sanitizers is a good idea. (Other brief notes) I have never actually encountered a case where the lack of frame pointers actually caused problems. As far as I'm concerned, any tool that breaks without them is a broken tool. (Theoretically they can speed up large traceback contexts if you're doing extensive profiling; good API design probably helps for the sanitizers case here) Rather than assembly int3, Unix-portable `SIGTRAP` is very useful for breakpoints; debuggers handle it specially. You can ignore it for non-debugged runs but get breakpoints when you are debugging without changing the binary or options! Alternatively you could leave it unignored if you have tooling that dumps core or something nicely for you.
- omoikane 2y agoDebugging experience aside, I found that "-O3" is generally worth it if you also set "-march=native". For example, here are some run times for computing SHA256, you can see that there is slightly more to be gained going from -O2 to -O3 with -march=native: -O2: 10.22 -O3: 9.82 -O2 -march=native: 9.86 -O3 -march=native: 9.43 This is basically SHA256 over ~8GB of data, averaged over 5 runs. The numbers are rather crude here since I measured them just now, but I remember it was more significant when I first did it last month for https://news.ycombinator.com/item?id=40687942 https://news.ycombinator.com/item?id=40687942
- josephg 2y agoYeah -march=native is amazing. I use it when compiling & benchmarking rust code. But - to anyone reading this later - please don’t do this blindly. You probably never want to distribute binaries with this flag set. It enables all the features available on the host CPU. So your build will change depending on the physical cpu you have installed. If you have a modern amd cpu, it may enable avx512 extensions and make your binary unusable on many Intel CPUs.
- dhashe 2y ago(author) > I believe GDB's built-in record / replay debugging tends to be further behind Yep, I hit this issue on gdb’s builtin stuff. I added a footnote linking here and saying that this is probably unnecessary for rr and Undo. Thanks!