9 ms·
Show HN: Using C++23 <stacktrace> to get proper crash logs in C++ programs
- catiopatio 3y agoReliable in-process crash reporting is exceptionally difficult. The code must be fully async-safe, which means you cannot use <stacktrace>. You also cannot acquire mutexes, use any of the standard allocators, etc etc etc.
- hoten 3y agoWhat's the benefit of in-process crash reporting compared to just using something like crashpad/breakpad? To the extent that in-process crash reporting is even possible... seems the most common class of crashes would be entirely unrecoverable.
- deleted 3y ago[deleted]
- rightbyte 3y agoI guess it is easier to print out some interesting process variables, compared to trying to save and then make sense of dumps etc.
- kevin_thibedeau 3y agoIt has value in embedded code where you can stop the world to handle or log fault conditions.
- aseipp 3y agoEase of integration, because having literally anything is typically better than nothing. Honestly Crashpad isn't fun to integrate unless you use a fork like backtrace's (which adds CMake support), which I think doesn't help. I don't know of any alternatives. A version of Crashpad or something like it with a single turnkey server for database dumps, a one-line "defaults are good enough" integration, would be a real great thing to see.
- hoten 3y agoI found Sentry's crash reporting (which uses crashpad) simple enough to configure into an existing CMake build within an afternoon. Building Sentry/crashpad from source in a few lines of CMake: https://github.com/ArmageddonGames/ZQuestClassic/commit/3471ec702c90aa9a759fd9c9a2683bd3c33d7cb0#diff-1e7de1ae2d059d21e1dd75d5812d5a34b0222cef273b7c3a2af62eb747f9d20aR559-R575 https://github.com/ArmageddonGames/ZQuestClassic/commit/3471... And a few lines in the main function: https://github.com/ArmageddonGames/ZQuestClassic/commit/3471ec702c90aa9a759fd9c9a2683bd3c33d7cb0 https://github.com/ArmageddonGames/ZQuestClassic/commit/3471...
- aseipp 3y agoNeat, I didn't know Sentry also had a good fork, I haven't tried it! But in contrast, here's an in-process fault library that I whipped up (from forking Phusion Passenger) about 10 years ago that I still reach for sometimes, which is surprisingly robust to most of the original complaints about async safety, but still not perfect: https://github.com/thoughtpolice/libfault https://github.com/thoughtpolice/libfault You add one C file and 6 lines of code in `main()`, and you can do this in pretty much any programming language with a tiny extra bit of glue. It takes 3 minutes to do this in any C/C++ codebase of mine. It is build system agnostic and works immediately, with zero outside deps. It's something, and that's better than nothing, in practice. So people reach for that. I reach for it. And not just because I wrote it. I want to be clear: Crashpad is 10000x better than mine in every way, except this one way. And I really wish it wasn't. To add onto this, I really don't like CMake for example, so this problem isn't just a "well I like my thing." I want something that will also work in my Java programs, or Rust programs, for instance! Sometimes they crash too. I don't need to add any dependencies except like 2 or 3 C function calls, which almost every langauge supports with a native FFI out of the box. The friction is extremely low. I'm reminded of something Yann Collet once said about the design of zstd, and getting people to adopt new compression technology. If you make a compressor and it's better than an alternative in one or more dimensions, but worse in another (size, decompressor speed), then friction is actually significantly increased by that one failure. But if you make it better in every dimension -- so it gives an equal ratio and compression and decompression are always better than alternatives -- the friction is eliminated and people will just reach for it. Even though you only did worse in one spot, people find ways to make it matter. It really makes people think twice. But if it's always better, in every way, then using and reaching for it is just instinctive -- it replaces the old thing entirely. So that's what I really wish we had here. I think that's what you would need to see a lot better crash handling and reporting become more widely used. There needs to be a version of Crashpad, or any robust out of process crash collector, that you can just drop into any language and any build system with a little C glue (or Rust! Sure! Whatever!) in 5 minutes and it should have a crash database server and crash handler process which should instantly work for most uses.
- zX41ZdbW 3y agoThe best way I've found is - patching LLVM's libunwind to make it fully async-signal safe, and sending the stack trace to another thread for symbolization. This is implemented in ClickHouse.
- einpoklum 3y ago1. Can you link to that? 2. Have these changes been offered as patch for libunwind or boost::stacktrace?
- zX41ZdbW 3y agohttps://github.com/ClickHouse/libunwind/ https://github.com/ClickHouse/libunwind/ There were multiple steps: 1. Avoid using malloc/free inside libunwind. 2. Avoid using FDECache that required a mutex. 3. Avoid using dl_iterate_phdr (a mutex inside libc). 4. Protection from dereferencing wrong pointers due to incorrect unwind tables. Most of the changes were integrated to libunwind, but not everything. Example: https://bugs.llvm.org/show_bug.cgi?id=48186 https://bugs.llvm.org/show_bug.cgi?id=48186
- einpoklum 3y agoWhat is async-unsafe in using `<stacktrace>`? As for not using standard allocators - not a problem, just have a fixed area set aside as a buffer for crash reporting. Yes, it might not fit an extremely long report, but it's not that much of an issue.
- catiopatio 3y agoIt’s not guaranteed to be async-safe. From the C++ proposal (P0881R7): > Note about signal safety: this proposal does not attempt to provide a signal-safe solution for capturing and decoding stacktraces. Such functionality currently is not implementable on some of the popular platforms. https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2020/p0881r7.html https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2020/p08... [edit] Replying here, because HN is doing its occasional obnoxious rate-limiting of replies: Signal-safe and async-safe are effectively the same thing, and “async-safe” absolutely isn’t the same thing as “thread-safe”. A code path that acquires a mutex can be thread-safe; that’s absolutely not async-safe. If boost implemented a fully async-safe stack unwinder, complete with DWARF expression support, Apple compact unwind encoding support, and all the other features required across platforms, then good for them — but that’s not what <stacktrace> is guaranteed to provide, and such a thing is still not sufficient to implement anything but the most barebones portion of a real crash reporter.
- einpoklum 3y ago1. signal-safe and async-safe/thread-safe is not quite the same thing, but fair enough. 2. The boost::stacktrace library (on which the standardization was mostly based IIANM) has a `safe_dump_to()` function for these cases. See here: https://github.com/boostorg/stacktrace/blob/develop/include/boost/stacktrace/safe_dump_to.hpp#L81 https://github.com/boostorg/stacktrace/blob/develop/include/...
- rightbyte 3y agoThis is my favorite macro: "#define WIN32_LEAN_AND_MEAN" Why is it that even though Github or what ever has cutsie unicorns (or whatever it is) as error messages it feels fake and contrived while this define just feels like some random dude at MS naming it before going off to write Solitaire?
- hoten 3y agoDon't forget `WIN32_EXTRA_LEAN`! Still no idea what that does/did. For those curious about WIN32_LEAN_AND_MEAN - it reduces compile time by not auto-including a number of windows headers: https://devblogs.microsoft.com/oldnewthing/20091130-00/?p=15863 https://devblogs.microsoft.com/oldnewthing/20091130-00/?p=15...
- dataflow 3y agoDo you mean VC_EXTRALEAN? Or is WIN32_EXTRA_LEAN also a thing?
- TeMPOraL 3y agoFirst time I hear of VC_EXTRALEAN, always used WIN32_EXTRA_LEAN.
- nikbackm 3y agoThe bad news is that WIN32_EXTRA_LEAN has no effect, you have to define VC_EXTRALEAN if you want to exclude extra stuff from the MFC and VC headers. The good news is also that it has no effect. :) You can verify this for yourself by grepping the Visual Studio headers for VC_EXTRALEAN and WIN32_EXTRA_LEAN.
- Tempest1981 3y agoBut it did in 2005? > VC_EXTRALEAN defines WIN32_LEAN_AND_MEAN and a number of NOservice definitions, such as NOCOMM and NOSOUND. > https://gamedev.net/forums/topic/367942-win32_lean_and_mean-vs-vc_extralean/3423682/ https://gamedev.net/forums/topic/367942-win32_lean_and_mean-...
- elsamuko 3y agoA question: Would it be possible to pass the stacktrace of the current thread to another, so that the stacktrace would be traceable across threadpools or worker threads?
- soulbadguy 3y agoI am not sure if i understand the question correctly. But once collected, stack traces are just regular object that can be passed around thread as other object. It's possible that some implementation have references to some stack addresses (like for example the address of a function parameter), in which case you would need to serialize the stack trace before storing them/ moving then another thread.
- mike_hock 3y ago> once collected, stack traces are just regular object that can be passed around thread as other object. > It's possible that some implementation have references to some stack addresses (like for example the address of a function parameter), in which case you would need to serialize the stack trace before storing them/ moving then another thread. So which of these two mutually exclusive options is it? As I understand it, that was the question.
- elsamuko 3y agoWhen I debug multithreaded programs, the stacktrace of a breakpoint usually ends somewhere in a worker thread. What I want is that the worker thread's stacktrace part is replaced by the one who put the work into it. Kinda like the program wasn't multithreaded at all.
- TylerGlaiel 3y agoif you actually wanted to you could probably wrap thread to pass the stacktrace of the spawning thread into the worker thread whenever you spawn a thread and then output that upon a crash as well. the library seems pretty simple and flexible.
- zX41ZdbW 3y agoThere are parts of C++ standard library that no one should ever use. The examples are: regex, iostreams, locale... My main concern - this can also become such a dead weight.
- the_svd_doctor 3y agoCan you expand on the problems with regex?
- mike_hock 3y agoThe implementations are bad and implementers are refusing to fix their own bad implementations so as to not break their own ABI, but that has nothing to do with C++ the standard.
- dkersten 3y agoFrom what I’ve heard, std::regex is notoriously inefficient (lots of memory allocations, no allocator support). But I’ve never used it myself.
- TylerGlaiel 3y agooh yeah C++ regex is stupidly inefficient, like "python is faster" inefficient. I tried to use it for text replacements and pretty much immediately abandoned it
- gpderetta 3y agoApparently typical implementations of std::regex are inefficient like "'popen("perl..")' is faster" is inefficient! I thing boost::regex is significantly faster although not particularly fast
- verall 3y agoWhat's wrong with std::regex? Seems to work fine for me. And iostreams - they're not great. Bad programming UI, poor performance, etc. Issues abound. But should you never use them? What do you use instead? *printf methods have lots of issues too. And so does depending on Boost::format. And so does writing your own Logger/wrapping code (which is what everyone does AFAICT). locale is bad though
- deleted 3y ago[deleted]
- evmar 3y agoThe code: //a decent amount of this was copied/modified from backward.cpp (https://github.com/bombela/backward-cpp) The license on the other side of that link: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
- londons_explore 3y agoThis is only an issue if the original author chooses to enforce. You don't know - this code may have been given special permission to omit the notice by the original author. It isn't up to random joe to find copyrights they think have been violated.
- Rexxar 3y agoBoth have "MIT License" and he explicitly acknowledge the source. He probably forgot that he had to add an additional line with "Copyright 2013 Google Inc." in his own license file.
- mlhpdx 3y agoI may be wrong, but as I recall it is good form to chain exception filter calls by making note of the return from `SetUnhandledExceptionFilter`. For example, if I want to use `<stacktrace>` and do copy-on-write using memory protection in the same program.
- r2vcap 3y agoYet another library feature that will be used by nobody... While having a standard library is good, system programming presents tricky aspects like async-safety that are not adequately addressed. Therefore, considering the challenges involved, I believe it would be better to utilize existing libraries like crashpad to handle such scenarios.
- gpderetta 3y agoI'm afraid you are right. Without async-signal safety, the library is not terribly useful.
- koyote 3y agoMaybe I am doing something wrong but the example code does not seem to work on VS 2022 (no output file after a div by zero or stack overflow crash).
- PaulDavisThe1st 3y agoIf you don't care about exotica like async or signal safety, and just need to see the callstack from arbitray points, this can do the job without C++23: https://github.com/Ardour/ardour/blob/master/libs/pbd/stacktrace.cc https://github.com/Ardour/ardour/blob/master/libs/pbd/stackt... (2 different implementations, one for POSIX-y systems with the execinfo.h header, and one for Windows) The demange() function is elsewhere.
- anarazel 3y agoIme the execinfo.h backtraces are unfortunately not that useful in practice, due to being unable to resolve symbol names of static functions. But FWIW, you can use it for async signals with the _fd variant.
- anarazel 3y agoThis does not even remotely look to be signal safe to me?
- simfoo 3y agoThe signal handler just saves the trace and then wakes up another thread which does the reporting. The signaled thread uses only <stacktrace> and mutex/cv, all signal-safe as far as I can tell.
- anarazel 3y agoThe crashing thread might hold a lock in the memory allocator - which could either self deadlock when saving the stack trace (which certainly seems to do memory allocation and thus isn't a-signal safe), or could deadlock with the crash reporting thread which definitely allocates memory all over. I also quite doubt that std::mutex, and even more so std::condition_variable are guaranteed to be signal safe.
- simfoo 3y agostack_trace allows you to specify a custom allocator, which would protect against a lock held in the allocator (never ran into this in the real world though). You're right about mutex & cv in the general case though
- anarazel 3y agoThat assumes the specified allocator actually controls all allocations - somewhat doubtful across all platforms as things like dl_iterate_phdr() IIRC allocate memory (and take locks). And there's a lot more to writing signal safe code than not calling malloc. Unless an interface documents to be signal safe you're take better of assuming it is not. FWIW I've run into malloc self dreadlocks due to rare signals plenty of time :(. In production workloads.
- gpderetta 3y ago> I also quite doubt that std::mutex, and even more so std::condition_variable are guaranteed to be signal safe. They are not. Use sem_t for signaling. edit: also even if they where, the signal handler might wait forever for a mutex owned by the blocked thread.