11 ms·
Error Stacking in Rust
- dpc_01234 2y agoProbably also worth mentioning: https://crates.io/crates/error-stack https://crates.io/crates/error-stack
- samanthasu 2y agoA good error report is not only about how it gets constructed, but what is more important, to tell what human can understand from its cause and trace. In this example, we analyzed and showed how to design stacked errors and what should be considered in this process.
- gfreezy 2y agoasync fn handle_request(req: Request) -> Result<Output> { let msg = decode_msg(&req.msg).context(DecodeMessage)?; // propagate error with new stack and context verify_msg(&msg)?; // pass error to the caller directly process_msg(msg).await? // pass error to the caller directly } async fn decode_msg(msg: &RawMessage) -> Result<Message> { serde_json::from_slice(&msg).context(SerdeJson) // propagate error with new stack and context } how to capture the virtual stack when `verify_msg` returns an error? Do you have some lint to make sure every error is attached with a context?
- shepmaster 2y agoI don't think you need a lint. When you define the error type returned by `handle_request`, you decide how the error type returned by `handle_request` will be incorporated. If you've decided to implement `From` then you've decided you don't want/need to add context. Otherwise, the compiler will give you an error when you use `?`. The time I can think this won't work is when you are reusing error types across places. Recently, I've been experimenting with creating a lot of error types, so far as one unique error type per function. I haven't done this for long enough to have a real report, but I haven't hated it so far.
- gfreezy 2y ago[dead]
- namjh 2y ago> Consequently, this also means you cannot define two error variants from the same source type. Considering you are performing some I/O operations, you won't know whether an error is generated in the write path or the read path. This is also an important reason we don't use thiserror: the context is blurred in type. This is true only if you add #[from] attribute to a variant. Implementing std::convert::From is completely optional. Personally I don't prefer it too as it ambiguates the context. I only use it for "trivially" wrapped errors like eyre::Report.
- skavi 2y agoYup. I absolutely would throw `#[from]` on everything when I started using thiserror, but now only do so in incredibly obvious cases like enum CarWontMove { EngineTroubles(EngineTroubles), WheelsFellOff(WheelsFellOff), } Even then, there’s often some additional context you can affix at that higher level.
- shepmaster 2y agoSNAFU follows much the same idea: we have an attribute you can add [0] when you want to allow directly implementing `From`. Like thiserror, you can also mark an error as transparent [1] when even the error existing doesn't provide useful information. [0]: https://docs.rs/snafu/latest/snafu/derive.Snafu.html#disabling-the-context-selector https://docs.rs/snafu/latest/snafu/derive.Snafu.html#disabli... [1]: https://docs.rs/snafu/latest/snafu/derive.Snafu.html#delegating-to-the-underlying-error https://docs.rs/snafu/latest/snafu/derive.Snafu.html#delegat...
- dgfitz 2y ago[flagged]
- speed_spread 2y agoRust's syntax is largely irrelevant to its purpose. If you don't see the need for it, might as well learn something else.
- more-nitor 2y agoidk if its about a few syntax, then it's possible to make a temp proc-macro for those
- jtrueb 2y agoWhich statically typed language do you find most agreeable? Same question for any language.
- GrantMoyer 2y agoThis seems to be a fairly common sentiment. I consider Rust's syntax fairly consistent and elegant for a curly brace language, but evidently I have some blind spots. What quibbles do you have with Rust's syntax?
- akira2501 2y agoThe explosion of single character sigils and the taint of C++'s template syntax.
- shepmaster 2y agoHey all, I’m the author of SNAFU (mentioned in the article). I’m off to bed now, but I’d be happy to try and answer any questions people might have sometime tomorrow. I’m glad to see SNAFU was useful to others!
- zamalek 2y agoIts looks really neat! Two questions: * can it be used as a build dependency (i.e symbols from the snafu crate don't appear in the generated code). * I assume you have to use one of the macros (ensure! or location!) when constructing an error that contains a location?
- shepmaster 2y agoIt can't be used as a literal build dependency [0], no. However, the fact that your crates uses SNAFU should [1] be completely hidden from your users. From the outside, you just return a regular enum or struct as your error type. If you were to look at the symbols in the resulting binary, I would expect that you could see references to the trait method `snafu::ResultExt::context` (and similar functions across similar types) depending on how well the code was inlined. If you use other features like `snafu::Location` or `snafu::Report`, those would definitely show up. You don't have to use the macros, no. When you define your error type, you can mark a field as `#[snafu(implicit)]` [2]. When the error is generated, that field will be implicitly generated via a trait method. The two types this is available for are backtraces and locations, but you could create your own implementations such as grabbing the current timestamp or a HTTP request ID. [0]: https://doc.rust-lang.org/cargo/reference/specifying-dependencies.html#build-dependencies https://doc.rust-lang.org/cargo/reference/specifying-depende... [1]: There's one tiny leak I'm aware of, which is that your error type will implement the `snafu::ErrorCompat` trait, which is just a light polyfill for some features not present on the standard library's `Error` trait. It's a slow-burn goal to remove this at some point, likely when the error "provider API" stabilizes. [2]: https://docs.rs/snafu/latest/snafu/derive.Snafu.html#controlling-implicitly-generated-data https://docs.rs/snafu/latest/snafu/derive.Snafu.html#control...
- Sytten 2y agoWhat is really annoying with thiserror is the wizard refusal to give us an easy way to print the error chain. No I dont want to convert it to anyhow just to print the error...
- lumost 2y agoRust is full of these, I’ve found the community simply falls back on user error to understand rust when vexed by in my opinion basic software operations. As someone who works extensively in cpp/java/python. I want so much to love rust, but unfortunately I haven’t found it to be productive after 6+ side projects.
- nixpulvis 2y agoRust's community is slightly more fragmented than it should be. The community being built while the language was changing so dramatically (e.g. async) didn't help, but it also is part of what lead to Rust in the first place. But it's still somewhat young, lots of stuff is being built. So some of the lack of productivity probably just comes from not knowing the right stacks yet.
- lumost 2y agoIt’s young, but my experience has been that developer ergonomics is not a focus, to the extent that c++ has a much stronger devex story.
- enigma101 2y agosame here.
- zote 2y agoURL changed, I think: https://greptime.com/blogs/2024-05-07-rust-error-handling https://greptime.com/blogs/2024-05-07-rust-error-handling
- joshka 2y agoIt's technically feasible to add SpanTrace support to thiserror fairly easily (30 mins work - Issue: https://github.com/dtolnay/thiserror/issues/400 https://github.com/dtolnay/thiserror/issues/400, PR: https://github.com/dtolnay/thiserror/pull/401 https://github.com/dtolnay/thiserror/pull/401). This would solve part of the problem in a way that is meaningfully good for that side of the ecosystem. I suspect you could probably do something similar for Snafu
- shepmaster 2y agoWithout deeply looking into it, I'd expect that to integrate with SNAFU, you could basically write something like this: struct SpanTraceWrapper(tracing_error::SpanTrace); impl snafu::GenerateImplicitData for SpanTraceWrapper { fn generate() -> Self { Self(tracing_error::SpanTrace::capture()) } } And then you can use it as #[derive(Debug, Snafu)] struct SomeError { #[snafu(implicit)] span_trace: SpanTraceWrapper, } This will capture the `SpanTrace` whenever `SomeError` is constructed (e.g. `thing().context(SomeSnafu)` or `SomeSnafu.fail()`.
- joshka 2y agoNeat :)
- lilyball 2y ago> Then, to be able to translate the stack pointer we will need to include a large debuginfo in our binary. In GreptimeDB, this means increasing the binary size by >700MB (4x compared to 170MB without debuginfo). Surely that's comparing full debuginfo, right? Backtraces just need symbols, not full debuginfo, and there's no way the symbols are 4x the size of the binary.
- dwattttt 2y agoThere's also split-debuginfo, which allows emission of debug info into a separate file, rather than needing to distribute it in the binary. Then they could capture stack traces, and resolve the symbols later if necessary. That would also address their concern about how long it takes to capture a stack trace, because just gathering the addresses themselves is quick.
- exDM69 2y agoWhy does adding `backtrace` to thiserror/anyhow require adding debug symbols? You'll certainly need it if you want to have human readable source code locations, but doesn't it work with addresses only? Can't you split off the debug symbols and then use `addr2line` to resolve source code locations when you get error messages from end users running release builds?
- delusional 2y agoYour binary usually won't get loaded at the same address in memory. The addresses would be useless without the memory map. That's solvable though. The bigger problem is how you unwind the stack. the stack is not generally unwindable, unless you're the compiler. Debug symbols include information from the compiler about the stack sizes and shapes to help backtrace with unwinding the stack. It's quite possible to include such symbols in the final binary without adding debug symbols, a lot of compilers just don't have a specification for that.
- exDM69 2y ago> Your binary usually won't get loaded at the same address in memory. The addresses you typically see in a backtrace error message (with debug syms disabled) are relative to the sections in the binary file, the runtime address it was loaded at has already been taken into account and subtracted. At least that's how you typically see a backtrace address in a typical native app on Linux. > The bigger problem is how you unwind the stack. Rust can unwind the stack on panic when built without debug symbols.
- School-Cotton 2y agoYou don’t need debug symbols to unwind the stack, you just need the .eh_frame section, which compilers emit by default regardless of whether you’re building with debug symbols. Source: I work on a profiler (Parca) that does stack unwinding. It works fine on Rust binaries with or without debug symbols.
- pornel 2y agoIt should be possible (it'd need to also save memory map), but for some reason Rust's standard library wants to resolve human-readable paths at runtime. Additionally, Rust has absurdly overly precise debug info. Even set to minimum detail, it's still huge, and still keeps all of the layers of those "zero-cost" abstractions that were removed from the executable, so every `for` loop and every arithmetic operation has layers upon layers of debug junk. External debug info is also more fragile. It's chronically broken on macOS (Rust doesn't test it with Apple's tools). On Linux, it often needs to use GNU debuginfo and be placed in system-wide directories to work reliably.
- starlite-5008 2y ago[dead]
- ekimekim 2y agoOk, so the original idea of Result<T, Error> was that you have to consider and handle the error at each place. But then people realised that 99% of the time you just want to handle the error by passing it upwards, and so ? was invented. But then people realised that this loses context of where the error occured, so now we're inventing call stacks. So it seems that what people actually want is errors that by default get transferred to their caller and by default show the call stack where they occured. And we have a name for that...exceptions. It seems that what we're converging towards is really not all that different from checked exceptions, just where the error type is an enum of possible errors (which can be non-exhaustive) instead of a list of possible exception types (which IIUC was the main problem with java's checked exceptions).
- tux3 2y agoIt does seem to be converging somewhere, but a major difference that I really like is pushing humans a little more to care about errors, instead of just letting whatever bubble up from wherever until a catch(...) somewhere. With checked exceptions, it's very common for the user to end up with only a cryptic message from a leaf function deep inside something, and that's very hard to interpret. Having a manual stack of meaningful messages that add context is so nice as a user. Even if I do get the stacktrace in a program that threw a deep exception, you typically won't understand anything as a user without access to the code, the stack trace for exceptions is just not meant for human consumption.
- shepmaster 2y ago> pushing humans a little more to care about errors This is 100% a reason that I like using SNAFU. The term I use for this is a "semantic stack trace" — a lot of the time, the person experiencing the error doesn't care that it occurred in "foo.rs" or "fn bar()" or "line 123". Instead, they care what the program is trying to do ("open the configuration file", "download the update file"). When I'm putting effort into my errors, I basically never use `snafu::Location` or `snafu::Backtrace`. My error stacks should always be unique — any stack can exactly point to a trace through my program.
- DavidWilkinson 2y agoInteresting approach! We had a similar journey at HASH to figuring out how we deal with stacked errors (as well as collecting parallel errors), developed the `error-stack` crate to solve for it. It works by abstracting over the boilerplate needed to stack errors by wrapping errors in a `Report`. Each time you change the context (which is equivalent to wrapping an error) the location is saved as well, with optional spantrace and backtrace support. It also supports supplying additional attachments, to enrich errors. We spent quite a bit of time on the user output, as well (both for `Debug` and `Display`) so hopefully the results are somewhat pleasant to work with and read.
- gregwebs 2y agoThis seems like a user implement of Zig error return traces: https://ziglang.org/documentation/master/#Error-Return-Traces https://ziglang.org/documentation/master/#Error-Return-Trace...
- k_bx 2y agoThe simplest thing that "just work" for me is replacing ? with .context(h!())? and this macro: #[macro_export] macro_rules! h { () => { concat!("at ", file!(), " line ", line!(), " column ", column!()) }; and then using anyhow::Result. Solves 99% problems in error handling
- dmart 2y agoUsing #[from] in a thiserror enum is an antipattern, IMO. I kind of wish it wasn't included at all because it leads people to this design pattern where errors are just propagated upwards without any type differentiation or additional context. You can absolutely have two different enum variants from the same source type. It would look something like: #[derive(Debug, Error)] pub(crate) enum MyErrorType { #[error("failed to create staging directory at {}", path.display())] CreateStagingDirectory{ source: std::io::Error, path: std::path::PathBuf, }, #[error("failed to copy files to staging directory")] CopyFiles{ source: std::io::Error, } } This does mean that you need to manually specify which error variant you are returning rather than just using ?: create_dir(path).map_err(|err| MyErrorType::CreateStagingDirectory { source: err, path: path.clone() })?; but I would argue that that is the entire point of defining a specific error type. If you don't care about the context and only that an io::Error occurred, then just return that directly or use a type-erased error.
- shepmaster 2y agoThis is one of the things I like about SNAFU: it makes this preferred pattern the default and makes it nicer to use. For example, your usage would look something like this with SNAFU: create_dir(path).context(CreateStagingDirectorySnafu { path })?; Note a few points: 1. No need to use the closure 2. No need to carry the source error over yourself (`context` does this for you) 3. No need to explicitly call `clone` on the path (`context` does this for you)
- carlsverre 2y agoInspired by this blog post I just added an `#[implicit]` field feature to the `thiserror` crate. It makes it easy to automatically annotate errors with things like code location (per this blog post), a timestamp, or a backtrace without requiring further modifications to the thiserror crate. I'm hoping that dtolnay will consider it. You can find my PR here: https://github.com/dtolnay/thiserror/pull/402 https://github.com/dtolnay/thiserror/pull/402