4 ms·
Well, rarely done isn't the same as never done. docs.rs has exactly the tradeoffs the parent comment is talking about: we have a long-running daemon thread that
by jynelson 6y ago
Well, rarely done isn't the same as never done. docs.rs has exactly the tradeoffs the parent comment is talking about: we have a long-running daemon thread that uses `catch_unwind` and builder threads that occasionally panic. We've had [issues with memory leaks](https://github.com/rust-lang/docs.rs/issues/656 https://github.com/rust-lang/docs.rs/issues/656) in the past - they weren't related to unwinding that I know of, but it's still possible that they were.
However I'm not in favor of the proposed solution - if the docs.rs server aborted every time a thread panicked we would have a lot of outages!
- steveklabnik 6y ago> Well, rarely done isn't the same as never done. Absolutely. But I do think the difference between "idiomatic" and "rarely done" is valuable, here. And as you said yourself, it's not clear that these leaks are caused by this kind of thing. If someone saw Rust code leaking memory all the time due to catching panics, I'd want to know about it, because it's very contradictory to my own experience, and I think that's interesting.
- cogman10 6y agoCorrect me if I'm wrong, but panics themselves are fairly rare. That's not the case in C++ where "throw" is a keyword you are expected to use for error handling. If you take this to a Java example, a panic is more like an "Error" and less like an "Exception". That is to say, in rust, if something panics it is a sign of a major bug that shouldn't have any option for recovery. With all that said, the same concept exists in C++. Try doing a 1/0 in C++ and see what happens. You don't get some nice exception to catch and you can't add a "noexcept" clause to stop it from happening. In the same original example, if you have that 1/0 error in the underlying method and you handle it instead of crashing by tying into the OS specific "div by zero" garbage. You to can create a memory leak in code thought to be safe. It goes to show that you can't (and shouldn't) expect to recover from everything. Crashing, IMO, is usually far safer than trying to fix things up and move forward.
- steveklabnik 6y agoConceptually, panics should be rare, because one firing means that some sort of unexpected problem has occurred. However, the real world is not "conceptually." I don't think there's any real data about how often they happen, but at least my experience is that tools I write in Rust rarely end up showing me panic output. > That's not the case in C++ where "throw" is a keyword you are expected to use for error handling. Yes, that's correct. The intended semantics of the two features are very, very different.
- matklad 6y agoA fun example here is rust-analyzer: we implement cancellation via unwinding. This is not technically a panic, but the mechanism is the same, and it more or less is invoked every time a user types something in a file.
- fluffything 6y agoWhere and why does rust-analyzer unwind from destructors ?
- richardwhiuk 6y agoThe solution for docs.rs is handling errors correctly and not panicing.
- jynelson 6y agoI was going to talk about some panics we've had recently, but those will hopefully be fixed soon so I don't think that quite fits my message. Instead I want to talk about why 'handling errors correctly and not panicking' wouldn't work in general. In order for that work, we'd have to have 0 panics - not just few, but none. That requires none of our code to panic, none of our dependencies to panic, and none of our uses of the standard library to panic. If you gave me a limited time frame to run the server - say a week - I think it's possible to make docs.rs that robust. However, for a server that's meant to run 24/7 for weeks on end, I just don't think that's realistic. What `catch_unwind` lets us do is localize panics to a single web request or crate build instead of it affecting the whole server. Of course we don't want to have 500s for any user, but we _especially_ don't want the whole server to be down until a team member has time to ssh in and restart it manually. Of course, after that there's whole question of whether making the server that robust is a good use of time in the first place. Is it better to fix a few 500s every week or to fix long-standing bugs? I don't think it's clear that the 500s are more important if they don't affect many users.
- nine_k 6y agoLook how Erlang handles this. Crashing on error is the encouraged policy. A managing process will notice a crashed process and restart it. Basically crashing is the safe way to release all resources in a problematic situation, at the cost of terminating the process. It's easy when you don't have shared resources at all (Erlang's case), and harder with threads: one threads crashes and frees its resources, another tries to use a shared resource already deallocated. If this can be avoided, life becomes vastly easier, at the cost of higher resource consumption.
- jynelson 6y agoI think we're saying the same thing from two different perspectives :) The 'managing process' is the daemon thread. The 'crashed process' is a worker thread. There's no need to worry about corrupted memory since all state is shared through the database.
- fluffything 6y ago> However I'm not in favor of the proposed solution - if the docs.rs server aborted every time a thread panicked we would have a lot of outages! Where is the proposed solution? Where and why does docs.rs unwind from destructors?