4 ms·
I was going to talk about some panics we've had recently, but those will hopefully be fixed soon so I don't think that quite fits my message. Instead I want to
by jynelson 6y ago
I was going to talk about some panics we've had recently, but those will hopefully be fixed soon so I don't think that quite fits my message. Instead I want to talk about why 'handling errors correctly and not panicking' wouldn't work in general.
In order for that work, we'd have to have 0 panics - not just few, but none. That requires none of our code to panic, none of our dependencies to panic, and none of our uses of the standard library to panic. If you gave me a limited time frame to run the server - say a week - I think it's possible to make docs.rs that robust. However, for a server that's meant to run 24/7 for weeks on end, I just don't think that's realistic.
What `catch_unwind` lets us do is localize panics to a single web request or crate build instead of it affecting the whole server. Of course we don't want to have 500s for any user, but we _especially_ don't want the whole server to be down until a team member has time to ssh in and restart it manually.
Of course, after that there's whole question of whether making the server that robust is a good use of time in the first place. Is it better to fix a few 500s every week or to fix long-standing bugs? I don't think it's clear that the 500s are more important if they don't affect many users.
- nine_k 6y agoLook how Erlang handles this. Crashing on error is the encouraged policy. A managing process will notice a crashed process and restart it. Basically crashing is the safe way to release all resources in a problematic situation, at the cost of terminating the process. It's easy when you don't have shared resources at all (Erlang's case), and harder with threads: one threads crashes and frees its resources, another tries to use a shared resource already deallocated. If this can be avoided, life becomes vastly easier, at the cost of higher resource consumption.
- jynelson 6y agoI think we're saying the same thing from two different perspectives :) The 'managing process' is the daemon thread. The 'crashed process' is a worker thread. There's no need to worry about corrupted memory since all state is shared through the database.
- fluffything 6y agoThe main practical difference is that if you use a process for isolation, the OS will garbage collect all process resources on termination (memory, file descriptors, threads, mutexes, etc.). If you are using a thread, you better not be leaking anything. Otherwise you are "ulimits" resources away from putting your "main" process into a crashing loop, e.g., if you leak 1 file-descriptor per crash, then you can crash your thread < 1000 times. Performance-wise, you are probably worse with threads as well. On linux, you can initialize a web-server on your main process, and spawn new processes by forking it. Forking isn't only pretty much instantaneous, your child process is initialized with the same state as the parent, so you instantaneously get a fully initialized web server (e.g. with multiple threads already started in your task pool, etc.).
- jynelson 6y agoHmm, this is an interesting point about the OS cleaning up resources. I'll see if it's feasible to switch from threads to processes. Thanks!