4 ms·
> I'm having a production problem at work right now in which some tasks just get stuck. To mitigate this kind of problem, at my company we use a library [1] th
by rav 3y ago
> I'm having a production problem at work right now in which some tasks just get stuck.
To mitigate this kind of problem, at my company we use a library [1] that allows regularly logging which tasks are running and what file/line numbers each task is currently at. It requires manually sprinkling our code with `r.set_location(file!(), line!());` before every await point, but it has helped us so many times to explain why our systems seem to be stuck.
[1] https://github.com/antialize/tokio-tasks/blob/main/src/run_token.rs https://github.com/antialize/tokio-tasks/blob/main/src/run_t... has set_location(), and task.rs has list_tasks()
- scottlamb 3y agoYeah, I can see how that'd be helpful. In my case, I suspect this is happening inside a third-party library I'd rather not have to vendor/patch extensively. So that method could confirm my suspicion but probably wouldn't easily allow me to drill down as far as I'd like. That said, I think the newest version of the third-party library might have some middleware hooks and/or tracing spans. With the right middleware impl / tracing subscriber, maybe I could accomplish something similar. This code also should be following the general distributed systems practice of setting deadlines/timeouts at the top level of each incoming request, propagating through to all dependent requests, and also setting timeouts on background ops. It's not. Fixing that is also on my list and might be enlightening...