4 ms·
I realize you have a different usecase in mind, so the following feedback might not be that useful to you. I also haven’t reviewed your API design, so perhaps
by binary132 3y ago
I realize you have a different usecase in mind, so the following feedback might not be that useful to you. I also haven’t reviewed your API design, so perhaps you have already considered this angle.
I picked up a habit in Go package design which I now apply everywhere: one of my immediate first thoughts when evaluating a library is how it creates or uses threads.
When I see that a library wants to consume all of the available cores of my system, or uses an internal task scheduler or threading mechanism, I worry that it will compete for CPU time with the other processes, threads, etc. that my system and process are using.
On the other hand, I appreciate when a library exposes an interface for me to offer it threads or a scheduler, or even better if it offers primitives that I can use directly in my own threads or tasks.
Perhaps this is mainly a problem for high-concurrency servers and UIs, but it’s always one of my first thoughts, and Rust is an ecosystem which needs to support those usecases. Many threaded libraries, even otherwise very good ones, are not designed with this in mind and it is a problem.
- eigenvalue 3y agoIt's a good point, and the way my library is designed now it would be a big resource hog. Not sure the best way to handle that given that it's supposed to be a super easy to use library for Python. Perhaps I could add an optional parameter "no_resource_hogging" or something, and if enabled that can ratchet down the amount of parallelism, or reduce the CPU priority of the threads to a very low level so it won't use up all the system resources. The problem with the latter approach is that it starts being system specific. For example you could do something like this: fn create_low_priority_pool() -> rayon::ThreadPool { ThreadPoolBuilder::new() .start_handler(|_| { #[cfg(target_os = "linux")] { use libc::{sched_param, sched_setscheduler, SCHED_IDLE}; unsafe { let mut param: sched_param = std::mem::zeroed(); param.sched_priority = 0; sched_setscheduler(0, SCHED_IDLE, ¶m); } } #[cfg(target_os = "windows")] { use winapi::um::processthreadsapi::SetThreadPriority; use winapi::um::winbase::THREAD_PRIORITY_LOWEST; unsafe { SetThreadPriority(std::ptr::null_mut(), THREAD_PRIORITY_LOWEST); } } }) .build() .unwrap() } But I don't know how that would work for MacOS. Anyway, it's something to think about.
- zacmps 3y agoI would say default to min(affinity, system threads, 16) and let the user input a specific number to override it. I had a problem recently where a library started a process per CPU thread I suspect the author didn't intend it to start start 256... Affinity is only easy to get on some platforms but you should use it instead of the system threads if available.
- binary132 3y agoOne approach I’ve used in the past is to offer an API which takes threads or a worker pool as a parameter, and a separate API which does that for you and calls the first one. Then in your Python wrapper you can use the wrapper, but your users who care can provide the pool they’ve already created. Bonus points if it’s an interface or trait that the user can adapt their own scheduler to, if they don’t use rayon. But, I can see how that could be tricky if the design of your library is tightly coupled with the threading solution. One way to frame it is that thread creation is a significant side effect.