4 ms·
Is it? I though that work stealing was a basic requirement for maintaining some semblance of cache locality.
by thinkharderdev 4y ago
Is it? I though that work stealing was a basic requirement for maintaining some semblance of cache locality.
- brrrrrm 4y agoFor what kind of work?
- pjscott 4y ago“Lots of tiny tasks” workloads, mostly. For more coarse-grained parallelism the cache thing matters a lot less. I still favor work-stealing schedulers as the default, since robustness across different types of workload is a very good thing, but they’re definitely not always needed.
- klempner 4y agoIf you have items running in the threadpool that are submitting more work to the threadpool, and you care about "high-performance", you probably want those items to run on the same thread for better cache (and memory locality) behavior. Really you want the same/nearby CPU, but in pure C++ "same thread" and relying on the OS is the closest you can get. It's not that you need "work stealing" but rather that you need separate queues per thread rather than a single central queue. At that point you need a work stealing story to keep a single long running item from blocking the work queued behind it. This matters less on systems which are "less NUMA" -- a single socket 6 core Intel home computer isn't going to care nearly as much if work migrates between CPUs unnecessarily, but the cache/NUMA effects are quite large when dealing with multi-socket servers, especially if you have a threadpool that wants to utilize the entire machine.
- PixelOfDeath 4y agoIt's also about prefetchability of the next task. If all cores e.g. do an atomic add on a single counter variable to get their next task-id just in time, then there is no chance for the cores to do any reasonable prefetching. Because most of the time another core will come in and write to this variable that all prefetching is based on. With short tasks, you can end up with a full pipeline stall each time a new task is started.