3 ms·
I think the argument against thread-per-core as the default can be made simply: - if you are CPU bound, work stealing will be better for most cases - if you a
by ko27 3y ago
I think the argument against thread-per-core as the default can be made simply:
- if you are CPU bound, work stealing will be better for most cases
- if you are IO bound, thread-per-core might work better, but again, you have enough CPU headroom that the performance doesn't really matter
IMO, work stealing is a better default to encode into language API
- layer8 3y ago> you have enough CPU headroom that the performance doesn't really matter But it would affect power consumption? (Just trying to understand.)
- basro 3y agoNo, performance and power consumption should go hand in hand in this case. If you are strongly IO bound, paying for the synchronization is not really going to matter much I believe. There are cases where you can be CPU bound and using the share nothing model would work out to your advantage. There's also the case where you only have one cpu core anyway (for example if you want to get all the juice out of a cheap single core VPS)
- vlovich123 3y agoThe main argument for work stealing is that it’s hard to achieve uniformity of work loads across all threads. The main argument for a single-threaded thread per core design is that it’s easier to code AND performs/scales way better than work stealing (including average and tail latencies, TPS etc). IMHO it’s a misconception that this is somehow tied to CPU or IO bound work. Take for example databases. You’d think that that’s the prime “I/O bound” use case. Except it’s not. There’s a talk about a DB researcher that analyzed that Postgres spends 70% of its time book keeping things within the database. And that makes sense. I/O is done in bulk with the cost amortized over a lot of transactions. That book keeping work? Extremely expensive because you have to acquire locks all over the place, do atomics, memory allocations etc. Atomics and memory allocations are extremely expensive in certain contexts and atomics also have a negative in that your scaling with number of CPUs is sub linear due to hard to remove false sharing of cache lines and CPU stalls to handle the synchronization. On the other hand, a shared nothing approach where you’re not allocating memory in your hot path is very hard to achieve and not suitable for all problems. Nor does everyone need that performance. So the work stealing approach is better in those use cases as it provides reasonable performance and the programming model is simpler in some ways since you don’t have to think about the data path as careful since everything has an Arc / Mutex in there.