4 ms·
> Even before getting into SIMD, try using Rust for concurrent, succinct, or external-memory data structures. It quickly becomes clear where the friction is. I
by pcwalton 2y ago
> Even before getting into SIMD, try using Rust for concurrent, succinct, or external-memory data structures. It quickly becomes clear where the friction is.
It's the exact opposite for me: I use concurrent data structures more often in Rust than I do in C++ because I don't have to worry about dumb data race bugs. If one of my Bevy systems is slow, I slap par_iter() on the query and if it compiles it probably works, or at least fails for a not-stupid reason.
- SleepyMyroslav 2y agoWhat do you think about 'par_iter' having to wait for work imbalance to return execution back? With 'when_all' like primitive one can continue execution on any thread without losing one for waiting. ps as someone who does not have rust job i would like to see an example how rust deals with task based systems if you have public one at hand ofc.
- pornel 2y agoThe rayon library uses work stealing for this. Its parallel iterators offer some control of splitting strategies and granularity, so you can tune the trade-off between full utilization and cost of moving things between threads. Additionally, in Bevy, independent queries (systems) are executed in parallel, so there's always something else to do, except your one worst loop.
- pclmulqdq 2y agoConcurrent data structures are rarely faster than "lock + non-concurrent data structure" and if you're putting constructs like par_iter() in a lot of places, there's a good chance you would be better off with the "dumb" pattern than the concurrent data structure. The same goes with Arc - if you're using it a lot there's a good chance that code with a lot of Arcs is slower than equivalent GC-ed code.
- 1932812267 2y agoWhile it's true that par_iter() uses a concurrent data structure under the hood, it's specifically designed to use work-stealing to avoid needing threads to communicate. Why would putting a lock over a global workqueue be faster than per-thread workqueues that don't require inter-thread communication (except in the case where work-stealing is required)?
- pclmulqdq 2y agoAtomics are very expensive operations. Lock/unlock is two atomics. Many concurrent data structures will end up doing many more atomic operations than you expect. The general wins of concurrent data structures come when you really are accessing them truly concurrently - as in when many threads on many cores are heavily contending for access and you need to make global progress.
- 1932812267 2y agoSure! However, the work-stealing queue in rayon [1] uses three atomic operations instead of the two atomic operations for a mutex for a global lock. The difference, however, is the three atomic operations for the thread-local queue should be uncontended, whereas a global lock on a global work queue would experience contention from every thread trying to access it for jobs. Between the choices of "single work sharing queue with a big mutex on it that all threads access for work" vs "per-thread work-stealing queue that's uncontended for the cost of one extra atomic," in what situations would the work-sharing queue with the global mutex outperform? Perhaps if there's a small number of jobs, and there's not enough time for the work-stealing algorithm to distribute jobs to the worker threads before the work-sharing algorithm has already finished. [1]: https://github.com/crossbeam-rs/crossbeam/blob/423e46fe204718785af99a2d68a52092463d0167/crossbeam-deque/src/deque.rs#L444-L481 https://github.com/crossbeam-rs/crossbeam/blob/423e46fe20471...
- pclmulqdq 2y agoRun a benchmark. With low contention, the lock will outperform. Atomics are very expensive assembly instructions. Fedor Pikus has a good talk on this at cppcon 2019.