4 ms·
Thank you for this article. The Rust and tokio folks are working on difficult and complex problems, I appreciate and thank them for the work they're doing to i
by samsquire 3y ago
Thank you for this article.
The Rust and tokio folks are working on difficult and complex problems, I appreciate and thank them for the work they're doing to improve server and desktop app performance everywhere. We all have multicore machines and it would be great if we could use more than 1/8 or 1/12 (or whatever high thread count of your beefy servers) of our hardware.
The Rust multithreaded thread memory management (and Arc and so on)* causes me to be uncomfortable because of a key lesson I've learned is that you cannot scale a program by adding threads and expect it to accelerate mutate access to the SAME memory location. Single threaded memory mutation performance is a fixed known quantity. Adding threads with contention for same memory location causes throughput and latency to be slower to a particular memory location at single threaded speeds because you need mutexes or a lock free algorithm to communicate safely.
To accelerate data fanout or storage (writing to memory from multiple threads) you need a shared nothing architecture or sharding.
This means that when you reach for threads and I'm guessing you're wanting to reach for threads for acceleration and performance you need to design your data structures to not share memory locations. You need to shard your data.
EDIT: I originally said Sync + Send.
- kzrdude 3y agoSend + Sync is maybe not enough but wouldn't it be possible to say: as long as the memory location is read-only you can parallelize access to it. Send + Sync helps pass the read-only data through without synchronization to all threads, while the rest of the Send and exclusive mutability system flags the tricky points for you. I can see that Send/Sync by itself does not tell you if the data is read only or just internally synchronizing mutation.
- nu11ptr 3y ago> We all have multicore machines and it would be great if we could use more than 1/8 or 1/12 (or whatever high thread count of your beefy servers) of our hardware. It is probably important here to realize that async solves concurrency, not parallelism. You can use async with a single threaded runtime for I/O concurrency and mix that with threads for computational parallelism for long running jobs. That said, there may be some benefit a multi-threaded runtime would have for the typical I/O bound app (after working around lifetime limitations by adding Send/Sync to data structures). This is because I/O bound programs and those requiring computation are not mutually exclusive and there is always some amount of computation going on, so there may still be some benefit. I doubt a synthetic benchmark would answer this as those typically don't measure any actual work performed, but just "requests/sec".
- bkolobara 3y ago> It is probably important here to realize that async solves concurrency, not parallelism. You can use async with a single threaded runtime for I/O concurrency and mix that with threads for computational parallelism for long running jobs. In my experience, it's impossible to mix threads and async tasks. They can't communicate or share state. Threads need locks, while async tasks require an awake mechanism. If you just stick to unbounded channels that don't block on send, you can get far, but in 99% cases you will need to decide upfront on a specific approach.
- thinkharderdev 3y agoThis has not been my experience at all. Delegating compute-intensive tasks to rayon inside a tokio runtime is not particularly hard (assuming you can pipeline things to separate IO and compute effectively). A pattern that has worked quite well for me is to use ``` struct ComputeTask { some_state: SomeState, completion: Sender<Output>, } impl ComputeTask { fn run(self) { .. do you compute-intensive stuff self.sender.send(output); } } async fn do_stuff() -> Result<Output> { let (tx,rx) = tokio::sync::oneshot::channel(); let task = ComputeTask { .., tx } rayon::spawn(move || task.run()); rx.await } ```
- galangalalgol 3y agoI still don't understand why async is faster. Sharding data can be as simple as a buffer per thread in a thread pool to catch incoming data. With async, don't you havev to allocate that input buffer each time? That seems hideously expensive.
- samsquire 3y agoI enjoyed this article from Cal Paterson my excolleague. It's about Python async not being faster: https://calpaterson.com/async-python-is-not-faster.html https://calpaterson.com/async-python-is-not-faster.html I think the idea is that while your blocking waiting for IO in one task you can serve a different task, potentially from a different user. Coroutines, green threads, communicating sequential processes as in Go or Occam.
- jerf 3y agoPure Python is a very slow language compared to Rust, with significant differences in orders of magnitudes of expenses. I would not expect information about Python performance to be particularly relevant to Rust without further evidence directly from Rust.
- galangalalgol 3y agoI think in retrospect, it makes sense to me that if you are io bound vs cpu bound (like my stuff usually is) that async could let you wait on more things at a time.
- jerf 3y agoI think the whole "IO bound" thing has taken on a life of its own and attained a legendary status that is not always an accurate reflection of reality. People often seem to model things as if "waiting on the DB" is all their system does and the code they wrote executes in exactly 0 nanoseconds, but that's not how it works. It isn't actually that hard to talk to a relatively local database with some well-optimized query and be doing CPU work either comparable to the wait you spent on the DB, or even greatly exceeding it, at which point your language's performance in fact does matter, potentially even dominates.
- tcfhgj 3y agoShared mutated memory isn't necessarily a problem, because it still is unknown how often access is required to that memory. E.g. having a thread that spends perhaps 1% of the time with state mutation vs 10 threads spending each 3% of the time with state mutation. You have smaller efficiency per thread, but still higher efficiency overal
- samsquire 3y agoThis reminds me of the whitepaper Scalability! But at what cost? http://www.frankmcsherry.org/assets/COST.pdf http://www.frankmcsherry.org/assets/COST.pdf Which I think is about how single threaded programs are faster than scalable but slow multithreaded systems. I think you might be right and that's why there is ReadWriteLock or RwLock for single writer taking turns. If you have a counter or a data structure you're mutating in every request then you'll hit lock contention.
- eximius 3y agoIf you're just using Arc<T> without any other parallelism primitives, then it's immutable and the cores can all read without blocking. The only thing it does is reference counting to know when to Drop. Blindly using Arc<Mutex<T>> without considering access patterns is a software architecture problem, not a problem with Arc or Mutex.