4 ms·
I'm only familiar with Rust's async story, though I think the following probably applies to other languages as well. Switching async tasks should have a smalle
by vlmutolo 5y ago
I'm only familiar with Rust's async story, though I think the following probably applies to other languages as well.
Switching async tasks should have a smaller overhead than switching threads. A context switch involves giving control to the OS and then the OS giving it back at some point. Both involve lots of cache trashing, as you said, and both switches involve bookkeeping that the OS has to do. This probably also involves instruction-cache evictions.
Switching async tasks involves loading the new task from memory onto the stack. That's it. The program can immediately start doing useful work again.
So, the short answer is that switching threads involves:
- more cache evictions
- bigger, slower cache evictions
- OS-level bookkeeping for threads
Also creating threads involves setting up a new, expandable stack in the process's memory space. Creating a new task involves allocating a fixed amount of memory on the heap once.
- ai-dev 5y ago> Switching async tasks involves loading the new task from memory onto the stack. That's it. One of the points on the list was that if the context switch is due to I/O readiness, then there's more work for the async task to do [0]: > Think about it this way—if you have a user-space thread which wakes up due to I/O readiness, then this means that the relevant kernel thread woke up from epoll_wait() or something similar. With blocking I/O, you call read(), and the kernel wakes up your thread when the read() completes. With non-blocking I/O, you call read(), get EAGAIN, call epoll_wait(), the kernel wakes up your thread when data is ready, and then you call read() a second time. > In both scenarios, you’re calling a blocking system call and waking up the thread later. 0: https://news.ycombinator.com/item?id=26110699 https://news.ycombinator.com/item?id=26110699
- roblabla 5y ago> this means that the relevant kernel thread woke up from epoll_wait() or something similar. With blocking I/O, you call read(), and the kernel wakes up your thread when the read() completes. With non-blocking I/O, you call read(), get EAGAIN, call epoll_wait(), the kernel wakes up your thread when data is ready, and then you call read() a second time. This is true for readiness-based IO (which, admittedly, most current async IO loops are using), but completion-based IO (such as IOCP or io_uring) don't suffer from this problem: you just add your IO operation to the queue, and do a single syscall that will return once one of the operation in your queue is completed AFAIK.
- ai-dev 5y ago> This is true for readiness-based IO (which, admittedly, most current async IO loops are using) Right, so if most async I/O frameworks use readiness-based IO (epoll, kqueue), then the context switch overhead is similar. So then most of the performance arguments for async I/O don't stand. That's where I'm confused :)
- drran 5y agoAsync programs are doing lazy evaluation, so they are better at discovering of critical path of execution. It's similar to out of order execution of instructions in CPU, but on a higher level. In theory, a compiler can (should) rearrange order of execution for non-async programs also.