4 ms·
>You can also write "actual" synchronous code which does the same thing in the kernel. :P And there are several CPU schedulers to choose from to fine-tune when
by acconsta 11y ago
>You can also write "actual" synchronous code which does the same thing in the kernel. :P And there are several CPU schedulers to choose from to fine-tune when things get woken up.
You can. But no matter how good your kernel scheduler is, that trip through it is going to cost you. Why? Because it's going to wreck your cache and TLB entries. I think the TLB is a big part of the story, given it's small, specific to the process, and particularly expensive to miss.
http://www.cs.cmu.edu/~chensm/Big_Data_reading_group/papers/flexsc-osdi10.pdf http://www.cs.cmu.edu/~chensm/Big_Data_reading_group/papers/...
Anyway, I'm sure you know this already. :) Out of curiosity, did you get a chance to take CS169 before you left Brown?
And I think the kernel already does pool/recycle stacks, but I guess not 200,000 of them given your results. An explicit thread pool would certainly work too.
- comex 11y agoKernel stacks are allocated by Linux in alloc_thread_info_node; user stacks by glibc in allocate_stack, which does have some sort of cache. Overall there's a fair bit of work going on that could be skipped... But yeah, certainly userland context switches are faster; it's just that your comment seemed to postulate "scheduling in response to IO events" as a benefit of coroutines over threads, as if the kernel had no way to distinguish between threads waiting for IO and threads needing to be woken up. I presumably misinterpreted it. Interesting paper. It seems like it could be a useful step toward my ideal imagined environment where the distinction between coroutines and threads would be meaningless - because scheduling would be a job shared between userland and the kernel, so userland could do fast thread switches, batch system calls, etc., while the kernel would still give them PIDs and let them take signals or be ptraced, and native tools (debuggers) would treat them as threads. Though the paper is from 2010; do you know if anyone has tried to implement anything along similar lines in production? And no, I didn't take that course.
- acconsta 11y agoOh OK, I definitely could have written that better. From what I can tell, M:N threading has been thoroughly abandoned by Linux, but it might be worth revisiting given the horrendous I/O scaling of Linux kernel threads. I would imagine scheduling user threads cooperatively simplifies things a lot. Google has a very interesting talk about cooperative native thread scheduling, but they haven't upstreamed their code: https://www.youtube.com/watch?v=KXuZi9aeGTw https://www.youtube.com/watch?v=KXuZi9aeGTw And here's what linux devs are actually doing (it's underwhelming, something like 50% of what the hardware is capable of): https://lwn.net/Articles/629155/ https://lwn.net/Articles/629155/ In the meantime, people who actually need the performance (HPC and even some enterprise servers) have to bypass the kernel entirely. It looks like that's going to be possible even in virtualized environments, thanks to hardware support: http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14.pdf http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14....