4 ms·
> Applications that create too many threads that are constantly fighting for CPU time (such as Apache's HTTPd or many Java applications) can waste considerable
by otabdeveloper1 10y ago
> Applications that create too many threads that are constantly fighting for CPU time (such as Apache's HTTPd or many Java applications) can waste considerable amounts of CPU cycles just to switch back and forth between different threads.
Completely wrong. Context switches happen at a set interval in the kernel. Creating more threads won't make context switches happen more frequently, each thread will just get a proportionally smaller share of CPU time.
Conceptually, the CPU scheduler can be thought of as a LIFO queue. An interrupt fires at a fixed interval of time and switches to the first ready thread (or process) on the queue.
What's more important is the fact that this queue is exactly equivalent to how 'asynchronous' operations are implemented in the kernel.
The only difference between creating threads and using async primitives is the thread prioritization algorithm. With threads you're using whatever scheduling heuristics are baked into the kernel, while with async you can roll your own. (And pay the penalty of doing this in usermode.)
- jandrewrogers 10y agoThis isn't really correct. In many (most?) multithreaded applications a thread will regularly yield, forcing a context switch, long before its kernel time slice expires. In modern systems it is not difficult to have a case where the time spent doing work is less than the time spent yielding to another thread i.e. almost all of the CPU time is spent context switching. This is in significant part a side effect of CPUs becoming much faster. You need to increase the amount of code that is run between potential thread yield points to offset the overhead of a yield or it can generate context switch storms (i.e. almost all time is spent context switching) and throughput will suffer. Modern server architectures take this to its logical conclusion and try to eliminate virtually all context switching, hence the rising prevalence of "process-per-core" software designs.
- wzdd 10y ago> Creating more threads won't make context switches happen more frequently This is generally wrong in practice, but for a slightly subtle reason. 1. You can force a context (or thread) switch whenever you like by performing a blocking operation: calling read() for example, or by waiting on a mutex (as this person was doing). There's no more work to be done until the blocking operation completes, so the kernel finds another process (or thread) to run on the now-spare hardware thread. 2. Applications which create a lot of threads or processes tend to use them as IO workers. 3. The IO worker pattern is to block on read(), do some processing, maybe write(), and then go back to blocking. This means that most of the time the worker thread won't be using its complete timeslice. 4. If lots of IO worker threads are continually doing this, you will get a lot of thread switching (or context switching for processes). 5. Thread switching is fast in Linux, but it's slower than a simple syscall, which is what you'd be doing with an async solution.
- otabdeveloper 10y agoRe: your point 1 -- waiting on a mutex happens regardless of whether you're using blocking or non-blocking IO. If your program only ever makes IO system calls and nothing else, then your point might make sense, but this isn't a realistic real-world assumption. Re: your point 3 -- not using the complete timeslice is not a realistic assumption. If you're creating threads then you're presumably doing some little bit of CPU-intensive work and not just copying bytes from one file descriptor to another. You're making terribly silly assumptions about what kind of work servers actually do, and ignoring the real benefit of async solutions. The real benefit is being able to control how threads switch by yourself instead of using the kernel's builtin 'black box' scheduling algorithms. The problem with the 'black box' is that the kernel might decide to penalize your threads for inscrutable reasons of 'fairness' and then you suddenly get inexplicable latency spikes. Of course rolling your own scheduling is an engineering boondoggle and most people just opt for a very primitive round-robin solution. (Which, incidentally, is what you want anyways if you want good latency.) In which case you might as well create a bunch of threads and schedule them as 'real-time' (SCHED_RR in Linux) and get the same result. (Seriously, try it -- benchmark an async server vs a SCHED_RR sync server and see for yourself.)
- chrisseaton 10y agoDoes the CPU interrupt threads running happily on cores even when there are no other threads which want to run or which have affinity that would allow them to run on that core? But creating more threads will cause each to run slower as their caches are ruined by each context switch aren't they?
- otabdeveloper 10y ago> Does the CPU interrupt threads running happily on cores even when there are no other threads which want to run or which have affinity that would allow them to run on that core? Yes. Even if you are careful to ever run only one process (so: no monitoring, no logging, no VM's, no 'middleware', etc.) and limit the number of threads to strictly equal the number of processors, you still have background kernel threads that force your process to context switch.
- gpderetta 10y agoTough you can instruct the kernel not to run anything (not even interrupt handlers) on specific cores, except for manually pinned processes.
- marcosdumay 10y agoYes, it does interrupt them, but if the kernel decides to keep the same thread there's no context switch. (In other words, there's a penalty, but it is much smaller.) Your conclusion is correct. By creating more CPU hungry threads than you have CPUs you will reduce the total throughput of the computer.
- jabl 10y agoNo, you're wrong. Or, you're correct as far as you have a bunch of threads doing only CPU-bound work (and even in that case having fewer threads can be useful due to cache effects). But if you have threads that block, then there will be more context switches. Consider something like apache httpd: Worker thread/process A writes to a socket, then the socket buffer fills up before the data has been sent over the network, the process blocks and goes to sleep (before it has used its time quantum and is preempted). CPU scheduler switches to apache worker thread B. B writes to its socket, socket buffer fills, B goes to sleep causing yet another context switch to apache worker C, etc. etc. See? Lots of context switches. Compare this to something like nginx: Worker process A writes to socket A1, socket A1 buffer fills up (errno EAGAIN, but no context switch), A switches to writing to socket A2, etc. Or for that matter consider a multithreaded process that manipulates some shared state. Thread A takes a lock X, and then the scheduler interrupt fires and the CPU scheduler decides to put A to sleep and switch to B. Well, B works for a little while, then tries to acquire lock X which is locked, and thus goes to sleep. CPU scheduler wakes up thread C, which works for a little while, then tries to acquire lock X. Oops, X is locked, so C goes to sleep and the scheduler wakes up thread D. Etc. etc. Again, lots of context switches.
- jsolson 10y agoYou are incorrect. If you have many threads each doing small amounts of work before blocking, the kernel will happily switch in another runnable thread when the current thread blocks. This is common in workloads that service network IO using a 1:1 thread:client model. It can absolutely be more efficient to retire work for multiple clients from a single thread. Your last paragraph ignores this. It also forgets that in addition to prioritization there is multi-core load balancing to deal with. That can leave runnable threads stranded for whatever the LB interval is (which is independent of the overall scheduler quantum), while a concurrent service queue would allow steady service of all clients from all workers (that said, I actually prefer the siloed model in most cases).