4 ms·
There's no such thing as a "context switching overhead". You don't understand how operating systems work. The OS context switches to the next available thread
by otabdeveloper1 8y ago
There's no such thing as a "context switching overhead". You don't understand how operating systems work.
The OS context switches to the next available thread at fixed time intervals. It doesn't matter if you have 5 threads or 5000, the number of context switches is the same.
The only difficulty is deciding which thread is the "next available" one. This is the job of the process scheduler.
Normally, the process scheduler is a complex heuristic affair that tries to pick a more deserving thread based on some black magic rules of thumb. Under certain loads this black magic can backfire, picking some threads more frequently while letting others starve.
If you know your workload, you can use a scheduler without magic heuristics. These kinds of schedulers are called "real time". Try them, you'll be surprised.
- viraptor 8y ago> There's no such thing as a "context switching overhead". So what do you call the time spent changing the system internal information about the currently executing thread, switching page tables via cr3 and the invalidation it causes, and the time wasted because of cache misses (although a lot of it can be reclaimed with CPU pinning). Because that's what people normally call "context/thread switching overhead" > The OS context switches to the next available thread at fixed time intervals. It doesn't matter if you have 5 threads or 5000, the number of context switches is the same. Not just at fixed intervals. If your threads are doing io work, more threads will switch more often because they will issue io requests forcing waits/yields. If they do pure-cpu work, that doesn't apply, but we're talking about networking mostly in these comments.
- ezdiy 8y agoTLB is only an indirect cause. This is because kernel scheduler preempts processes fairly infrequently (100 or 1000hz, or dynamic, but still capped to a small number). Scheduling quantums are so large precisely to keep TLB flush overhead of a context switch low. If a network mandates more interaction (say, 100k req/s across all workers), each quantum tick must process a queued bundle of 1000 requests which piled up while asleep. This works as designed - you're supposed to use up all of your quantum, and not terminate it early by issuing blocking IO per request. One prerequisite for this is that your network/disk protocol must be pipelineable (most are because thats how we deal with network/seek latencies). But at certain point the overhead of this pipelining itself becomes so great (message queues too deep) you have to switch to threading. Hardcore threading advocates on the other hand, need to account for overhead of atomics (for locking, or for "lockless" algorithms). An atomic must wait for all pending writeback flush. Threading gets a lot of bad rep not because "kernels suck at it", but because person making such a statement wrote their program as an exercise in lock contention and/or too much write cache pollution per single atomic. Threading vs process tradeoff = deep pipeline overhead vs frequent queue flush+locking overhead tradeoff. Typically, you need to meet somewhere in the middle for best performance, which is when you end up with threads with job queues - those basically emulate process-induced queues within thread model.
- dmitrygr 8y ago> Threading gets a lot of bad rep > not because "kernels suck at it", > but because person making such a > statement wrote their program as > an exercise in lock contention Well put. I'm going to have this printed on a plaque and hung above my desk.