4 ms·
Thank you for this blog post. I like the thread per core design because kernel context switches are expensive. But userspace context switching scheduling (such
by samsquire 3y ago
Thank you for this blog post.
I like the thread per core design because kernel context switches are expensive. But userspace context switching scheduling (such as an unbuffered golang channel) for a 64 bits of data at a time makes me uncomfortable too unless it represents a large amount of work backed by a pointer to work data.
I think I like to multiplex sockets over threads and multiplex IO over threads so that you can do CPU work while IO is going on and you can process multiple clients per thread.
You cannot scale memory mutation by adding threads. You want ideally one thread to own the data and be safe to mutate it uncontendedly.
When you send data to another thread, don't refer to that data again. Transfer ownership to that thread.
If you can divide your request into phases of expansion (map) and contraction (synchronization) you can do intrarequest parallisation.
I've been looking into runqueues of go and tokio where you have a local runqueue without a mutex and a global runqueue with a mutex.
I've been trying to think how IO threads that run liburing or epoll can wake up a Coroutine or an async task on a worker thread without the mutex.
The worker thread is looking for tasks to resume that are unblocked and needs to be notified when there is IO finished. I think you can have a Coroutine that is always runnable to read from a lock free ringbuffer. You can have that Coroutine that checks for finished IO's ringbuffer yield if its contended by writes by the IO thread that is trying to make work available to the worker threads.
- ori_b 3y ago> I like the thread per core design because kernel context switches are expensive You can do a millions to tens of millions kernel entries/exits per second per core. They're on the same order of magnitude as a single digit number of full cache misses. So, expensive, but a lot less expensive than many people assume. Edit: for some (outdated, 2018) measurements, https://eli.thegreenplace.net/2018/measuring-context-switching-and-memory-overheads-for-linux-threads/ https://eli.thegreenplace.net/2018/measuring-context-switchi...
- cmrdporcupine 3y agoAnd all our assumptions get blown up with every new hardware generation, too. Hardware engineers are building systems for the applications of today. And the kernel developers are optimizing as well. It's all a moving target. Dogma has no place here.
- kldx 3y agoThat's an interesting view and one I agree with. Surely if context switch = expensive is common knowledge for a long time in computing, the kernel people and CPU companies would have tried optimizing it to the fullest, no?
- insanitybit 3y agoSyscalls have gotten quite a lot slower after Specter/Meltdown mitigations. Though I suspect that if you're writing a web service that wants TPC you should just disable these - I've told that to Scylla Cloud before, they have no reason to keep them enabled and plenty of reasons to disable them. The other thing is 'io_uring', which is trying to come at the problem by removing context switching altogether by providing a cheap primitive for communicating with the kernel.
- kldx 3y agoSadly io_uring has its share of vulnerabilities as well
- insanitybit 3y agoOf course, I'm very aware of that fact. But it is designed to minimize system calls exactly because they can add significant overhead.
- amluto 3y agoYou’re replying to a comment about context switches, not syscalls. Linux on x86 cannot do tens of millions of context switches per second per core. (I expect that, at high clock speed, carefully tuned, under ideal and possibly unrealistic conditions, you might get 2M/s — it’s been a while since I benchmarked this.) But this is silly — shared nothing will have no substantial difference in the rate of context switches (or of syscalls) compared to a thread-per-core-shared-everything architecture. The architecture that actually loses is many-threads-per-core if it ends up in a small-batch mode in which it context switches once per request or so.
- deleted 3y ago[deleted]
- TexasMick 3y agoI have seen context switches be a killer in some embedded contexts. Especially on FPGA soft core applications where you often are dealing with heavy I/O, if you just naively spin up threads for I/O then you will suffer heavily.