6 ms·
I wonder if this kind of approach could be generalized to all syscalls - pushing them asynchronously to shared memory, then receiving output in arbitrary order
by d33 8y ago
I wonder if this kind of approach could be generalized to all syscalls - pushing them asynchronously to shared memory, then receiving output in arbitrary order once kernel core takes care of it. Is this feasible? Does it make sense to expand this beyond IO? My reasoning is that we already usually have more cores than we need, why not dedicate some specifically for kernel stuff?
- hornetblack 8y agoThere's a concept called exception-less syscalls, which might be similar to what your thinking.
- d33 8y agoMy idea is to avoid context switching like in VDSO, but with a generalized solution that would allow this sort of behavior for all system calls. Could you point me to any write-ups about exception-less syscalls?
- anonymousDan 8y agoSee the FlexSC paper at the bottom here: https://www.usenix.org/conference/osdi10/flexsc-flexible-system-call-scheduling-exception-less-system-calls https://www.usenix.org/conference/osdi10/flexsc-flexible-sys... or more recently SCONE in the context of Intel SGX: https://www.usenix.org/conference/osdi16/technical-sessions/presentation/arnautov https://www.usenix.org/conference/osdi16/technical-sessions/...
- int0x80 8y agoThere is mmap for read/write (some of the most used syscalls). There is also the vdso page [1] for things like gettimeofday etc. [1] http://man7.org/linux/man-pages/man7/vdso.7.html http://man7.org/linux/man-pages/man7/vdso.7.html
- blattimwind 8y agoThat's almost literally IOCPs.
- d33 8y agoInteresting. Where could I read up on that?
- rrdharan 8y agohttps://en.wikipedia.org/wiki/Input/output_completion_port https://en.wikipedia.org/wiki/Input/output_completion_port
- kbwt 8y agoI'm relatively sure ICOP still requires two syscalls per IO.
- blattimwind 8y agoYes, that's why I wrote almost. IOCPs do the other half (receiving results in arbitrary order in a thread pool managed by the kernel). I don't see any particular reason why a batched syscall wouldn't work for it. (Note that Windows has Read/WriteFileScatter/Gather, which is multiple reads from the same handle [as opposed to multiple reads from different handles]).
- kbwt 8y agoICOP is the opposite of polling. I really don't see the resemblance here.
- nwmcsween 8y agoI had this exact idea long ago, the syscall and interacting language would have to be different due to a mismatch, a graph language of some sort to describe dependent calls and batch and a language that makes async less painful. The non obvious benefit is a compiler could optimize across protection domains.
- convolvatron 8y agodo you think that model can usefully be extended to distributed services to reduce round trips?
- kevingadd 8y agoTalking with GPUs works a lot like this using modern APIs. You fill a command buffer with stuff and then asynchronously get results back "later", using synchronization primitives and memory mapping. There are syscalls to kick off the command buffers but those can (and often are) be implemented in userspace by the driver.
- bennofs 8y agoI have seen this technique being used to implement syscalls for processes running on intel SGX [0]. In that cases, it makes even more sense because exiting SGX, making a syscall and then reentering SGX is a lot of overhead for a single syscall, while writing to shared memory is doable without exiting the SGX environment. [0]: https://www.usenix.org/system/files/conference/osdi16/osdi16-arnautov.pdf https://www.usenix.org/system/files/conference/osdi16/osdi16...
- marcosdumay 8y agoI believe that describes the L4 messages. So, yes, it can be generalized, and there is a widely used kernel family out there that only uses it. But you will lose some performance on some of the syscalls. While an interruption is expensive, memory access is expensive too. When you have only one of the options, you will have stuff you can't optimize. The thing with IO is that the slower part of the memory access can be done by a co-processor, so you get the entire CPU available for more important work.
- naasking 8y agoModern L4 use synchronous messaging because asynchronous messages leave you vulnerable to DoS.