6 ms·
Typically a system call only leads to a context switch if the call will block, or if the caller has used up its timeslice; otherwise, it's just a couple of rela
by dhess 18y ago
Typically a system call only leads to a context switch if the call will block, or if the caller has used up its timeslice; otherwise, it's just a couple of relatively cheap protection domain crossings.
- tptacek 18y agoA read from the disk is always going to sleep the process.
- dhess 18y agoYes, regardless of the mechanism that causes the disk read. Given a particular access pattern, whether the disk gets read or not is independent of the choice to use seek+read or mmap. System calls that don't block don't cause a context switch (1). Therefore, the occurrence of context switches due to disk reads is also independent of the choice to use seek+read or mmap, so trying to avoid context switches isn't a reason to use one or the other. (1) modulo the timeslice expiring during the system call, which could actually help the system call case since there's no need to take an interrupt from user space when the timeslice expires due to a timer: the system call can be predicted since it's in the code stream, but the interrupt can't be. Anyway, that's all in the noise.)
- tptacek 18y agoI agree with you, but you just refuted the argument you made one comment back. =)
- dhess 18y agoNo, I didn't :) All I did in my original comment was simply to point out to ajross that system calls don't always cause context switches. It's a common misconception, not unique to ajross, of course, that a protection domain crossing requires a context switch, so I was making what I hoped was a helpful observation. I didn't say anything at all about disk reads in that original comment, nor about the merits of mmap vs. seek+read! I'm glad we agree on my second comment, in any case. :)
- ajross 18y agoI think we're quibbling about terminology. I'm not sure of the difference between a "protection domain crossing" (not a common term, AFAIK) and a "context switch" (pervasive jargon). You seem to be indicating that the latter always involves a separate process, which of course is more expensive (on most architectures) due to cache issues. I'm using the term to mean a switch between two execution domains.
- dhess 18y agoSorry, you're right, "protection domain crossing" probably isn't common jargon. It's a holdover from my time as a processor architect on IA-64. It means privilege escalation. Anyway, in any common parlance of the term "context switch" I've ever seen, the amount of state switched to take a system call is much less (edit: oops, originally said 'greater' ;) than the amount of state switched in a context switch. At the very least, a context switch should mean the save and restore of user-level register state, which isn't necessary for a system call. You certainly wouldn't want to do that just to make a system call on a typical RISC architecture with 30-something registers! After all, from the perspective of the caller, it's just a special procedure call. The compiler and system call entry point can even use the same argument-passing convention. I wrote at least part of (maybe most, can't remember anymore :) the "recommended" context switch handler for IA-64. It was significantly more expensive than the system call handler. IA-64 dedicated quite a bit of hardware to making system calls cheap. If I recall correctly, besides a few privileged registers, the only thing we changed was the location of the register stack engine's backing store. That's just a single user-level register. (You don't want to spill registers with potentially privileged state into user space.) Anyway, this is way off-topic now and I don't think anybody's interested other than the three of us, so that's the last I'll say about it.
- tptacek 18y agoWell, here's where this gets fun: most of the kernels we work with are structured like event loops, and disk I/O is asynchronous. When you issue a read vnop in a Unix kernel, your process will probably yield back to the scheduler loop, incurring the out switch and the in switch. This mmap vs. read argument is as old as the hills. I picked a side a long time ago; keep it simple.
- ajross 18y agoA system call is a pair of context switches: swap the registers, swap the stack, change the TLB mappings, on some architectures flush the cache. Even the fastest system calls (Linux on amd64) take a thousand clock cycles or more.
- tptacek 18y agoTechnically, a system call (on i386) is only changing CS, EIP, and ESP. You don't have to change CR3 or flush the TLB. "Swapping stacks" isn't expensive.
- ajross 18y agoIt's a thousand clock cycles more than a function call (seriously, I'm not making this up: time it sometime). I guess whether that is "expensive" or not is context dependent. For some applications (e.g. database servers) that make lots of random access reads from large files, using a mapping can be overwhelmingly faster. I guess I'm stunned that this turns out to have been such a controversial notion...
- tptacek 18y agoCan you be specific about which part of the system call you're talking about? Because you said "swapping the stack" (which doesn't take a thousand clock cycles) and "flushing the TLB" (which doesn't happen in OSs where the kernel occupies a permanent part of the VM address space). I'm sorry, this isn't controversial. All I said was, "mmap isn't faster than read in typical use cases". But then we all got really specific talking about ESP and CS and CR3 and now we have something to go back and forth on, which is kind of fun, and I might learn something new. Didn't mean to snipe at you.
- ajross 18y agoTo be clear: I didn't say "flush the TLB". Nonetheless unless you happen to be lucky and have those kernel mapping in there, those TLB entries need to be faulted in. That's expensive. Likewise the new stack isn't in L1 cache and needs to be read in from main memory. Likewise the kernel code to execute the handler needs to be read into the instruction cache, etc... Given that main memory reads are pushing 100 cycles on a modern box, all those things add up. A context switch (my usage: meaning a bounce to a non-local, non-current execution environment) is really expensive.