5 ms·
So from my understanding, that's a kind of "secure" (I'd like to know more about the security model tbh) module that runs with kernel privilege with no scheduli
by Iv 7y ago
So from my understanding, that's a kind of "secure" (I'd like to know more about the security model tbh) module that runs with kernel privilege with no scheduling (so it runs until completion). These are supposed to be short and I am assuming, can't call libs and can't allocate memory (outside a predefined stack I would guess?)
Aren't they very similar to interrupts? What is the difference there? The kernel API?
- GrayShade 7y agoMy understanding is that they're program fragments that can be attached to kernel code and are generally used for debugging and observability. It's using a virtual machine that's designed to sandbox these, and they're limited in power so that a buggy filter can't hang the kernel. This is different form hardware interrupts because they don't run in response to hardware events, and they don't have side effects. There is a similarity in that you want the filters to run quickly, though.
- wisty 7y agoIIR I think it's a VM(?) that has certain limits, and because of those limits it's OK to run it in kernel-space. IIRC it's not Turning complete, and has a fixed run time. So Linux can just say "OK run this now" and not worry about scheduling. It's like putting a green thread in the kernel but to do this safely you need very strict restrictions (finite memory it can access, finite number of steps, etc). You can thus get speed-ups because there's no API, memory management, even scheduling.
- Iv 7y agoGot it, thanks! I did not realize it runs in a VM, ok, that indeed makes sense to call it a different kind of application.
- maxdamantus 7y agoI think it's more to do with avoiding overheads typically associated with system calls (presumably involving some interrupt and disabling/enabling/changing paging behaviour). Here's an example of a syscall-heavy command on my system: $ time dd if=/dev/zero bs=1 count=10M of=/dev/null 10485760+0 records in 10485760+0 records out 10485760 bytes (10 MB, 10 MiB) copied, 7.09089 s, 1.5 MB/s real 0m7.092s user 0m2.123s sys 0m4.968s 3 million system calls per second seems quite slow on a 3.2 GHz CPU when all it should really be doing is dereferencing a couple of pointers until it finds some functions that simply write a zero byte to a buffer (the "/dev/zero" descriptor handler) and ignore bytes from a buffer (the "/dev/null" descriptor handler). If you have a safe bytecode format for representing operations that are performed in a loop, the kernel can just perform those operations without having to switch back and forth to userspace.
- megous 7y agoRe-try with mitigations=off
- joosters 7y agoHow much of that time is really spent in the system call interface? You've got 4.968s of system time there (i.e. broadly the time spent in kernel code) and 2.123s of user time. Given that the user-space program is effectively a tight loop around read() and write() calls, we can assume that almost all of those 2 seconds are spent going through the syscall plumbing. Now, there's going to also be some of the kernel-side time spent in the syscall plumbing too, but there's also a lot of I/O, buffer and filesystem layer code executing there. All of which will be in use with a BPF program too. So it's unclear how much of the effective time can be shaved off.
- maxdamantus 7y agoThere shouldn't really be any significant filesystem code involved, since once `dd` has opened the files, it should have handlers for those devices more-or-less directly in its descriptor table. Once you have a descriptor to a pipe or device, there shouldn't be any filesystem-level checking in the middle of your reads/writes; all you're doing is filling/emptying buffers. And given that I can write a program that makes 132 million calls per second to the glibc `putchar` function (which also buffers), I'm pretty sure there's a lot of time that can be shaved off as we start to replace the system call mechanism with plain function calls.
- barrkel 7y agoHave you forgotten about Meltdown, Spectre, and all the other cache attacks?
- de_watcher 7y agoIt's one of the two phases. We're back to the other one, wait for a couple of months.
- maxdamantus 7y ago