10 ms·
EBPF is turning the Linux kernel into a microkernel
- jfkebwjsbx 6y agoEverything would be a microkernel if adding some kind of VM or interpreter is enough to get that name, no? With that logic, could we argue loadable kernel modules (perhaps with proper memory separation) are a sign of a microkernel architecture?
- jacquesm 6y agoNo, a microkernel is only 'the real thing' when 'kernel modules' are simply called 'user processes'.
- tedunangst 6y agoTurn your linux into a microkernel with this one weird trick: run fuse!
- rantwasp 6y agoyes. the author of that deck is playing it pretty loose when it come to the definition of a microkernel. normally the microkernel means the minimum needed primitives to implement the OS and after that everything is build on top of that, not pluggable modules. For all intents and purposes the Linux kernel is a monolithic one and the eBPF capability make it more extensible / less of a pain to do certain things but definitely do not turn it into a microkernel.
- naasking 6y ago> normally the microkernel means the minimum needed primitives to implement the OS and after that everything is build on top of that Sure, the minimum amount of full trust code. In this case, the full trust code is the eBPF VM which enforces protection boundaries instead of the MMU as in a classic microkernel. I'm not sure a microkernel classification ought to depend on the MMU specifically, it's a general system design philosophy.
- rantwasp 6y agoit’s not just the memory protection. it’s the scheduling, IPC, etc. the eBPF vm uses the capabilities of the kernel, it is not the kernel. No kernel, no nothing. also, following your train og thought I could say that containers make this a microkernel. it would be a claim that would get you laughed out of a room.
- naasking 6y agoA kernel provides trusted runtime services for an operating system. A microkernel provides a minimal set of trusted runtime services for an operating system, and relies on some protection mechanism for isolating subsystems to avoid corrupting the trusted core. Preemptive scheduling is not necessarily part of it; depends whether your system requires "time" to be a protected resource. eBPF is a kernel service, just like processes, scheduling, IPC. If eBPF can isolate subsystems and supports safe collaboration of eBPF programs despite all running at ring 0, then the eBPF VM in the Linux kernel could qualify as a microkernel once you remove everything else. > also, following your train og thought I could say that containers make this a microkernel. If you could run all of the device drivers in containers such that they couldn't corrupt the kernel's data, then sure, you could run it as a microkernel because you wouldn't have anything left in the kernel except essential services like threading, IPC and containers.
- stefan_ 6y agoeBPF are vendor kernel modules on steroids: now instead of getting compile failures trying to build your out-of-tree module, your stuff just blows up at runtime.
- rhinoceraptor 6y agoBut, you have the huge advantage that if they crash, they don't bring down your system.
- bitcharmer 6y agoeBPF has been invaluable in my field (low-latency linux applications) and it changed a lot. If you had problems working with kernel modules before, you probably should expect struggling with writing correct code for eBPF too. It's not for everyone.
- bibabaloo 6y agoCan you share the sort of things you've been doing with it?
- bitcharmer 6y agoMost recently I used ebpf to track which other threads were stealing cpu time (and how much) from my latency sensitive cpu-pinned thread. You can do almost anything, the level of introspection into the kernel internals is amazing.
- smallstepforman 6y agoPfft, had that with 1987 Amiga 1.3, took Linux another 26 years to get there.
- teleforce 6y agoTioga editor in Xerox's Cedar already had a native structural active document capability back in 1987, but the most successful commercial Microsoft Office applications with billions of dollars budget still do not have this capability, your point exactly?
- rjsw 6y agoSun did some experiments with building a JVM into their kernel so that you could write device drivers in Java.
- smitty1e 6y agoTanenbaum lives!
- rhinoceraptor 6y agoTechnically, Linux is just a guest OS, running on top of Minix :)
- jdub 6y agoeBPF is turning Linux into a microkernel like drinking Gatorade is turning me into a Super Bowl quarterback. (I tried to localise this for a predominantly US audience.)
- badrabbit 6y agoTrue, this should have a disclaimer: "* For a very flexible definition of a microkernel"
- bregma 6y ago"given a sufficiently large value of 'micro'"
- numlock86 6y agoAre you telling me that when I use Axe deodorant my house won't be flooded by hundreds of nearby - alleged beautiful - women in their early 20s within a couple of seconds? Outrageous!
- eeZah7Ux 6y agoThe Internet is not the US.
- numlock86 6y ago> The Internet is not the US. Well, and HN is not the internet. What is your point?
- dingo_bat 6y ago> Rebooting 20,000 servers takes a very long time without risking extensive downtime. With eBPF, hot-patching servers will take a very short time to start the extensive downtime, plus the consequent reboot of 20,000 servers.
- gazsp 6y agoIt's not.
- fefe23 6y agoI don't think that word means what you think it means. Microkernel = move all the code OUT OF the kernel. These slides are about moving all the code INTO the kernel. Putting your application logic into the kernel would be more like a unikernel I guess?
- rurban 6y agoNo, it's similar to what Microsoft did. Put all the attack vectors into the kernel, because it's so much faster and we rather redo it again, we don't want to take a proven and secure existing solution. Just play the rust game and call it secure. People believe everything if you constantly repeat it.
- saagarjha 6y agoIt's disingenuous to call this the same as putting "attack vectors into the kernel", as BPF programs are sandboxed, unlike Windows kernel components. I don't know of any existing proven and secure solutions to this besides BPF, by the way.
- rurban 6y agoAs we saw with CPU's and VM's those sandboxing schemes are never secure. Eg eBPF arrays can be abused for cache attacks. The white paper and security guarantees never thought of that. The secure solution is to disable it, as well as hyperthreading. And use a secure, non-backdoored CPU.
- naasking 6y ago> Microkernel = move all the code OUT OF the kernel. Move it out of the kernel to isolate and protect. If you can isolate and protect code within the kernel's memory protection boundary, I'm not sure that that should disqualify it as a microkernel. In other words, I'm not sure that the microkernel design depends on memory protection boundaries specifically, it's a more general philosophy, akin to, "a microkernel is an operating system design which runs the minimum amount of code needed for an OS with full trust".
- layoutIfNeeded 6y agoAs always, worse is better™!
- ThePhysicist 6y agoEBPF is a super interesting technology but it’s so painfully hard to use it for application development. There are some tools based on LLVM to compile EBPF programs using C as a source language (which is much easier to reason in than the low-level code), but there is a lot of room for improving the developer workflow.
- rhinoceraptor 6y agobpftrace is getting pretty good lately, they've added support for stack arguments, so you can do things like trace golang function calls, and get arguments with a one-liner.
- MaxBarraclough 6y agoI'm not seeing how this helps solve the API stability problem faced by ordinary kernel modules. There must be some difference between this project, and a project that simply creates a more stable wrapper/subset of the APIs available to kernel modules, but it's not clear to me what it is. Also, why use JIT rather than offline verification and ahead-of-time compilation? Aside: the idea that the web delivers on the requirement of Programmability must be provided with minimal overhead is pretty laughable. Think Microsoft Teams (a chat application) would consume 600MB of memory if it were built with C++ rather than Electron? I realise not every JIT-powered technology needs to be as bloated as the web, but it seems a poor example.
- barrkel 6y agoHow can the kernel trust your offline verification? At best, what you're arguing for sounds like signed binary blobs. How do you dynamically instrument things? How do you write programs which decide, at run time, to move compute closer to the hardware?
- zozbot234 6y ago> How can the kernel trust your offline verification? You can use proof-carrying code. There is a residual "online" verification of course, but it ought to be quick and efficient.
- MaxBarraclough 6y agoYou're right, but you're way ahead of me. I'd misunderstood the emphasis of the project, and was thinking I'd be a superuser, trusted by the kernel.
- zozbot234 6y agoWell, you would still need "superuser" privileges for things like adding new capabilities to the proof verifier. Of course this might open you up to security problems if you're relying on incorrect assumptions while doing that. But then, this project also has trusted components of its own, such as the JIT. A proof verifier can be a lot simpler than a JIT.
- RMPR 6y agoI was thinking as EBPF as a way to enter in the Linux kernel development with a modern language, but I'm kinda confused by I read in the comments, it's not quite a thing?
- rhinoceraptor 6y agoeBPF is just an in-kernel VM. You can do a lot of things with it, which makes it hard to figure out what to do with it. Original BPF is in most Unix kernels, it was just a way of writing simple packet filtering programs that run in-kernel. For example, tcpdump is effectively just a frontend that emits BPF bytecode. eBPF expands the capabilities of the VM, but it still has tight restrictions on what can run: no unbounded loops, arbitrary memory access, etc. I would recommend trying out bpftrace as a first step: https://github.com/iovisor/bpftrace https://github.com/iovisor/bpftrace
- justinsaccount 6y agoThe link should be changed to https://docs.google.com/presentation/d/1AcB4x7JCWET0ysDr0gsX-EIdQSTyBtmi6OAW7bE0jm0/preview https://docs.google.com/presentation/d/1AcB4x7JCWET0ysDr0gsX... currently it links to the 2nd to last slide and not the beginning.
- dathinab 6y agoWhile the sites are interesting and Linux gets some functionalities known mostly from micro kennels it's not really turning Linux into a micro kennel at all. It just provided a new _additional_ extension mechanism which is sandboxed and much nicer to use. But to make the Linux kennel into a micro kennel eBPF would need to have the capability to replace _all_ existing kernel modules. Including file system drivers, and graphic drivers. Which is not something it's cable of sand at least currently it's only meant for new kennel functionality in to of the "core" which we have. This maybe could change at some point in the (not very close by) future. But for now it doesn't yet turn Linux into a micro kennel.
- riffraff 6y agoI appreciate this comment and agree with it but there are so many typos/autocorrectisms that it's painful to read. Dathinab, maybe do an "edit" pass? :) EDIT: fixed my own mess, thanks :)
- ghostpepper 6y ago> there are some many typos I think you meant "so many"?
- wolfgang42 6y agoMuphry's law strikes again.
- hinkley 6y agoThat damn Muphry, always messing with people.
- smileybarry 6y ago"Micro kennel" has a nice ring to it.
- yellowapple 6y ago
- aey 6y agoEBPF is ridiculously awesome. It’s safe enough to jit in ring-0! We built a rust tool chain that can output ebpf elfs :). https://github.com/solana-labs/rust-bpf-builder https://github.com/solana-labs/rust-bpf-builder
- snvzz 6y agoRunning even more code in supervisor mode != turning into a microkernel.
- perlgeek 6y agoCan device drivers be written in EBPF?
- monocasa 6y agoIt's turning it into an exokernel. Check out xok, it had three in kernel virtual machines. https://github.com/monocasa/exopc/tree/master/sys https://github.com/monocasa/exopc/tree/master/sys
- exabrial 6y agoI saw a link on HN a few months back that was going to do the same thing with WASM.
- brendangregg 6y agoI don't see anyone sharing it, but the video for this talk is here: https://www.infoq.com/presentations/facebook-google-bpf-linux-kernel/?utm_source=twitter&utm_medium=link&utm_campaign=helpcampaign https://www.infoq.com/presentations/facebook-google-bpf-linu...
- musicale 6y agoMore like an exokernel.
- seangrogg 6y agoThis is perhaps the most apt description available.
- peter_d_sherman 6y ago"A thorough introduction to eBPF" https://lwn.net/Articles/740157/ https://lwn.net/Articles/740157/ Excerpts: "While eBPF was originally used for network packet filtering, it turns out that running user-space code inside a sanity-checking virtual machine is a powerful tool for kernel developers and production engineers." [...] "The eBPF virtual machine more closely resembles contemporary processors, allowing eBPF instructions to be mapped more closely to the hardware ISA for improved performance." [...] "Originally, eBPF was only used internally by the kernel and cBPF programs were translated seamlessly under the hood. But with commit daedfb22451d in 2014, the eBPF virtual machine was exposed directly to user space." [...] "What can you do with eBPF? An eBPF program is "attached" to a designated code path in the kernel. When the code path is traversed, any attached eBPF programs are executed. Given its origin, eBPF is especially suited to writing network programs and it's possible to write programs that attach to a network socket to filter traffic, to classify traffic, and to run network classifier actions. It's even possible to modify the settings of an established network socket with an eBPF program. The XDP project, in particular, uses eBPF to do high-performance packet processing by running eBPF programs at the lowest level of the network stack, immediately after a packet is received. Another type of filtering performed by the kernel is restricting which system calls a process can use. This is done with seccomp BPF. eBPF is also useful for debugging the kernel and carrying out performance analysis; programs can be attached to tracepoints, kprobes, and perf events. Because eBPF programs can access kernel data structures, developers can write and test new debugging code without having to recompile the kernel. The implications are obvious for busy engineers debugging issues on live, running systems. It's even possible to use eBPF to debug user-space programs by using Userland Statically Defined Tracepoints." There, now you understand eBPF. It is not a Microkernel. It is an in-kernel Virtual Machine, with access to all of the kernel, whose programs can register for, receive, filter, and optionally act upon or act to moderate, kernel events. Quite the powerful tool indeed -- but not a Microkernel...