5 ms·
The one time I tried to use eBPF it wasn't expressive enough for what I needed. Does the limited flexibility it provides really justify the added kernel space
by katzinsky 2y ago
The one time I tried to use eBPF it wasn't expressive enough for what I needed.
Does the limited flexibility it provides really justify the added kernel space complexity? I can understand it for packet filtering but some of the other stuff it's used for like sandboxing just isn't convincing.
- knorker 2y agoThere are other technologies for this, such as DTrace. The kernel's choice isn't eBPF or nothing, it's eBPF or something else like it. You may not use it much, but some people use it all day. I think FAANG engineers have said that they run tens (hundreds?) of these things on all servers, all the time. And that's excluding one-offs. And FAANG has full time kernel coders on staff, so they're also funding this complexity that they use. But also yes, I've solved problems by using eBPF. Problems that are basically unsolvable by non-kernel-gurus without eBPF. I rarely need it. But when I need it, there's nothing else that does the trick. In some cases, even for kernel gurus, it's a choice between eBPF or maintaining a custom kernel patch forever.
- katzinsky 2y agoI'm not sure "Google engineers use it" is a very good counter argument. They have a very high tolerance for complexity and like most large corporations what actually gets built and used tends to be driven more by internal politics than technical merit.
- eggnet 2y agoGoogle would maintain a kernel patch or upstream a patch if that was the right choice for a given problem.
- katzinsky 2y agoThat's really begging the question. I don't believe they would as they have consistently over engineered solutions in the past.
- DaiPlusPlus 2y ago> Google would maintain a kernel patch I look forward to seeing that patch on Google Graveyard in a couple years' time.
- knorker 2y agoI don't mean it as a counter argument, or I don't think the way you mean it, at least. You may not use it at your smaller scale. But there are millions of machines out there that do use it, and the alternative for the same functionality is much worse. I bet you never use SCTP sockets either. eBPF is used much more than SCTP. And its users "fund" its development, so it's not a burden to those who don't use it. But are you sure your systems don't use it? Run "bpftool prog" to see. Whatever you see there someone thought was better than the alternative.
- lynxmachine 2y ago> I've solved problems by using eBPF. Problems that are basically unsolvable by non-kernel-gurus without eBPF. I rarely need it. Would you mind giving some examples? I recently started learning about ebpf's from Liz Rice's book and is curious about what makes ebpf the correct choice in a particular scenario.
- znpy 2y ago> There are other technologies for this, such as DTrace. The kernel's choice isn't eBPF or nothing, it's eBPF or something else like it. To add on this point: I successfully used SystemTap a few years ago to debug an issue i was having. Before going further: keep in mind that my point of view (at the time) was the one of somebody working as a devops engineer, debugging some annoyances with containers (managed by Kubernetes) going OOM. I'm no kernel developer and I have a basic-good understanding of the C language based on first-years university course and geekyness/nerdyness. So in this context I'm a glorified hobbyist. Learning SystemTap is easier in my opinion. I followed a tutorial by RedHat to get the hang of the manual parts but after that I remember being fairly easy: 1. Try to reproduce the issue you're having (fairly easy for me) 2. Skim the source code of the linux about the part that you think might be relevant (for me it was the oom killer) 3. Add probes in there, see if they fire when you reproduce the issue 4. Look back at the source code of the kernel and see what chain of data structures and fields you can follow to reach the piece of information you need 5. Improve your probes 6. If successful, you're done 7. Goto 4 I think it took like one or two days between following the tutorial and getting a working probe. It was a pleasant couple of days.
- fch42 2y agoDTrace and eBPF are "not so different" in the sense that dtrace programs / hooks are also a form of low-level code / instruction set that the kernel (dtrace driver) validates at load. It's an "internal" artifact of dtrace though, https://github.com/illumos/illumos-gate/blob/master/usr/src/lib/libdtrace/common/dt_cc.c https://github.com/illumos/illumos-gate/blob/master/usr/src/... and to my knowledge, nothing like a clang/gcc "dtrace target" exists to translate more-or-less arbitrary higher-level language "to low-level dtrace". The additional flexibility eBPF gets from this is amazing really. While dtrace is a more-targeted (and for its intended usecases, in some situations still superior to eBPF) but also less-general tool. (citrus vs. stone fruit ...)
- cryptonector 2y agoDTrace's bytecode machine is also very very limited. eBPF's is much less limited. Limiting the scope of what a probe can do is very important.
- bcantrill 2y agoYes, thank you. Long before eBPF existed, we spent a ton of time on the safety of DTrace[0][1] -- there's a bunch of subtlety to it. The proof is in the pudding, however: thanks to our strict adherence to the safety constraint, we have absolute confidence in using DTrace in production. [0] https://bcantrill.dtrace.org/2005/07/19/dtrace-safety/ https://bcantrill.dtrace.org/2005/07/19/dtrace-safety/ [1] https://www.usenix.org/legacy/publications/library/proceedings/usenix04/tech/general/full_papers/cantrill/cantrill.pdf https://www.usenix.org/legacy/publications/library/proceedin..., §3.3
- saagarjha 2y agoI’m curious which part of these tenets would feel would have prevented the bug demonstrated, besides “oh we tried harder”? I don’t see any of those that seem unique to DTrace other than limiting where probes can be placed.
- 2y ago
- deleted 2y ago[deleted]
- ssahoo 2y agoWouldn't even the classic loadable kernel mode driver be a better choice than a patch and eBpf? I know they are unsafe but people who deal with it, know the power comes with responsibility.
- tptacek 2y agoNo? SREs roll eBPF programs on the fly just in the process of debugging problems; if you tried to do that with an LKM, you'd almost certainly blow up your system. People who write Linux kernel code routinely crash their systems in the process of development.