9 ms·
How we found and fixed an eBPF Linux kernel vulnerability
- 285094302 2y ago[flagged]
- deleted 2y ago[deleted]
- katzinsky 2y agoThe one time I tried to use eBPF it wasn't expressive enough for what I needed. Does the limited flexibility it provides really justify the added kernel space complexity? I can understand it for packet filtering but some of the other stuff it's used for like sandboxing just isn't convincing.
- knorker 2y agoThere are other technologies for this, such as DTrace. The kernel's choice isn't eBPF or nothing, it's eBPF or something else like it. You may not use it much, but some people use it all day. I think FAANG engineers have said that they run tens (hundreds?) of these things on all servers, all the time. And that's excluding one-offs. And FAANG has full time kernel coders on staff, so they're also funding this complexity that they use. But also yes, I've solved problems by using eBPF. Problems that are basically unsolvable by non-kernel-gurus without eBPF. I rarely need it. But when I need it, there's nothing else that does the trick. In some cases, even for kernel gurus, it's a choice between eBPF or maintaining a custom kernel patch forever.
- katzinsky 2y agoI'm not sure "Google engineers use it" is a very good counter argument. They have a very high tolerance for complexity and like most large corporations what actually gets built and used tends to be driven more by internal politics than technical merit.
- eggnet 2y agoGoogle would maintain a kernel patch or upstream a patch if that was the right choice for a given problem.
- katzinsky 2y agoThat's really begging the question. I don't believe they would as they have consistently over engineered solutions in the past.
- DaiPlusPlus 2y ago> Google would maintain a kernel patch I look forward to seeing that patch on Google Graveyard in a couple years' time.
- knorker 2y agoI don't mean it as a counter argument, or I don't think the way you mean it, at least. You may not use it at your smaller scale. But there are millions of machines out there that do use it, and the alternative for the same functionality is much worse. I bet you never use SCTP sockets either. eBPF is used much more than SCTP. And its users "fund" its development, so it's not a burden to those who don't use it. But are you sure your systems don't use it? Run "bpftool prog" to see. Whatever you see there someone thought was better than the alternative.
- lynxmachine 2y ago> I've solved problems by using eBPF. Problems that are basically unsolvable by non-kernel-gurus without eBPF. I rarely need it. Would you mind giving some examples? I recently started learning about ebpf's from Liz Rice's book and is curious about what makes ebpf the correct choice in a particular scenario.
- znpy 2y ago> There are other technologies for this, such as DTrace. The kernel's choice isn't eBPF or nothing, it's eBPF or something else like it. To add on this point: I successfully used SystemTap a few years ago to debug an issue i was having. Before going further: keep in mind that my point of view (at the time) was the one of somebody working as a devops engineer, debugging some annoyances with containers (managed by Kubernetes) going OOM. I'm no kernel developer and I have a basic-good understanding of the C language based on first-years university course and geekyness/nerdyness. So in this context I'm a glorified hobbyist. Learning SystemTap is easier in my opinion. I followed a tutorial by RedHat to get the hang of the manual parts but after that I remember being fairly easy: 1. Try to reproduce the issue you're having (fairly easy for me) 2. Skim the source code of the linux about the part that you think might be relevant (for me it was the oom killer) 3. Add probes in there, see if they fire when you reproduce the issue 4. Look back at the source code of the kernel and see what chain of data structures and fields you can follow to reach the piece of information you need 5. Improve your probes 6. If successful, you're done 7. Goto 4 I think it took like one or two days between following the tutorial and getting a working probe. It was a pleasant couple of days.
- fch42 2y agoDTrace and eBPF are "not so different" in the sense that dtrace programs / hooks are also a form of low-level code / instruction set that the kernel (dtrace driver) validates at load. It's an "internal" artifact of dtrace though, https://github.com/illumos/illumos-gate/blob/master/usr/src/lib/libdtrace/common/dt_cc.c https://github.com/illumos/illumos-gate/blob/master/usr/src/... and to my knowledge, nothing like a clang/gcc "dtrace target" exists to translate more-or-less arbitrary higher-level language "to low-level dtrace". The additional flexibility eBPF gets from this is amazing really. While dtrace is a more-targeted (and for its intended usecases, in some situations still superior to eBPF) but also less-general tool. (citrus vs. stone fruit ...)
- cryptonector 2y agoDTrace's bytecode machine is also very very limited. eBPF's is much less limited. Limiting the scope of what a probe can do is very important.
- bcantrill 2y agoYes, thank you. Long before eBPF existed, we spent a ton of time on the safety of DTrace[0][1] -- there's a bunch of subtlety to it. The proof is in the pudding, however: thanks to our strict adherence to the safety constraint, we have absolute confidence in using DTrace in production. [0] https://bcantrill.dtrace.org/2005/07/19/dtrace-safety/ https://bcantrill.dtrace.org/2005/07/19/dtrace-safety/ [1] https://www.usenix.org/legacy/publications/library/proceedings/usenix04/tech/general/full_papers/cantrill/cantrill.pdf https://www.usenix.org/legacy/publications/library/proceedin..., §3.3
- saagarjha 2y agoI’m curious which part of these tenets would feel would have prevented the bug demonstrated, besides “oh we tried harder”? I don’t see any of those that seem unique to DTrace other than limiting where probes can be placed.
- 2y ago
- deleted 2y ago[deleted]
- ssahoo 2y agoWouldn't even the classic loadable kernel mode driver be a better choice than a patch and eBpf? I know they are unsafe but people who deal with it, know the power comes with responsibility.
- tptacek 2y agoNo? SREs roll eBPF programs on the fly just in the process of debugging problems; if you tried to do that with an LKM, you'd almost certainly blow up your system. People who write Linux kernel code routinely crash their systems in the process of development.
- techwiz137 2y agoIn my country we have a saying. "Porcupine in the pants". Sounds like for all the good it can do, it isn't written safely and carefully.
- deskr 2y agoWith experience you'll realise that despite things being done safely and carefully, mistakes can and do pop up.
- bugtodiffer 2y agoTrue. There are some nasty bugs in some very well written code.
- tptacek 2y agoA reminder that on the platforms eBPF is most commonly used, verifier bugs don't matter much, because unprivileged code isn't allowed to load eBPF programs to begin with. Bugs like this are thus root -> ring0 vulnerabilities. That's not nothing, but for serverside work it's usually worth the tradeoff, especially because eBPF's track record for kernel LPEs is actually pretty strong compared to the kernel as a whole. In the setting eBPF is used today, most of the value of the verifier is that it's hard to accidentally crash your kernel with a bad eBPF program. That is comically untrue about an ordinary LKM.
- chc4 2y agoThe PoC uses eBPF maps as their out-of-bounds pointer, but it sounds like it would also be exploitable via non-extended BPF programs loadable via seccomp since it's just improper scalar value range tracking, which doesn't require any privileges on most platforms. And, of course, root -> ring0 is less of a problem with unprivileged user namespaces where you can make yourself "root", as we've seen in every eBPF bug PoC since distros started turning that on (and have since turned it off again, mostly)
- tptacek 2y agoI just want to say that this is a hell of a nerd snipe.
- chc4 2y agoLMAO Ok that's fair. check_seccomp_filter actually has a more restrictive list than just "BPF with no backwards jumps", and in particular doesn't allow BPF_IND in the BPF_LDX, so you can't read out of bounds because you can't use a dynamic displacement...but BPF_STX is allowed, so you can probably write out of bounds? BPF_W is the seccomp_data address and the control flow diagram they show to compute incorrect scalar ranges doesn't require any backwards jumps...
- tptacek 2y agoI feel like I just played the Uno Reverse card on the nerd snipe.
- mrbluecoat 2y ago> “Uno no es ninguno” (One is none) I believe that translates to "One is not none" https://bughunters.google.com/blog/6303226026131456/a-deep-dive-into-cve-2023-2163-how-we-found-and-fixed-an-ebpf-linux-kernel-vulnerability#-uno-no-es-ninguno-one-is-none- https://bughunters.google.com/blog/6303226026131456/a-deep-d...
- DanielVZ 2y agoThats the direct translation but for some reason in spanish our double negations are usually just negations.
- kmarc 2y agoIt doesn't; It translates to "One is none" This is the infamous double negation many foreign speakers (including me) struggles with. https://spanish.stackexchange.com/questions/26777/how-does-double-negation-using-no-hay-and-ning%C3%BAn-ninguna-work https://spanish.stackexchange.com/questions/26777/how-does-d...
- samatman 2y agoPerhaps we should translate this as "one ain't nothin'".
- TacticalCoder 2y ago> “Uno no es ninguno” (One is none) Literally "One not is none", aka "One is not none".
- jolmg 2y agoIn Spanish, it's common for double negatives to not actually be double negatives. For example, if you wanted to say "there's nothing here", you'd say "no hay nada aquí", which word-for-word means "there's not nothing here". Checking out the Royal Spanish Academy, here's what they say about it: https://www.rae.es/espanol-al-dia/doble-negacion-no-vino-nadie-no-hice-nada-no-tengo-ninguna https://www.rae.es/espanol-al-dia/doble-negacion-no-vino-nad... > The so-called "double negation" is due to the obligatory negative agreement that must be established in Spanish, and other Romance languages, in certain circumstances (see New Grammar, § 48.3d), which results in the joint presence in the statement of the adverb no and other elements that also have a negative meaning. > The concurrence of these two "negations" does not annul the negative meaning of the statement.