9 ms·
I’m new to Linux kernel programming & eBPF (just started last week) and I’m having major troubles with eBPF verifier. I honestly feel like it would be easier fo
by DethNinja 3y ago
I’m new to Linux kernel programming & eBPF (just started last week) and I’m having major troubles with eBPF verifier. I honestly feel like it would be easier for me to write a kernel module than eBPF code.
I do wonder if this is the case for many people. It seems verifier is a bit unpredictable and makes eBPF programming quite painful.
- tptacek 3y agoIt would be much easier to write a kernel module than an eBPF program. But the eBPF program is unlikely to panic your machine, and the kernel module is almost certain to.
- israrkhan 3y agoAlso In cloud environments, I would much rather trust an eBPF program, than a kernel module.
- anonymousDan 3y agoWhy in cloud environments?
- Temporary_31337 3y agoYou don’t want kernel panic affecting other users
- cassianoleal 3y agoWhy would your cloud instance panicking affect other users of the cloud provider? Or do you mean something else?
- bionsystem 3y agoAnecdotically, I had the case on VMWare 4 (that was in 2012 or 2013) that a Solaris 11 VM managed to reboot the entire ESX it was hosted on. Very weird bug where ESX passed through some interrupt or something. But in this case I think they mean on the same machine. "In production" would be more accurate than "in a cloud environment". And yeah I wouldn't load custom kernel modules in production just to do observability.
- kqr 3y agoCloud goes beyond rented VMs. Fully managed cloud services have thousands or millions of production customers on the same node. They have to be very careful about what they run as root.
- vocram 3y agoI understand your point, but millions sounds an exaggeration- I have a hard time believing a single node can handle millions of concurrent users
- cookiengineer 3y agoCan confirm, it is quite painful. The bpftools maintainers tell you to learn the bytecode format when you ask them what the errors mean, because they expect you to understand what the verifier means when it tells you "unknown scalar" on every single goddamn line of code. Something like "ebpf coding rules, what to use and what not" would be very helpful. All ebpf examples that are older than say, 3 months, already don't work anymore. Not even the official ones from the XDP tutorial project (and the libxdp maintainers because the kinda are splitting off a lot of headers into a separate xdp library as it seems). Most userspace code still relies on the 5 years old bpf-helpers.h, which meanwhile is not supported anymore because it doesn't use the __helper methods from the kernel (they also refactored the kernel in the meantime, and force you to use e.g. __u128 instead of native data types). Oh boi, did I underestimate what "bytecode vm" means when kernel developers talk about it. Also, always use llvm, and remember to build two bpf files for each endianness, and use -g for debug symbols. Otherwise you will try to find out what the bpftool errors mean for days, because of shitty mailing list answers.
- tptacek 3y agoI mean, the flip side of this is: do you really expect people on a mailing list to debug your custom XDP code for you? You get what you pay for, and you can pay people to help you with this stuff, or not.
- diarrhea 3y agoReading more about that… How does the verifier detect infinite loops anyway? Halting problem and all. It must use some rather crude heuristics, no?
- 1propionyl 3y agoIt doesn't. Flip things around and you get something tractable, though incomplete. Unsolvable problem: reject any loop that is provably infinite. Solvable problem: reject any loop that isn't provably finite. The trade off is that there will always be some loops that in fact always do terminate, but that the verifier can't prove do.
- blipvert 3y agoYou need to slightly modify your code. Rather than: while (condition) { … } Do: #define MAX 1000 for n = 0; n < MAX; n++ { if !condition break; … } Unroll all loops, don’t allow any backward jumps and limit to (say) 1m instructions.
- kqr 3y agoIncidentally, iteration limits are a good idea for production code anyway. If you don't imagine any input needing more than 50 k iterations, throw a user-friendly exception after something like 10 M iterations. Prevents much more annoying problems than it causes.