9 ms·
Exploiting null-dereferences in the Linux kernel
- deleted 4y ago[deleted]
- jeffbee 4y agoThe article is about the exploitability of the flaw but really the flaw should not exist. Printing /proc/$pid/smaps is not on any conceivable performance-critical hot path. It can stand to have bounds checks and safety. The call to print out smaps should be well-encapsulated in some non-C language.
- tedunangst 4y agoWhat does your safe language do when it accesses a null object? Does it oops?
- monocasa 4y agoIdeally it has a iterator construct built in so it views an empty linked list chain truly as an empty list without derefencing the first (null) item preemptively.
- deathanatos 4y agoMy safe language doesn't have "null", more or less. What it has is Option<T>, and I cannot turn that into a T without handling the failure case: there is literally no way to construct the code otherwise¹. One must handle the failure path. (That might be way of explicit panic/abort/oops, but it's then right there in the code: that branch will panic … and safely.) ¹(this example is using safe Rust. There's unsafe Rust too and there I can chase the null pointer all I want with that, but the parent's point is that we should be sticking to safe interfaces for stuff like this. And I'm using Rust as an example, but Option is hardly unique to Rust, heck, Rust stole the idea from its predecessors.)
- tedunangst 4y agoWhat happens when I call unwrap?
- SpaghettiCthulu 4y agoYou shouldn't be allowed to in the kernel
- deathanatos 4y agoThat's the explicit handling of the None case I mentioned in the comment: it causes an explicit, and safe, abort. By "explicit", I mean the .unwrap() call will be right there, in the method that needs to turn an Option<T> into a T, and visible to a code reviewer. In the larger context here of kernel code, it should raise the eyebrow on the reviewer: "wait, this function shouldn't abort, it needs to handle the edge cases!". (But for some userland app, aborting might be acceptable. The kernel is in a bit of a bind, since an abort — a kernel panic — means the user loses computer until they reboot, and the work along with it.) Vs. a C pointer … all uses are more or less equally suspect; any given use, you hope the code has done it's homework for ensuring they're not NULL, and if they are, the consequence is UB. (And in Rust, and in the languages Rust steals the idea of Option from, you're only using/passing Options where "None"/null/nil is a possibility. If it's not, or you've verified or handled that at some outer stack frame, then you just pass a reference to a T, which is statically guaranteed to point to a valid object¹.) ¹again, barring buggy code using unsafe Rust, in the example of Rust, or calling into C code that fails to maintain its invariants, etc. Take the example in the article, where the code does, priv->mm->mmap->vm_start while trying to generate the output for smaps_rollup. That's compilable, but buggy, C, because mmap can be null, but we failed to check for it. Vs., if mmap were an Option<T>, where T is whatever type that pointer points to. Let's say our coder attempts to write, priv->mm->mmap->vm_start (In some imaginary language, because C doesn't have Option, AFAIK.) The compiler would say, no, you can't "->vm_start", because "mmap" could be None (whatever you call the "nothing here" value/variant; I'm going to call it None, to distinguish it from the null pointer). In the case of unwrap, the coder could do something like (this is psuedo-code) (priv->mm->mmap).unwrap().vm_start It would then be obvious there is an abort there. Their reviewer would not be pleased with that, I suspect: we don't want kernel panics or oops or aborts while generating a file in /proc. And likely our imaginary coder would know this too, and when the compiler errored the first time, saying, "hey, mmap is an Option", they'd raise an eyebrow, say something like, "wait, it is? When would mmap be None?" and then proceed to properly handle that case. (E.g., by treating it as if it where the empty list.)
- wyldfire 4y agoIn order to have the safe language I believe you would need to decompose the code shown here in show_smaps_rollup(). If the null deref occurred in the unsafe portion it would likely still do an oops. If the null deref occurred in the safe portion it would likely exit safely and cause the syscall to return some errno that describes a kernel fault.
- pjmlp 4y agoNot if the safe language requires to test for nullabillity before use.
- howinteresting 4y agoYou actually work on the OpenBSD kernel?
- david2ndaccount 4y agoC could support the concept of nullable vs non-null pointers. Clang even already has this as an extension: https://clang.llvm.org/docs/AttributeReference.html#nullability-attributes https://clang.llvm.org/docs/AttributeReference.html#nullabil... There is also an associated nullability sanitizer. I use this in my own C code all the time and null pointer errors vanish if you faithfully annotate every pointer. There’s also a pragma to make non null pointers the default in a file. GCC devs would have to be convinced to add this to GCC and then nullability annotations would need to be added to the kernel. You can then do static analysis/compile error if you do an unguarded check of a nullable pointer.
- chc4 4y agoYes, in an ideal world your kernel shouldn't have any bugs. We don't live in an ideal world. Security engineering is the field of practical mitigations - given that there are, in fact, null pointer dereferences in the kernel, mmap_min_addr and adding count limit to kernel oops provides defense in depth to help prevent them from being exploitable.
- chlorion 4y agoI don't think they are arguing that there should be no bugs. >Security engineering is the field of practical mitigations Somehow I've managed to use bounds checking in anything I create, and I'm not even an engineer, just a hobbyist!
- deleted 4y ago[deleted]
- roguebantha 4y agoThankfully this isolated flaw was quite easy to fix. And yes this code isn't likely to be on any hot paths, and code can always stand to have bounds/sanity checks (and it always should). But unfortunately encapsulating all non-hot-paths in Linux kernel that might have these sorts of bugs in a memory-safe language is at best a very long term goal and at worst a pipe-dream. The real goal of the blog post was not to push for any sort of rewrite, but rather to note how even the simplest and most innocuous of bugs can lead to security-relevant primitives. And also to make sure kernel developers and bug fixers have strategies like this in mind when they evaluate other bugs in the future. TLDR: However honorable the end-goal is, this blog post is not the ammo you need to push for a big rewrite of various kernel<->userland interfaces into memory safe languages.
- nix0n 4y ago> Printing /proc/$pid/smaps is not on any conceivable performance-critical hot path. I disagree, for profiling memory usage it's useful to get memory map data multiple times per second.
- jeffbee 4y agoIf you think smaps is performance-critical that raises of the question of its ridiculous textual format. Clearly, it would be vastly more efficient to pass the information as a protobuf or whatever. Believe me, as the person who had to refactor the smaps-reading library on cost/efficiency grounds at Google, this issue is nearer and dearer to me than to probably anyone else.
- roguebantha 4y agoThe irony isn't lost on me that the one conceivable hotpath is this exploit technique...
- azakai 4y agoIIUC steps 5-7 in the exploit cause around 2^32 oopses. I don't know much about the Linux kernel - could it perhaps have a limit on the number of oopses before it halts the entire system? The article explains why it is important to not do that in general, as an oops allows debugging and recovery etc. But 2^32 of them seems suspicious.
- roguebantha 4y agoYes, there is now an oops limit, specifically because of this technique - see the conclusion paragraph.
- azakai 4y agoAh, thanks! I should finish reading the entire article before commenting, sorry...
- spacepinball 4y agoYou'd think that be the rule; reflect, THEN discuss.
- roguebantha 4y agoTo the contrary, I appreciate azakai's enthusiasm for seeking solutions, and also credit the fact that they came up with exactly the same solution we used to mitigate the issue.
- high_byte 4y ago8 days to exploit :) pretty neat. and 2 years on servers? still worth a shot. I bet it can be much faster in certain scenarios.
- roguebantha 4y agoIt can be a lot slower too - but it's most dependent on how the kernel handles console logging. Serial consoles slow it down dramatically, which is why it's so much quicker on your typical GUI enabled setup.
- nineteen999 4y agoI'm really curious why there aren't more enterprise-grade, production ready kernels at this point. Isn't Rust nearing maturity? Doesn't the community have tonnes of enterprise ready C code that could be used as a reference (ie. Linux, BSD) of "what not to do"? I'm not trying to start an argument here, I think the world knows that C/C++ make it way too easy to shoot ourselves in the foot by now. I know that writing operating systems is hard and takes a long time, i've written my own prototype single and multitasking operating systems for x86_32, 68k, Z80, 6502 etc. I'm aware that Rust support has been added to recent Linux kernels, for the limited use case of writing secure device drivers. None of these things are news to me, so please don't regurgitate these points. But given the great body of reference that is available, the enthusiasm in the Rust community for the promise of more secure operating system kernels, I'm genuinely suprised that things aren't further along. Yes I'm aware of Redox, but it seems more aimed at desktop use, and last time I tried it didn't even boot. Projects in C/C++ seem to be making much faster progress eg. SerenityOS than the Rust community. What is holding Rust back in this area? This is a genuine question, not intending to inflame the discussion. I'm spending some time learning Rust as I can afford, but am not opinionated one way or the other yet. Where are all the Linux replacements that I would have imagined to be up and running by now given Rust's maturity? What am I missing here? Happy to be genuinely informed. I kind of expected there to be a bunch of projects in flight by now, ala bazaar style, with the Rust community starting to conglomerate around the strongest contenders and move them forward at a rapid pace.
- newaccount2021 4y ago[dead]
- sph 4y agoYou're asking for enterprise-grade, production ready kernels, and immediately after say Rust is nearing maturity. There you go. Writing a kernel is like building a cathedral. To have a production ready kernel, set aside 15 years, or the equivalent in dollars, at least a couple billion. This is why stuff like Serenity, which is an outstanding achievement, is not much more than a toy. We will have a Rust based kernel that is memory safe. Not this decade though. By that time Linux will have replaced more and more subsystem with Rust already. Remember, enterprise grade means boring and stable. It is nonsense to want a new, unproved kernel to provide that level of safety out of a language that has reached 1.0 not very long ago. Personally, I think Linux is plenty good enough, but we have seen the best UNIX can offer. It's time to move on and try something new.
- mappu 4y agoSee also this recent (Nov 2022) LWN article regarding the new oops limit: https://lwn.net/Articles/914878/ https://lwn.net/Articles/914878/
- Vecr 4y agoWhen they say "map the zero page" in the article, it appears they are talking about the page with index zero, not the page with all zeros in it. Does anyone know if this is correct?
- bennettnate5 4y agoYes, you have that right. By mapping memory to the page at index zero (the page a 'null' pointer would point to), an attacker can leverage a null dereference to achieve much more interesting/dangerous outcomes than just a crash.
- pjmlp 4y agoYep, that is common OS speak.
- seba_dos1 4y agoI don't know any context where "the zero page" means "the page with all zeros in it". It's a rather commonly used term.
- Dylan16807 4y agoGoogle linux zero page. The results about the page at address zero and the results about the page of all zeros are about equally common. The page full of zeros is used constantly when processes allocate memory, and is set to copy-on-write.
- roguebantha 4y agoYes I had considered it rather apparent I was referring to the page at the virtual address zero, but I've definitely referred to the CoW page used for private anonymous mappings as the "zero page" too, so I understand the confusion. I probably should have made it more clear! Sorry about that.