9 ms·
-fbounds-safety: Enforcing bounds safety for C
- clarabennett26 8mo ago[dead]
- adrianN 8mo agoThere is GWPAsan that has lower overhead than asan but still is not super popular.
- hoyhoy 8mo agoASAN/LSAN is amazing. It absolutely monkey-hammers performance though.
- lelanthran 8mo ago> ASAN/LSAN is amazing. It absolutely monkey-hammers performance though. It's not so bad; until the sanitisers arrived all we had was valgrind :-/ The sanitisers are about 10x to 50x faster than valgrind.
- vlovich123 8mo agoBecause it can only catch a subset of issues, it’s not guaranteed to catch issues (probabilistic), even issues it “could” catch may not be caught due to temporal distance of the free and a subsequent use, and requires the use of a different allocator that supports it. It’s also unclear to me how it know whether a given free is for a sampled or unsampled region - I suspect it must capture all free/realloc to accomplish that but it does imply all of these are sampled. It’s nowhere near the same as robust bounds checking.
- hoyhoy 8mo agoI looked at trying to implement -fbounds-safety and -Wunsafe-buffer on a reasonably large codebase (4,000 C and C++ files), and it's basically impossible. You have to instrument every single file. It can be done in stages though. Just turn the flag on one-by-one for each file. The xnu kernel is _mostly_ instrumented with -fbounds-safety.
- jimmaswell 8mo agoThis sounds like the kind of low-thought pattern-based repetitive task where you could tell an LLM to do it and almost certainly expect a fully correct result (and for it to find some bugs along the way), especially if there's some test coverage for it to verify itself against. If you're skeptical, you could tell it to do it on some files you've already converted by hand and compare the results. This kind of thing was a slam dunk for an LLM even a year or two ago.
- safercplusplus 8mo agoPlug: In theory you could auto-convert to a memory-safe subset of C++ as a build step. Auto-converted code would have some run-time overhead, but you can mark any performance-sensitive parts of the code to be exempt from conversion. And you get lifetime and type safety too. For full coverage, performance-sensitive parts of the code can be manually converted to the safe subset to minimize overhead. (Interfaces in extern C blocks remain unconverted by default to maintain ABI compatibility.) [1]: https://duneroadrunner.github.io/scpp_articles/PoC_autotranslation_of_wget https://duneroadrunner.github.io/scpp_articles/PoC_autotrans...
- musicale 8mo agoI want an OS distro where all C code is compiled this way. OpenBSD maybe? or a fork of CheriBSD? macOS clang has supported -fbounds-safety for a while, but I"m not sure how extensively it is used.
- 1over137 8mo ago>I want an OS distro where all C code is compiled this way. You first have to modify "all C code". It's not just a set and forget compiler flag.
- musicale 8mo agoIndeed. I still want it.
- wyldfire 8mo agoYou need to annotate your program with indications of what variable tracks the size of the allocation. So, sure, but first work on the packages in the distro. Note that corresponding checks for C++ library containers can be enabled without modifying the source. Google measured some very small overhead (< 0.5% IIRC) so they turned it on in production. But I'd expect an OS distro to be mostly C. [1] https://libcxx.llvm.org/Hardening.html https://libcxx.llvm.org/Hardening.html
- bombcar 8mo agoGet gentoo, add this to CFLAGS and start fixing everything that breaks. Become a hero.
- pezgrande 8mo agodoes any distro uses clang? I thought all linux kernels were compiled using gcc.
- zmodem 8mo agoNot a Linux distro, but FreeBSD uses Clang. And Android uses Clang for its Linux kernel. -fbounds-safety is not yet available in upstream Clang though: > NOTE: This is a design document and the feature is not available for users yet.
- nananana9 8mo agotemplate <typename T> struct Slice { T* data = nullptr; size_t size = nullptr; T& operator[](size_t index) { if (index >= size) crash_the_program(); return data[index]; } }; If you're considering this extension, just use C++ and 5 lines of standard, portable, no-weird-annotations code instead.
- zmodem 8mo agoThe extension is for hardening legacy C code without breaking ABI.
- baq 8mo agoand if you write directly in assembly you don't even need a C++ compiler
- nananana9 8mo agoThat's an objectively correct statement, but I don't see how it makes sense as a response to my comment, as I'm advocating to use the more advanced feature-rich tool over the compiler-specific-hacks one.
- zephen 8mo ago> I don't see how it makes sense as a response to my comment Your comment started out with "just." As if there are never any compelling reasons to want to make existing C code better. But instead of taking that as an opportunity to reflect on when various tools might be appropriate, > as I'm advocating to use the more advanced feature-rich tool over the compiler-specific-hacks one. You've simply doubled down.
- yjftsjthsd-h 8mo agoIf you're advocating switching languages, then there's no reason to stop at C++. It's more common to propose just converting the universe to Rust, but assembly also enjoys the possibility of being fairly easy to drop in on an existing C project.
- 8mo ago
- nimbus-hn-test 8mo ago[dead]
- taminka 8mo agothis is amazing, counter to what most ppl think, majority of memory bugs are from out of bounds access, not stuff like forgetting to free a pointer or some such
- Retr0id 8mo agoI think UAFs are more common in mature software
- q3k 8mo agoOr type confusion bugs, or any other stuff that stems from complex logic having complex bugs. Boundary checking for array indexing is table stakes.
- michh 8mo agotable stakes, but people still mess up on it constantly. The "yeah, but that's only a problem if you're an idiot" approach to this kind of thing hasn't served us very well so it's good to see something actually being done. Trains shouldn't collide if the driver is correctly observing the signals, that's table stakes too. But rather than exclusively focussing on improving track to reduce derailments we also install train protection systems that automatically intervene when the driver does miss a signal. Cause that happens a lot more than a derailment. Even though "pay attention, see red signal? stop!" is conceptually super easy.
- q3k 8mo agoI'm not saying it's not important, it is. I just don't believe that '[the] majority of memory bugs are from out of bounds access'. That was maybe true 20 years ago, when an unbounded strcpy to an unprotected return pointer on the stack was super common and exploiting this kind of vulnerabilities what most vulndev was. This brings C one tiny step closer to the state of the art, which is commendable, but I don't believe codebases which start using this will reduce their published vulnerability count significantly. Making use of this requires effort and diligence, and I believe most codebases that can expend such effort already have a pretty good security track record.
- worldsavior 8mo agoVery cool. I always wondered why there isn't something like this in GCC/LLVM, it would obviously solve uncountable of security issues.
- ndiddy 8mo agoHas any progress been made on this? I remember seeing this proposal 3 or 4 years ago but it looks like it still hasn't been implemented. It's a shame because it seems like a useful feature. It looks like Microsoft has something similar (https://learn.microsoft.com/en-us/cpp/code-quality/understanding-sal?view=msvc-170 https://learn.microsoft.com/en-us/cpp/code-quality/understan...) but it would be nice to have something that worked on other platforms.
- Someone 8mo agohttps://discourse.llvm.org/t/the-preview-of-fbounds-safety-is-now-accessible-to-the-community/84221 https://discourse.llvm.org/t/the-preview-of-fbounds-safety-i...: “-fbounds-safety is a language extension to enforce a strong bounds safety guarantee for C. Here is our original RFC. We are thrilled to announce that the preview implementation of -fbounds-safety is publicly available at this fork of llvm-project. Please note that we are still actively working on incrementally open-sourcing this feature in the llvm.org/llvm-project . To date, we have landed only a small subset of our implementation, and the feature is not yet available for use there. However, the preview does contain the working feature. Here is a quick instruction on how to adopt it.” “This fork” is https://github.com/swiftlang/llvm-project/tree/stable/20240723 https://github.com/swiftlang/llvm-project/tree/stable/202407..., Apple’s fork of LLVM. That branch is from a year ago. I don’t know whether there’s a newer publicly available version. There is a GSoC 2026 opportunity on upstreaming this into mainline LLVM (https://discourse.llvm.org/t/gsoc-2026-participating-in-upstreaming-fbounds-safety/89649 https://discourse.llvm.org/t/gsoc-2026-participating-in-upst...)
- groos 8mo agoMicrosoft's SAL annotations are meant to inform the static analyzer how the parameters are meant to be used so any violations of the contract can be diagnosed at compile time. The LLVM proposal is different in that it is checked at run time and will stop your program before it makes an out of bounds access. Static analyzers can obviously use the information in the type to help diagnose a subset of such problems at compile time.
- mrpippy 8mo agoApple is shipping code built with this, and is supporting it for developers to use (see https://developer.apple.com/documentation/xcode/enabling-enhanced-security-for-your-app#Adopt-bounds-checking-in-C https://developer.apple.com/documentation/xcode/enabling-enh...)
- cranberryturkey 8mo agoThe real question is adoption friction. The annotation requirement means this won't just slot into existing codebases — someone has to go through and mark up every buffer relationship. Google turning on libcxx hardening in production with <0.5% overhead is compelling precisely because it required zero source changes. The incremental path matters more than the theoretical coverage. I'd love to see benchmarks on a real project — how many annotations per KLOC, and what % of OOB bugs it actually catches in practice vs. what ASAN already finds in CI.
- favorited 8mo agoThe WebKit folks have apparently been very successful with the annotations approach[0]. It's a shame that a few of the loudest folks in WG21 have decided that C++ already has the exact right number of viral annotations already, and that the language couldn't possibly survive this approach being standardized. [0]https://www.youtube.com/watch?v=RLw13wLM5Ko https://www.youtube.com/watch?v=RLw13wLM5Ko
- manbash 8mo agoExciting! It doesn't imply that we should now sprinkle the new annotations everywhere. We still should keep working with proper iterators and robust data structures, and those would need to add such annotations.
- hoyhoy 8mo agoXcode (AppleClang) has had -fbounds-safety for a while now. What is the delay getting this into merged into LLVM?
- jcalvinowens 8mo ago> As local variables are typically hidden from the ABI, this approach has a marginal impact on it. I'm skeptical this is workable... it's pretty common in systems code to take the address of a local variable and pass it somewhere. Many event libraries implement waiting for an event that way: push a pointer to a futex on the stack to a global list, and block on it. They address it explicitly later: > Although simply modifying types of a local variable doesn’t normally impact the ABI, taking the address of such a modified type could create a pointer type that has an ABI mismatch That breaks a lot of stuff. The explicit annotations seem like they could have real value for libraries, especially since they can be ifdef'd away. But the general stack variable thing is going to break too much real world code.
- rbanffy 8mo agoI would imagine variables that are passed to functions would be considered ABI-visible. If the compiler is smart enough, it can keep the pointer wide when it’s passed to a function that’s also being compiled and act accordingly on the other side, but that worries me because this new meaning of “pointer” is propagating to parts of the code that might not necessarily agree with it.
- menaerus 8mo agoI don't understand this example: you're taking an address of local-scope stack object, storing it into a global list, and then use this address elsewhere in the code, possibly at different time-point, to manipulate with the object? I am obviously missing something because this cannot work unless this object lives on the stack of main().
- jandrese 8mo agoYep, it's a straight up error in C to return the address of a local variable from a function outside of main. Valgrind will flag this as use of an uninitialized value. The problem is that as long as it's something where the calling function checks it immediately after the function exits and never looks again (something like an error code or choosing a code path based on the result) they often get away with it, especially in single threaded code. I'm running into this at this very moment as I'm trying to make my application run cleanly, but some of the libraries are chock full of this pattern. One big offender is the Unix port of Microsoft's ODBC library, at least the Postgre integration piece. I also blame the Unix standard library for almost having this pattern but not quite. Functions that return some kind of internal state that the programmer is told not to touch. Later they had to add a bunch of _r variants that were thread safe. The standard library functions don't actually have this flaw due to how they define their variables, but from the outside it looks like they do. It makes beginning programmers think that is how the functions should work and write their code in a similar manner.
- tandr 8mo agoNiklaus Wirth died in 2024, and yet I hope he is having a major I-told-you-so moment about people blaming Pascal's bounds checking to be unneeded and making things slow.
- nmz 8mo agoTo this day, FPC uses less ram than any C compiler, A good thing in today's increasingly ramless world and they've managed this with way less developers working on it than its C compiler equivalent, I can't even imagine what it would look like if they had the same amount of people working on it. C optimization tricks are hacks, the fact godbolt exists is proof that C is not meant to be optimizable at all, it is brute force witchcraft. At a certain point though, something's gotta give, the compiler can do guesswork, but it should do no more, if you have to add more metadata then so be it it's certainly less tedious than putting pragmas and _____ everywhere, some C code just looks like the writings of an insane person.
- inkyoto 8mo ago> […] C optimization tricks are hacks, the fact godbolt exists is proof that C is not meant to be optimizable at all, it is brute force witchcraft. > At a certain point though, something's gotta give, the compiler can do guesswork, but it should do no more, if you have to add more metadata then so be it it's certainly less tedious than putting pragmas and _____ everywhere, some C code just looks like the writings of an insane person. There is not even a single correct or factual statement in cited strings of words. C optimisation is not «hacks» or «witchcraft»; it is built on decades of academic work and formal program analysis: optimisers use data-flow analysis over lattices and fixed points (abstract interpretation) and disciplined intermediate representations such as SSA, and there is academic work on proving that these transformations preserve semantics. Modern C is also deliberately designed to permit optimisation under the as-if rule, with UB (undefined behaviour) and aliasing rules providing semantic latitude for aggressive transformations. The flip side is non-negotiable: compilers can't «guess» facts they can't prove, and many of the most valuable optimisations require guarantees about aliasing, alignment, loop independence, value ranges, and absence of UB that are often not derivable from arbitrary pointer-heavy C, especially under separate compilation. That is why constructs such as «restrict», attributes and pragmas exist: they are not insanity, they are explicit semantic promises or cost-model steering that supply information the compiler otherwise must conservatively assume away. «metadata instead» is the same trade-off in a different wrapper, unless you either trust it (changing the contract) or verify it (reintroducing the hard analysis problem). Godbolt exists because these optimisations are systematic and comparable, not because optimisation is impossible. Also, directives are not new, C-specific embarrassment: ALGOL-68 had «pragmats» (the direct ancestor of today’s «pragma» terminology), and PL/I had longstanding in-source compiler control directives, so this mechanism is decades older than and predates modern C tooling.
- matheusmoreira 8mo agoAmazing, this is a life saving feature for C developers. Apparently it's not complete yet? I will apply this to my code once the feature is included on LLVM and GCC. Would be nice if the annotations could also be applied to structure fields. struct bytes { size_t count; unsigned char * __counted_by(count) pointer; }; void work_with(struct bytes);
- zokier 8mo agocounted_by for struct fields actually is actually the part that afaik works today: https://embeddedor.com/blog/2024/06/18/how-to-use-the-new-counted_by-attribute-in-c-and-linux/ https://embeddedor.com/blog/2024/06/18/how-to-use-the-new-co...
- matheusmoreira 8mo agoThat's amazing. Thanks for that reference. If it's good enough for the kernel, then it's good enough for me to start using in my own projects. It's really cool that the kernel is using this. The compiler must be generating simple bounds checking code with traps instead of crazy stuff involving magical C standard library functions. Perfect for freestanding nostdlib projects.
- uecker 8mo agoClang has this and upcoming GCC will also have this: https://godbolt.org/z/KETrPEnT1 https://godbolt.org/z/KETrPEnT1
- matheusmoreira 8mo agoThis is awesome!!
- kazinator 8mo ago> To tackle this issue, the model incorporates the concept of a “wide pointer” (a.k.a. fat pointer) – a larger pointer that carries bounds information alongside the pointer value. Bounds checking with fat pointers existed as a set of patches for GCC in the early 2000's. (C front end only). https://sourceforge.net/projects/boundschecking/ https://sourceforge.net/projects/boundschecking/