19 ms·
Kernel bugs hide for 2 years on average. Some hide for 20
- sedatk 9mo agoFirefox bugs stay in the open for that long.
- steveklabnik 9mo agoOne of my favorite Firefox bugs was some I don’t quite remember the details of, but went something like this: “There’s a crash while using this config file.” Something more complex than that, but ultimately a crash of some kind. Years later, like 20 years later, the bug was closed. You see, they re-wrote the config parser in Rust, and now this is fixed.” That’s cool but it’s not the part I remember. The part I always think about is, imagine responding to the bug right after it was opened with “sorry, we need to go off and write our own programming language before this bug is fixed. Don’t worry, we’ll be back, it’s just gonna take some time.” Nobody would believe you. But yet, it’s what happened.
- nurettin 9mo agoTo be fair, any rewrite could have fixed it, didn't have to wait for Rust.
- Xss3 9mo agoBut that take ruins all the intrigue of their comment... But youre spot on. They fabricated a story.
- Yossarrian22 9mo agoNo, Graydon Hoare took one look at the config code, went “fuck this” and decided to create a new language instead.
- steveklabnik 9mo agoI didn’t say otherwise. Rust is not the point here.
- mmooss 9mo agoAll software has long-lived bugs. None are bug-free, at any point in their existance, so it's almost inevitable. Have you seen Windows' bug tracker? The anti-Firefox mob really is striving to take shots at it. The point of the article isn't a criticism of Linux, but an analysis that leads to more productive code review.
- snvzz 9mo agoMillions of lines of code, all running in supervisor mode. One bug is all it takes to compromise the entire system. The monolithic UNIX kernel was a good design in the 60s; Today, we should know better[0][1]. 0. https://sel4.systems/ https://sel4.systems/ 1. https://genode.org/ https://genode.org/
- windowssuperfi 9mo agoYeah cause windows is amazing Or maybe macos? Ignore their freebsd parts of course.
- DowsingSpoon 9mo agoYes. As far as kernels go, NT was pretty damn good. So is Mach, by the way, if you can afford the microkernel performance overhead.
- dundarious 9mo agoIf you include all the drivers too (which surely makes the comparison more accurate), is that still the case?
- nobodyandproud 9mo agoWindows NT 3.x was a true microkernel. Microsoft ruined it but the design was quite good and the driver question was irrelevant, until they sidestepped HAL. The Linux kernel was and is a monstrosity.
- WalterGR 9mo agoWhat do you meant by them sidestepping the HAL?
- avadodin 9mo agoI think the biggest one is that the whole GDI library was moved into the Kernel in 3.5x because the performance was terrible at the time. I don't think they ever intended to keep all drivers strictly userland, though. Just the service side.
- esseph 9mo agoImagine if no one outside a select circle ever got to examine the code.
- immibis 9mo agoEverything is open source if you're skilled with Ghidra. We call AI models "open source" if you can download the binary and not the source. Why not programs?
- deleted 9mo ago[deleted]
- KK7NIL 9mo ago> We call AI models "open source" if you can download the binary and not the source. Who's "we"? There's been quite a lot of pushback on this naming scheme from the OSS community, with many preferring the term "open weights".
- serf 9mo ago>We call AI models "open source" if you can download the binary and not the source. Why not programs? the weights of a model aren't equivalent to the binary output of source code, no matter how you try to stretch the metaphor. >why not because we aren't beholden to change all definitions and concepts because some guy at some corp said so.
- immibis 9mo agoUnless that corp is OSI, right?
- heavyset_go 9mo agoBinaries and AI models can be inscrutable. They're meant to be interpreted by machines. We want human readable, comprehensible, reproducible and maintainable sources at minimum when we say open source.
- dspillett 9mo ago
- eulgro 9mo agoFrom the stats we see that most bugs effectively come from the limitations of the language. Impressive results on the model, I'm surprised they improved it with very simple heuristics. Hopefully this tool will be made available to the kernel developers and integrated to the workflow.
- NewsaHackO 9mo agoIt may be just my system, but the times look like hyperlinks but aren't for some reason. It is especially disappointing that the commit hashes don't link to the actual commit in the kernel repo.
- Telaneo 9mo agoThey're <strong> tags with color:#79635c on hover in the CSS. A really weird style choice for sure, but semantically they aren't meant to be links at all.
- NewsaHackO 9mo agoI know, I am saying they should be links, as it is what one would expect from an article like this.
- Fiveplus 9mo agoBefore the "rewrite it in Rust" comments take over the thread: It is worth noting that the class of bugs described here (logic errors in highly concurrent state machines, incorrect hardware assumptions) wouldn't necessarily be caught by the borrow checker. Rust is fantastic for memory safety, but it will not stop you from misunderstanding the spec of a network card or writing a race condition in unsafe logic that interacts with DMA. That said, if we eliminated the 70% of bugs that are memory safety issues, the SNR ratio for finding these deep logic bugs would improve dramatically. We spend so much time tracing segfaults that we miss the subtle corruption bugs.
- anon-3988 9mo ago> It is worth noting that the class of bugs described here (logic errors in highly concurrent state machines, incorrect hardware assumptions) wouldn't necessarily be caught by the borrow checker. Rust is fantastic for memory safety, but it will not stop you from misunderstanding the spec of a network card or writing a race condition in unsafe logic that interacts with DMA. Rust is not just about memory safety. It also have algebraic data types, RAII, among other things, which will greatly help in catching this kind of silly logic bugs.
- JuniperMesos 9mo agoYeah, Rust gives you much better tools to write highly concurrent state machines than C does, and most of those tools are in the type system and not the borrow checker per se. This is exactly what the Typestate pattern (https://docs.rust-embedded.org/book/static-guarantees/typestate-programming.html https://docs.rust-embedded.org/book/static-guarantees/typest...) is good at modeling.
- aw1621107 9mo ago> It is worth noting that the class of bugs described here (logic errors in highly concurrent state machines, incorrect hardware assumptions) While the bugs you describe are indeed things that aren't directly addressed by Rust's borrow checker, I think the article covers more ground than your comment implies. For example, a significant portion (most?) of the article is simply analyzing the gathered data, like grouping bugs by subsystem: Subsystem Bug Count Avg Lifetime drivers/can 446 4.2 years networking/sctp 279 4.0 years networking/ipv4 1,661 3.6 years usb 2,505 3.5 years tty 1,033 3.5 years netfilter 1,181 2.9 years networking 6,079 2.9 years memory 2,459 1.8 years gpu 5,212 1.4 years bpf 959 1.1 years Or by type: Bug Type Count Avg Lifetime Median race-condition 1,188 5.1 years 2.6 years integer-overflow 298 3.9 years 2.2 years use-after-free 2,963 3.2 years 1.4 years memory-leak 2,846 3.1 years 1.4 years buffer-overflow 399 3.1 years 1.5 years refcount 2,209 2.8 years 1.3 years null-deref 4,931 2.2 years 0.7 years deadlock 1,683 2.2 years 0.8 years And the section describing common patterns for long-lived bugs (10+ years) lists the following: > 1. Reference counting errors > 2. Missing NULL checks after dereference > 3. Integer overflow in size calculations > 4. Race conditions in state machines All of which cover more ground than listed in your comment. Furthermore, the 19-year-old bug case study is a refcounting error not related to highly concurrent state machines or hardware assumptions.
- silver_sun 9mo agoTheir section on "Dataset limitations" says that the study "Only captures bugs with Fixes: tags (~28% of fix commits)." Just worth noting that it is a significant extrapolation from only "28%" of fix commits to assume that the average is 2 years.
- tremon 9mo agoWhy? A sample size of 28% is positively huge compared to what most statistical studies have to work with. The accuracy of an extrapolation is mostly determined by underlying sampling bias, not the amount of data. If you have any basis to suggest that capturing "only bugs with fixes tags" creates a skewed sample, that would be grounds to distrust the extrapolation, but simply claiming "it's only 28%" does not make it worth noting.
- dpc_01234 9mo agoMight be obviously, but there is definitely a lot of biases in the data here. It's unavoidable. E.g. many bugs will not be detected, but they will be removed when the code is rewritten. So code that is refactored more often will have lower age of fixed bugs. Components/subsystems that are heavily used will detect bugs faster. Some subsystems by their very nature can tolerate bugs more, while some by necessity will need to be more correct (like bpf).
- a3w 9mo agoThe kernel this speaks of is probably linux. Does windows have a similar round time?
- pixl97 9mo agoI mean, yea. Here is a device driver bug that was around 11 years. https://www.bitdefender.com/en-us/blog/hotforsecurity/google-reveals-windows-kernel-bug-exploited-in-the-wild-thats-been-around-since-2009 https://www.bitdefender.com/en-us/blog/hotforsecurity/google...
- deleted 9mo ago[deleted]
- maximgeorge 9mo ago[dead]
- ValdikSS 9mo agogrsecurity project has fixed many security bugs but did not contribute back, as they're profiting from selling the patchset. It's not uncommon for the bugs they found to be rediscovered 6-7 years later. https://xcancel.com/spendergrsec https://xcancel.com/spendergrsec
- __bjoernd 9mo ago> as they're profiting from selling the patchset Profiting from selling their patchset is not the whole story, though. grsec was public and free for a long time and there were many effects at play preventing the kernel from adopting it.
- woliveirajr 9mo agoBut the patchset should use the same license as the original code, shouldn't?
- ValdikSS 9mo agoIt is: https://grsecurity.net/faq https://grsecurity.net/faq
- staticassertion 9mo agoThis implies (or states, hard to say) that they don't upstream specifically in order to profit. That is nonsense. 1. Tons of bugs are reported upstream by grsecurity historically. 2. Tons of critical security mitigations in the kernel were outright invented by that team. ASLR, SMAP, SMEP, NX, etc. 3. They were completely FOSS until very recently. 4. They have always maintained that they are entirely willing to upstream patches but that it's a lot of work and would require funding. Upstream has always been extremely hostile towards attempts to take small pieces of Grsecurity and upstream them.
- zkmon 9mo agoA bug is a piece of code that doesn't agree with requirements or architecture. The misalignment can not be attributed to code alone.
- jackfranklyn 9mo ago[flagged]
- MarleTangible 9mo agoOne of the iOS 26 Core Audio bug (CVE-2025-31200) is about synchronizing two different arrays with each other and the assumption mistakes that were made trusting dimensional information which could be coming from the user. https://youtu.be/nTO3TRBW00E https://youtu.be/nTO3TRBW00E
- pixl97 9mo ago>hide for years in application code Yea, it's pretty common. We had a customer years ago that was having a rare and random application crash under load. Never could figure out where it was from. Quite some time later a batch load interface was added to the app and with the rate things were input with it the crash could be triggered reliably. It's something else that's added/changed in the application that eventually makes the bug stand out.
- sureglymop 9mo agoOnly tangentially related but maybe someone here can help me. I have a server which has many peripherals and multiple GPUs. Now, I can use vfio and vfio-pcio to memory map and access their registers in user space. My question is, how could I start with kernel driver development? And I specifically mean the dev setup. Would it be a good idea to use vfio with or without a vm to write and test drivers? How to best debug, reload and test changing some code of an existing driver?
- MORPHOICES 9mo ago[dead]
- fsflover 9mo ago> To what extent do you trust "well-tested" code? I don't, which is why I use Qubes OS providing security through compartmentalization.
- hun3 9mo agoThen the question becomes: to what extent do you trust Xen and Qubes RPC?
- fsflover 9mo agoI do have to somewhat trust Xen, but Qubes' isolation relies on hardware virtualization (VT-d), which statistically has much less security issues than Xen itself. Most Xen advisories do not affect Qubes: https://www.qubes-os.org/security/xsa/ https://www.qubes-os.org/security/xsa/
- saagarjha 9mo ago> Undefined behavior-related bugs are permanently hidden. No they are often found and fixed.
- burnt-resistor 9mo agoSpeaking of nasty kernel bugs although on another platform, there's a nasty one in either Microsoft's Win 11 nwifi.sys handling of deadlock conditions or Qualcomm's QCNCM865 FastConnect 7800 WCN785x driver that panics because of a watchdog failure in nwifi!MP6SendNBLInternal+0x4b guarded by a deadlocked ndis!NdisAcquireRWLockRead+0x8b. It "BSODs" the system rather than doing something sane like dropping a packet or retransmitting. Am I the only unreasonable maniac who wants a very long-term stable, seL4-like capability-based, ubiquitous, formally-verified μkernel that rarely/never crashes completely* because drivers are just partially-elevated programs sprinkled with transaction guards and rollback code for critical multiple resource access coordination patterns? (I miss hacking on MINIX 2.) * And never need to reboot or interrupt server/user desktop activities because the core μkernel basically never changes since it's tiny and proven correct.
- giamma 9mo agoIs the intention of the author to use the number of years bugs stay "hidden" as a metric of the quality of the kernel codebase or of the performance of the maintainers? I am asking because at some point the articles says "We're getting faster". IMHO a fact that a bug hides for years can also be indication that such bug had low severity/low priority and therefore that the overall quality is very good. Unless the time represents how long it takes to reproduce and resolve a known bug, but in such case I would not say that "bug hides" in the kernel.
- staticassertion 9mo ago> IMHO a fact that a bug hides for years can also be indication that such bug had low severity/low priority and therefore that the overall quality is very good. It doesn't seem to indicate that. It indicates the bug just isn't in tested code or isn't reached often. It could still be a very severe bug. The issue with longer lived bugs is that someone could have been leveraging it for longer.
- galangalalgol 9mo agoWorst case is that it doesn't even cause correctness issues in normal use, only when misused in a way that is unlikely to happen unintentionally.
- staticassertion 9mo agoI guess because I work in security the "unintentionally" doesn't matter much to me.
- SAI_Peregrinus 9mo agoBut it matters for detection time, because there's a lot more "normal" use of any given piece of code than intentional attempts to break it. If a bug can't be triggered unintentionally it'll never get detected through normal use, which can lead to it staying hidden for longer.
- gjfr 9mo agoInteresting! We did a similar analysis on Content Security Policy bugs in Chrome and Firefox some time ago, where the average bug-to-report time was around 3 years and 1 year, respectively. https://www.usenix.org/conference/usenixsecurity23/presentation/franken https://www.usenix.org/conference/usenixsecurity23/presentat... Our bug dataset was way smaller, though, as we had to pinpoint all bug introductions unfortunately. It's nice to see the Linux project uses proper "Fixes: " tags.
- staticassertion 9mo ago> It's nice to see the Linux project uses proper "Fixes: " tags. Sort of. They often don't.
- YouAreWRONGtoo 9mo ago[dead]
- deleted 9mo ago[deleted]
- jmyeet 9mo agoThe lesson here is that people have an unrealistic view of how complex it is to write correct and safe multithreaded code on multi-core, multi-thread, assymmetric core, out-of-order processors. This is no shade to kernel developers. Rather, I direct this at people who seem to you can just create a thread pool in C++ and solve all your concurrency problems. One criticism of Rust (and, no, I'm not saying "rewrite it in Rust", to be clear) is that the borrow checker can be hard to use whereas many C++ engineers (in particular, for some reason) seem to argue that it's easier to write in C++. I have two things to say about that: 1. It's not easier in C++. Nothing is. C++ simply allows you to make mistakes without telling you. GEtting things correct in C++ is just as difficult as any other language if not more so due to the language complexity; and 2. The Rust borrow checker isn't hard or difficult to use. What you're doing is hard and difficult to do correctly. This is I favor cooperative multitasking and using battle-tested concurrency abstractions whenever possible. For example the cooperative async-await of Hack and the model of a single thread responding to a request then discarding everything in PHP/Hack is virtually ideal (IMHO) for serving Web traffic. I remember reading about Google's work on various C++ tooling including valgrind and that they exposed concurrency bugs in their own code that had lain dormant for up to a decade. That's Google with thousands of engineers and some very talented engineers at that.
- wordisside 9mo ago[dead]
- aw1621107 9mo ago> The implementations of sort in Rust are filled with unsafe. Strictly speaking, the mere presence of `unsafe` says nothing on its own about whether "it" is easier in C++. Not only does `unsafe` on its own say nothing about the "difficulty" of the code it contains, but that is just one factor of one side of a comparison - very much insufficient for a complete conclusion. Furthermore, "just" writing a sorting algorithm is pretty straightforwards both in Rust and C++; it's the more interesting properties that tend to make for equally interesting implementations, and one would need to procure Rust and C++ implementations with equivalent properties, preferably from the same author(s), for a proper comparison. Past research has shown that Rust's current sorting algorithms have different properties than C++ implementations from the time (e.g., the "X safety" results in [0]), so if nothing substantial has changed since then there's going to be some work to do for a proper comparison. Edit: forgot to add the reference [0]: https://github.com/Voultapher/sort-research-rs/blob/main/writeup/sort_safety/text.md https://github.com/Voultapher/sort-research-rs/blob/main/wri...
- deleted 9mo ago[deleted]
- ryukoposting 9mo agoThis is fascinating stuff, especially the per-subsystem data. I've worked with CAN in several different professional and amateur settings, I'm not surprised to see it near the bottom of this list. That's not a dig against the kernel or the folks who work on it... more of a heavy sigh about the state of the industries that use CAN. On a related note, I'm seeing a correlation between "level of hoopla" and a "level of attention/maintenance." While it's hard to distinguish that correlation from "level of use," the fact that CAN is so far down the list suggests to me that hoopla matters; it's everywhere but nobody talks about it. If a kernel bug takes down someone's datacenter, boy are we gonna hear about it. But if a kernel bug makes a DeviceNet widget freak out in a factory somewhere? Probably not going to make the front page of HN, let alone CNN.
- pixl97 9mo agoThere is a general rule on bugs is that the more devices they are on, the more apt they are to trigger. A CAN with 10,000 machines total and relatively fixed applications is either going to trigger the bug right off the bat and then work around it, or trigger the bug so rarely it won't be recognized as a kernel issue. General purpose systems running millions and millions of units with different workloads are an evolutionary breeding ground for finding bugs and exploits.
- eab- 9mo agoI'd find this article a bit more compelling if it was used to find current introduced bugs, instead of just using a holdout set
- michaelcampbell 9mo agoThank goodness for reader mode. The transparent background where the text is with the wiggly line background is... challenging.
- GaryBluto 9mo agoWhat's with the odd scribbles in the background?
- kmavm 9mo agoIt's an easter egg on the website that usually goes unnoticed. It's our first time on the front page of HN, so it's a little overutilized right now. Capital-C clears it.
- redleader55 9mo agoI don't think the problem is the kernel. Kernel bugs stay hidden because no one runs recent Kernels. My Pixel 8 runs kernel a stable minor from 6.1, which was released more than 4 years ago. Yes, fixes get backported to it, but the new features in 6.2->6.19 stay unused on that hardware. All the major distros suffer from the same problem, most people are not running them in production Most hyperscalers are running old kernel versions on which they do backports. If you go Linux conferences you hear folks from big companies mentioning 4.xx, 3.xx kernels, in 2025.
- calebm 9mo agoIt's interesting to consider that the same phenomenon may also hold true for humanity's psychological software.
- Adrian-ChatLocl 9mo agoStill probably a lot better than Windows.
- blueboo 9mo ago“In a sufficiently complex system, malfunction or even total non-function may go undetected for long periods, if ever” John Gall, The Systems Bible