6 ms·
> I know next to nothing about kernel programming, but I'm not sure here what Linus' objection to the comment he is responding to here is. You should read the
by arinlen 4y ago
> I know next to nothing about kernel programming, but I'm not sure here what Linus' objection to the comment he is responding to here is.
You should read the email thread, as Linhas explains in clear terms.
Take for instance Linus's insightful followup post:
https://lkml.org/lkml/2022/9/19/1250 https://lkml.org/lkml/2022/9/19/1250
- ChrisSD 4y agoWhat is better: continuing to "limp along" in some unknown corrupted state (aka undefined behaviour) or in a well defined (albeit invalid) state?
- throw827474737 4y agoHad the same topic often on MCUs: limp along to hopefully get the error out somehow, otherwise it won't be noticed if not with JTAG debugger attached (default in field). So I can understand where Linus comes from.
- gmueckl 4y agoYes. You could still hard reset after the error is reported if you wanted to. And if system availability matters, a hardware watchdog would handle the case where the error handling doesn't finish.
- mlindner 4y agoLimping along is what the salesman and the business people want as failures look bad. Engineers should want the immediate stop, because that's safer, especially in safety critical situations.
- wtallis 4y agoThe kernel is not the whole system. The kernel needs to offer the "limping along" option so that the other parts of the system can implement whatever graceful failure method is appropriate for that system. There's no one size fits all solution for the kernel to pick.
- warinukraine 4y agoYou sound like you code websites or something. Real engineers, like say the people who code the machines that fly in mars, don't want "oops that's unexpected, ruin the entire mission because that's safer". Same for the Linux kernel.
- niscocity35 4y agoWhat are you talking about? Should planes stop flying when they encounter an error? Safety critical systems will try to recover to a working state as much as possible. It is designed with redundancy that if one path fails, it can use path 2 or path 3 towards a safe usable state.
- Someone1234 4y agoThis question is answered in Linus' emails fully and better than I'm going to do. But to restate briefly, the answer varies wildly between kernel and user programs, because a user program failing hard on corrupt state is still able to report that failure/bug, whereas a kernel panic is a difficult to report problem (and breaks a bunch of automated reporting tooling). So in answer: Read the discussion.
- ChrisSD 4y agoYou seem to have misunderstood me. The distinction I'm making is not between kernel panic or undefined behaviour. The distinction is between undefined behaviour and defined behaviour. That defined behaviour can be anything, even including "limping on" somehow.
- yencabulator 4y agoWhat is better for a desktop user: 1) needing to reload a wifi driver to reinitialize hardware (with a tiny probability of memory corruption) OR choosing to reboot as soon as convenient (with a tiny probability of corrupting the latest saved files) 2) to lose unsaved files for sure and not even know what caused the crash
- Jweb_Guru 4y agoThe latter, because the "tiny probability of memory corruption" can easily become a CVE.
- P5fRxh5kUvp2th 4y agoWe have a term for this. FUD
- Jweb_Guru 4y agoLinux has numerous CVEs, and a large percentage stem from memory corruption. That's not FUD, I'm afraid.
- scoutt 4y agoIt's FUD. And not only that. The fear of constantly being attacked by an external entity is also paranoic.
- Jweb_Guru 4y agoUnfortunately, whether you personally care about this sort of thing isn't good enough anymore. Owned Linux boxes on IoT devices are now being marshaled into massive botnets used to perform denial of service attacks, while other vulnerabilities are exploited to enable ransomware. You having negligent security on your own unpatched box because you don't personally feel like it's a good tradeoff has many negative external consequences. Fortunately, the decision isn't actually up to you (and having fewer vulnerabilities won't influence you negatively anyway, so I'm not sure why you're so angry about it).