14 ms·
Rust for Filesystems
- gritzko 2y agoFrom the minutes I conclude that Rust-in-the-kernel looks like an additional complexity tax. I mean, if you write an OS from scratch, you can use the full power of your language. Plastering it to the side of an already vast codebase creates additional issues, as we see here.
- thesnide 2y agoWhile I agree there are benefits to rust, I tend to think all reason cannot fight hype. The tax will be seen as a necessity to embrace future and progress. I'm wondering why do not restrict ourselves to a safe subset instead of jumping into a huge bandwagon of unknown bugs and tradeoffs
- Yoric 2y agoSafe subset of what?
- pjc50 2y agoThere is no "safe subset" of C. MISRA is fairly close, but all sorts of things that you might need, like integer arithmetic, have potential UB in C. (The best current effort is https://sel4.systems/ https://sel4.systems/ , which is written in C but has a large proof of safety attached. The language design question is basically: should the proof be part of the language?)
- regularfry 2y agoGiven that undefined behaviour just means "undefined by the standard", do you get usefully closer to being able to identify a safe subset with the (MISRA/alternative, specific compiler, specific architecture) triple?
- pjc50 2y agoNo, undefined behavior does not mean "not defined by the standard", it means those places where the standard says "undefined behavior". And then the long and complicated war over "the compiler may assume that UB does not happen and then optimize on that basis". You might be able to tighten it up in some specific cases, and those battles are being fought elsewhere, but there's stuff like lock lifetimes which you cannot do without substantial extra annotations inside or outside the language.
- regularfry 2y agoSorry, yes - poor wording on my part.
- PhilipRoman 2y agoI found frama-c to be pretty good, including all the integer quirks
- germandiago 2y agoI like Linus view on evolution. Evolution will tell what the most sensible choices are over time. It is like "the market" in some way. Let everyone make their bets, wait, see, analyze, research. That's it.
- mnau 2y ago> additional complexity tax Yes, but that should be offsetted by easier driver development. See the blog about Rust GPU driver for asahii linux, done in one month. EDIT: Google "tales of the m1 gpu" (author has a very negative opinions about hacker news, read if you like by clicking the link https://asahilinux.org/2022/11/tales-of-the-m1-gpu/ https://asahilinux.org/2022/11/tales-of-the-m1-gpu/) Is it universal? We'll see in coming years.
- eru 2y agoAlas, that link just gets you a rant about politics, if you click on it directly. Copy-and-pasting works.
- olivermuty 2y ago"Rant about politics", haha. Or as other people like to call it: "A real concern described in an apt manner". I have observed these inflammatory sub-graphs of comments myself and have thought to myself that this must be a huge growing grounds for unmoderated and unwanted behaviour because it more or less becomes invisible once flagged enough.
- aoieu 2y ago[flagged]
- vollbrecht 2y agoOne can argue that any additional code is introducing complexity, not only writing Rust. Does that mean we should just stop innovating and go into an indefinite state of maintenance, since we are already so vast? A tax in one place may not be a net negative, if it's used like in the real world to offset other problems. And just saying it will not offset any problems because of a single discussion, that does not have a definite conclusion, comes of as a short argument.
- another2another 2y ago>Does that mean we should just stop innovating and go into an indefinite state of maintenance If you mean that not using Rust (or maybe some other languages e.g. Zig or Ada?) means that there can be no innovation in the Linux kernel, I would have to disagree since there's been plenty of progress in plain old c (see for instance io_uring), not to mention the fact that the c language itself could change to make developer ergonomics better - since that seems to be the nub of the problem. It also raises the question of what happens in the future when Rust is no longer the language du jour - how do we know it's going to last the course? And now there's 2 different codebases, potentially maintained by 2 different diminishing sets of active maintainers.
- vollbrecht 2y ago> If you mean that not using Rust (or maybe some other languages e.g. Zig or Ada?) means that there can be no innovation in the Linux kernel, I would have to disagree since there's been plenty of progress in plain old c. No i didn't mean that. If i understand OP correctly here, he argued that it is a tax to use rust, a tax is always bad, and thous should be avoided. We obviously can't now the future. We also can't now how future maintainers look like, and if there will be a bigger abundance of people understanding kernel level C or kernel level Rust or both. I also don't think that any one developer can claim to fully get every part of the Linux Kernel. So if one person want's to work on a particular subsection they need to make themself familiar with it, independent of the language used. And then we are back at the argument, is the additional tax bad, or what does it bring to the table.
- deleted 2y ago
- Already__Taken 2y agoreads a lot like letting perfect be the enemy of good.
- aoieu 2y ago[flagged]
- tucnak 2y agoThey simply wish the actual kernel developers just surrendered & weren't in the way of new Rust code. Maybe if these guys wrote C for a living some 10 years or so, became maintainers in their own right, and THEN brought these ideas forward—they would have the chance. But you can't come to the other people's projects, and seriously expect them to just nod ahead to everything you have to say, surrender your concerns, and act like they owe you sommat. I'm glad this conversation is happening, though.
- paavohtl 2y agoTo my knowledge most if not all of these people driving Rust adoption in Linux are seasoned Linux contributors and/or maintainers. They are not outsiders "coming to other people's projects".
- sprash 2y ago[flagged]
- lelanthran 2y ago> To my knowledge most if not all of these people driving Rust adoption in Linux are seasoned Linux contributors and/or maintainers. They are not outsiders "coming to other people's projects". Not in this case. It's the Rust evangelist who is the newcomer. FTA: > Almeida said that he is not trying to keep the C API static; his goal is to get the filesystem developers to explain the semantics of the API so that they can be encoded into Rust. The responses from the kernel team members to the proposal seem reasoned and mature to me: > In addition, when the C code changes, the Rust code needs to follow along, but who is going to do that work? > As the C code evolves, which will happen more quickly than with the Rust code, at least initially, there will be a need to keep the two APIs in sync. > The object lifecycles are being encoded into the Rust API, but there is no equivalent of that in C; if someone changes the lifecycle of the object on one side, the other will have bugs. > Encoding a single lifecycle understanding into the API means that its functions will not work for some filesystems. > Part of the problem, Ted Ts'o said, is that there is an effort to get "everyone to switch over to the religion" of Rust; that will not happen, he said, because there are 50+ different filesystems in Linux that will not be instantaneously converted. > Bottomley said that as more of those semantics get encoded into the bindings, they will become more fragile from a synchronization standpoint > But Ts'o pointedly said that not everyone will learn Rust; if he makes a change, he will fix all of the affected C code, but, "because I don't know Rust, I am not going to fix the Rust bindings, sorry".
- ysw0145 2y agoHaving more options available in the Linux kernel is always beneficial. However, Rust may not be the solution for everything. While Rust does its best to ensure its programming model is safe, it is still a limited model. Memory issues? Use Rust! Concurrency problems? Switch to Rust! But you can't do everything that C does without using unsafe blocks. Rust can offer a fresh perspective to these problems, but it's not a complete solution.
- drdo 2y agoBut unsafe blocks are available! And you should use them when you have to, but only when you have to. Using an unsafe block with a very limited blast radius doesn't negate all the guarantees you get in all the rest of your code.
- sanxiyn 2y agoNote that unsafe blocks don't have limited blast radius. Blast that can be caused by a single incorrect unsafe block is unlimited, at least in theory. (In practice there could be correlation of amount of incorrectness to effect, but same also could be said about C undefined behavior.) Unsafe blocks limit amount you need to get correct, but you need to get all of them correct. It is not a blast limiter.
- weinzierl 2y agoYes, they don't contain the blast, but they limit the places where a bomb can be, and that is their worth.
- foldr 2y agoGenerally speaking yes, but there could be a logic error somewhere in safe code that causes an unsafe block to do something it shouldn’t. For example, a safe function that is expected to return an integer less than n is called within an unsafe block to obtain an index, but the return value isn’t actually less than n. In that case the ‘bomb’ may be in the unsafe block, but the bug is in the safe code.
- pjmlp 2y agoThe disconnect section of the article is a good example of exactly on how not to do the things, and how things can turn out sour if the existing community isn't taken for the ride.
- pornel 2y agoI don't get how can each file system have a custom lifecycle for inodes, but still use the same functions for inode lifecycle management, but apparently with different semantics? That sounds like the opposite of an abstraction layer, if the same function must be used in different ways depending on implementation details. If the lifecycle of inodes is filesystem-specific, it should be managed via filesystem-specific functions.
- phkahler 2y ago>> I don't get how can each file system have a custom lifecycle for inodes, but still use the same functions for inode lifecycle management, but apparently with different semantics? I had the same question. They're trying to understand (or even document) all the C APIs in order to do the rust work. It sounds like collecting all that information might lead to some [WTFs and] refactoring so questions like this don't come up in the first place, and that would be a good thing.
- crest 2y agoI assume it's supposed to work by having the compiler track the lifetime of the inodes. The compiler is expected to help with ephemeral references (the file system still has to store the link count to disk).
- sandywaffles 2y agoI understood it as they're working to abstract as much as is generally and widely possible in the VFS layer, but there will still be (many?) edge cases that don't fit and will need to be handled in FS-specific layers. Perhaps the inode lifecycle was just an initial starting point for discussion?
- deleted 2y ago[deleted]
- DSMan195276 2y ago> but still use the same functions for inode lifecycle management I'm not an expert by any means but I'm somewhat knowledgeable, there's different functions that can be used to create inodes and then insert them into the cache. `iget_locked()` that's focused on here is a particular pattern of doing it, but not every FS uses that for one reason or another (or doesn't use it in every situation). Ex: FAT doesn't use it because the inode numbers get made-up and the FS maintains its own mapping of FAT position to inodes. There's then also file systems like `proc` which never cache their inode objects (I'm pretty sure that's the case, I don't claim to understand proc :P ) The inode objects themselves still have the same state flow regardless of where they come from, AFAIK, so from a consumer perspective the usage of the `inode` doesn't change. It's only the creation and internal handling of the inode objects by the FS layer that depends based on what the FS needs.
- hu3 2y agoSome of the comments below the lwn.net page are rather disrespectful. Imagine getting this comment about the open source project you contribute to: "Science advances one funeral at a time"
- deleted 2y ago[deleted]
- gwbas1c 2y agoMaybe they are asking the wrong questions? Does Rust need to change to make it easier to call C? I've done a bit of Rust, and (as a hobbyist,) it's still not clear (to me) how to interoperate with C. (I'm sure someone reading this has done it.) In contrast, in C++ and Objective C, all you need to do is include the right header and call the function. Swift lets you include Objective C files, and you can call C from them. Maybe Rust as a language needs to bend a little in this case, instead of expecting the kernel developers to bend to the language?
- tupshin 2y agoThis is not a notable challenge in rust, nor relevant to the article. The article is about finding ways of using rust to actually implement kernel fs drivers/etc. Note that any rust code in the kernel is necessarily consuming C interfaces. Bindgen works quite well for the use case that you are thinking. https://github.com/rust-lang/rust-bindgen https://github.com/rust-lang/rust-bindgen
- moomin 2y agoYeah, the Rust proponents are being significantly more ambitious. Not just the ability to code a file system in Rust, but do it in a way that catches a lot of the correctness issues relating to the complex (and changing) semantics of FS development.
- codetrotter 2y agoI’ve written Rust code that called C++ It wasn’t completely straightforward, but on the whole I figured out everything I needed to within a few days in order to be able to do it. Calling C would surely be very similar.
- duped 2y agoIt's actually pretty easy. All you need is declare `extern "C" fn foo() -> T` to be able to call it from Rust, and to pass the link flags either by adding a #[link] attribute or by adding it in a build.rs. You can use the bindgen crate to generate bindings ahead of time, or in a build.rs and include!() the generated bindings. Normally what people do is create a `-sys` crate that contains only bindings, usually generated. Then their code can `use` the bindings from the sys crate as normal. > in contrast, in C++ and Objective C, all you need to do is include the right header and link against the library.
- BiteCode_dev 2y agoGiven how those discussions usually go, and the scale of the change, I find that discussion extraordinarily civil. I disagree with the negative tone of this thread, I'm quite optimistic given how clearly the parties involved were able to communicate the pain points with zero BS.
- nickparker 2y agoI found myself reading this more for the excellent notetaking than for the content. I suspect the discussion was about as charged, meandering, and nitpicky as we all expect a PL debate among deeply opinionated geeks to be, and Jake Edge (who wrote this summary) is exceptionally good at removing all that and writing down substance.
- BiteCode_dev 2y agoCertainly. We are talking about extremely competent people who worked on a critical piece of software for years and invested a lot of their lives in it, with all pain, effort, experience, and responsibilities that come with that. That this debate is inscribed is a process that is still ongoing, and in fact, progressing, is a testament to how healthy the situation is. I was expecting the whole Rust thing to be shut down 10 times, in a flow of distasteful remarks, already. This means that not only Rust is vindicated as promising for the job, but both teams are willing and up to the task of working on the integration. Those projects are exhausting, highly under-pressure situations, and they last a long time. I still find that the report is showing a positive outcome. What do people expect? Move fast and break things? A barrage of "no" is how it's supposed to go.
- 0cf8612b2e1e 2y agoI am definitely of the opinion we need to rush away from C. Rust, Go, Zig, etc does not matter, but anything which can catch some of the repeated mistakes that squishy humans keep repeating. That being said, the file system is one of those infrastructure bits where you cannot make a mistake. Introduce a memory corruption bug leading to crashes every Thursday? Whatever. Total loss of data for 0.1% of users during a leap year at a total eclipse? Apocalypse. There is no amount of being too careful when interfacing with storage. C may have a lot of foibles, but it is the devil we know.
- sandywaffles 2y agoI wasn't clear and am not familiar enough with the Linux FS systems to know if this Rust API would be wrapping or re-implementing the C APIs? If it's re-implementing (or rather an additional API) it seems keeping the names the same as the C API would be problematic and lead to more confusion over time, even if initially it helped already-familiar-developers grok whats going on faster.
- CGamesPlay 2y ago> Almeida put up a slide with the equivalent of iget_locked() in Rust, which was called get_or_create_inode(). Seems like the answer is that it's reimplementing and doesn't use the same names.
- swfsql 2y agoI'm not familiar with those functions, but I had the impression they actually shouldn't have the same name. Since the Rust function has implicit/automatic behavior depending on how it's state is and how it's used by the callsite, and since the C one doesn't have any implicit/automatic behavior (as in, separate/explicit lifecycle calls must be made "manually"), I don't even see the reason for them to have the same name. That is to say, having the same name would be somehow wrong since the functions do and serve for different stuff. But it would make sense, at least from the Rust site, to have documentation referring to the original C name.
- brodouevencode 2y ago> about the disconnect between the names in the C API and the Rust API, which means that developers cannot look at the C code and know what the equivalent Rust call would be Ah, the struggle of legacy naming conventions. I've had success in keeping the same name but when I wanted an alternative name I would just wrap the old name with the new name. But yeah, naming things is hard.
- adastra22 2y agoOne of the two major problems in computer science (the other two being concurrency and off-by-one errors).
- simon04 2y agotl;dr?