5 ms·
Rust sits in a really weird spot. It's too high-level for a lot of low-level work, and too low-level for a lot of high-level work. Example for the first case:
by rbehrends 9y ago
Rust sits in a really weird spot. It's too high-level for a lot of low-level work, and too low-level for a lot of high-level work.
Example for the first case: Writing a garbage collector runtime in Rust has most of the same problems in Rust as in C, because you have to write most of it in unsafe code, where Rust inherits much of C's undefined behavior w.r.t. pointers via LLVM. In short, you have largely the same problems and have added a hard dependency on Rust.
For high-level work, almost all [1] of what Rust gives you is memory safety and that comes at the price of dealing with a LOT of extra language complexity. But aside from dynamic memory management, memory safety isn't hard (we did that back in the 1970s and 1980s), and for dynamic memory management, we can get memory safety with a garbage collector and much less complexity. So Rust is primarily of interest for those use cases where garbage collection is not an option.
While that still gives you plenty of interesting use cases for Rust, there are also plenty of programming niches that it serves poorly.
[1] People will also mention "fearless concurrency", but guaranteeing the absence of data races is not hard. That more languages don't do it is partly because they simply neglected that aspect [2], but also because any mechanism – including Rust's – for doing so inherently constrains your options w.r.t. concurrency [3]. Plus, avoiding data races is the easy part of getting concurrency right.
[3] Concurrent Pascal had guaranteed absence of data races in the absence of pointers in the 1970s, Eiffel had done it with pointers in the 1980s, and there was a plethora of research in the 1990s to do it in various other ways.
[3] For example, there are plenty of use cases, such as certain idempotent operations, where data races are not only perfectly safe, but also desired for performance. There are also use cases where you can prove that no data races occur, but a type system cannot easily capture that.
- nine_k 9y agoYour critique of Rust can be largely applied to C++. Maybe the latter is a niche language, but that niche was not served by many offerings up until recently, and C++ is still going strong, despite being less safe than Rust.
- tdbgamer 9y agoIt's too high-level for a lot of low-level work, and too low-level for a lot of high-level work. This is true in the very specific cases that you gave, but I believe that is the minority of use cases, not the majority. Even the example of writing a GC that requires tons of unsafe code, that is not a good argument for making all the code unsafe. All the unsafe GC code would be abstracted away into a module and would be more obvious to those looking at it that they will need to be watchful for undefined behavior. Now you can proceed writing the rest of the project in safe, simple Rust. People will also mention "fearless concurrency", but guaranteeing the absence of data races is not hard Maybe for developers that are very familiar with the race conditions of parallel code, but definitely not for most people. Even seasoned developers will make mistakes with simple multithreaded code. Also, the reasoning behind "x is easy so why do I need my language to check it for me" is questionable. The whole point is that you have a guarantee. Have you never had a compiler catch a stupid mistake before it happened and felt relieved? I doubt it. Now imagine if instead of debugging stupid data races in your parallel code you can spend that time optimizing and improving it. I fail to see how this can be viewed as negative. Sure Rust doesn't cover 100% of use cases, but it definitely covers more than you're implying. It's low-level enough that Redox OS can be written in Rust, but high-level enough that Firefox is now outpacing other browsers and parallelizing everything with Rust.
- rbehrends 9y ago> All the unsafe GC code would be abstracted away into a module and would be more obvious to those looking at it that they will need to be watchful for undefined behavior. That code that could be "abstracted away" would be "virtually all the code" in my example. > Maybe for developers that are very familiar with the race conditions of parallel code, but definitely not for most people. Even seasoned developers will make mistakes with simple multithreaded code. I'm not talking about manually guaranteeing absence of data races. I mean absence of data races as a language feature. > Also, the reasoning behind "x is easy so why do I need my language to check it for me" is questionable. This is not at all what I was talking about. You completely misunderstood me.
- steveklabnik 9y agoI think you'd be surprised, even operating systems, the canonical unsafe activity, has a relatively low percentage of unsafe code. For example, https://doc.redox-os.org/book/introduction/unsafes.html https://doc.redox-os.org/book/introduction/unsafes.html says > A quick grep gives us some stats: the kernel has about 70 invocations of unsafe in about 4500 lines of code overall.
- rbehrends 9y ago> I think you'd be surprised, even operating systems, the canonical unsafe activity, has a relatively low percentage of unsafe code. My example was a GC runtime, not an OS kernel. If I have only very little unsafe code, then I could just do that in C and the rest in whatever other high-level language suits my project and not see any difference. The bigger problem – where Rust failed to pick some low-hanging fruit, IMO – is that "unsafe" is too much like the bad parts of C. There is no medium position between "everything is defined and memory-safe" and "everything may explode at a moment's notice". My most practical need for a low-level language is a language that is in that in-between position: semantics that remain easy to comprehend and predictable even if there are no static guarantees, and where I have to use a different strategy for software assurance. The point here is that for such a language I can resort to alternate validation tools (think Ada and SPARK for an example). Rust's unsafe mode does not handle that situation well because (like C) it does not provide a foundation for alternate validation strategies. It's perhaps also worth pointing out that I have a formal methods background. In short, I've done formal specifications/proofs for software before. In this context, safe Rust has a fairly high cost for only providing memory safety (and few other guarantees), and unsafe Rust is not a good foundation (or at least, not much better than C) for bringing advanced tools to bear.
- adwhit 9y agoYou might be interested in this paper [0] wherein the authors implement a high-performance GC in Rust. Quoting the abstract: We find that Rust’s safety features do not create significant barriers to implementing a high performance collector. Though memory managers are usually considered low-level, our high performance implementation relies on very little unsafe code, with the vast majority of the implementation benefiting from Rust’s safety. We see our experience as a compelling proof-of-concept of Rust as an implementation language for high performance garbage collection. [0] http://users.cecs.anu.edu.au/%7Esteveb/downloads/pdf/rust-ismm-2016.pdf http://users.cecs.anu.edu.au/%7Esteveb/downloads/pdf/rust-is...
- rbehrends 9y agoThat seems to be a bit misleading. What they seem to do, inter alia, is expose memory addresses as a safe type in Rust, with pointer arithmetic and dereferencing simply declared safe without it actually being so. There is no check that the underlying address actually points to valid memory, satisfies aliasing rules, etc.
- burntsushi 9y agoDid you read the paper? Dereferencing is not "simply declared safe." There's an entire section of the paper that goes over the API of the Address type, and explicitly points out that dereferencing is considered unsafe. Their conclusion runs directly contrary to your stated claims: > We found that the Rust programming model is quite restrictive, but not needlessly so. In practice we were able to use Rust to implement Immix. We found that the vast majority of the collector could be implemented naturally, without difficulty, and without violating Rust’s restrictive static safety guarantees. In this paper we have discussed each of the cases where we ran into difficulties and how we overcame those challenges. Our experience was very positive: we enjoyed programming in Rust, we found its restrictive programming model helpful in the context of a garbage collector implementation, we appreciated access to its standard libraries (something missing when using a restricted language such as restricted Java), and we found that it was not difficult to achieve excellent performance. Our experience leads us to the view that Rust is very well suited to garbage collection implementation.
- deleted 9y ago[deleted]
- bitwize 9y ago> But aside from dynamic memory management, memory safety isn't hard (we did that back in the 1970s and 1980s), and for dynamic memory management, we can get memory safety with a garbage collector and much less complexity. At a severe cost in performance. Static object lifetimes cover 99.9% of a garbage collector's use cases, without the performance cost of GC, nor the nondeterministic runtimes. We first saw static object lifetimes come into their own with C++'s value semantics; Rust refines and clarifies the idea and makes memory safety an inherent part of the language itself. Static object lifetime is to GC what static types are to dynamic types. Lisp is 1960s tech. It has failed, and been replaced with something much better.
- rbehrends 9y agoI think your terminology is off. Static lifetime means that objects exist for the entire duration of the program and do not affect memory management at all, automated or manual. If you're talking about automatic variables, then things already become more complicated. For starters, we have to assume that we don't deal with value types (which will end up on the stack, one way or the other), but with local variables that reference heap objects. Second, we have to distinguish between tracing and reference-counting GCs. A modern tracing garbage collector will have cost for such temporary allocations comparable to `alloca()` and those allocations will typically be inlined. The cost of deallocating a short-lived stack object is zero (yes, zero). This is possible because GCs (unlike manual memory management schemes) are compacting. Whether one approach or the other comes out on top is very situational. More importantly, I dispute the 99.9% as a vast exaggeration. There are plenty of important use cases (such as persistent data structures, shared caches, etc.) where unique ownership is insufficient; Rust requires you to use either copying or reference counting when you run into shared ownership scenarios, both of which are more expensive than tracing GC (naive reference counting is already one of the more expensive memory management methods known, and atomic reference counting is especially expensive). If you use reference-counting GC, then for any program that satisfies Rust's borrow checker, the optimizer can eliminate reference counts that satisfy the same conditions (assuming that the optimizer knows about them because they're part of the language semantics). This is largely what Swift does, for example. Finally, there is deferred reference counting, which incurs only trivial overhead for objects with automatic lifetime (on the order of a fraction of a percent). This is because this algorithm incurs real cost only when pointers are written to global or heap locations; this is also why it's seen limited use in practice: it's excellent for objects with automatic lifetimes and does not rely on the generational hypothesis, but those do not constitute 99.9% of all use cases. If they were, deferred reference counting would have a far more prominent role. This does not even account for the fact that when there is overhead, that overhead is generally trivial in an imperative language with value types. There are use cases, of course, where a tracing GC is an inappropriate choice, but that is not because of throughput. Tracing GCs make interoperability with other GCs different, for example, and have implicit memory overhead that may be prohibitive in large applications such as a web browser (that can easily consume gigabytes of memory on a laptop or desktop machine). That said, there are alternative approaches to garbage collection that do not have those problems.