6 ms·
This seems a bit unfair to me. The author is not entirely wrong, but it feels like they don't give the full picture. The truth is in Rust you often don't have t
by nu11ptr 3y ago
This seems a bit unfair to me. The author is not entirely wrong, but it feels like they don't give the full picture. The truth is in Rust you often don't have to heap allocate at all and it is very good at avoiding this much of the time. Obviously, if building something like an AST in a parser, a naïve implementation is going to be calling the boxed allocator quite a bit which will have a lot of overhead, but for things like this I would just use an arena allocator (like bumpalo). So yes, sometimes you have to do a little more allocation, but I think it is greatly over stated. I do really like OCaml, however, and agree a compacting GC with bump allocator can be quite nice at times as well (and a compiler would probably be a good use case for one).
- BulgarianIdiot 3y agoThe default abstraction in Rust, in retrospect, will be considered poorly chosen. Moving, read borrowing, readwrite borrowing, copying, cloning are an unnecessary complication over what truly happens at both your hardware level, and between components of it, and on the network between machines. Everything is copied, end of the story, no modes. And if something is copied, but it doesn't change, you can defer the copy until it's changed (copy on write). When something goes out of scope, it's freed. That's it. It's so simple, but very few platforms truly go with it, despite it's pervasive underneath the platform and much simpler than anything else we're doing.
- dathinab 3y ago> Everything is copied, end of the story, no modes [..] (copy on write) [..] something goes out of scope, it's freed. this fundamentally doesn't work, in no language Because there is a fundamental difference between something containing a owning a sub-resource and something being allowed to use a resource as a sub-resource without owning it. And copying implicitly turns a onwed reference in a shared on (and in some cases). And while not all languages do express this difference, it still matter in how you use things, which things are exposed in public APIs, which things are thread save etc. For example on networks everything is _serialize_ which in rust term is more a clone then anything else (but actually more then just a clone in rust). While in the CPU things are normally "rust-Copy copied". But not always, hardware is complicated and trying to do "what hardware does" is basically never ever a good idea for high level abstractions (not doing so is somewhat the whole point of highlevel abstractions). As a side note: The fundamental abstraction of Rust is ownership (split into: owning, lend exclusively, lend shared) + scoped lifetimes (things get destructed when they go out of scope). Combined with semi-manual memory management. Scoped lifetimes is needed to archive a semi-manual memory management which was a must have design decision for the goals rust tried to archive. And move semantics are a direct consequence of this choices. The first is fundamental to pretty much any programming but not explicitly expressed in most. The difference between Copy and Clone is also not very rust specific, in many other languages (including GCed ones) it would be the distinction between a flat and a deep clone. There is pretty much no way to get ride of the abstraction (due to ownership of nested resources) but I agree that it could have been done a bit better the in rust, not that it's poorly designed IMHO, just not perfect. No one thing which is a sub-optimal decisions (and was known as such since they where done) are the specific traits signatures used for some of this mechanics. Through the problem is the "better" designs need features rust doesn't even fully have today and they where needed to be stabilized ~8 years ago. This is also where I see a potential rust 2.0 (in many years), the same core compiler but a changed standard library and a automatic transpiration step which allows you to use most of the rust 1.x crates without any changes to them (i.e. like a rust edition, but with a much larger change and just mostly backward compatible).
- BulgarianIdiot 3y agoYou said the fundamental concept in Rust is ownership. What I described is also fundamentally ownership, but much simplified. You can use something without owning it, because to "use it" you send a message to it, and get a message back (i.e. send input, get output). There are two ways to message something in this system: 1. To own it and message it directly. 2. To own a reference identifier to it, and message your own owner to message it via the reference and return you the resulting message. The distinction between Copy and Clone is Rust specific. It's not about deep vs non-deep, it's about whether something is moved by default or copied implicitly by default, and they wanted to copy some simpler scalars, like numbers, say, for pragmatic reason. But this odd exception, i.e. does something move or copy by default is very much a defect in Rust. It's a core behavior that differs based on whether the type implements copy or not. In the simplified model I propose everything is copied (deep) by default, but you don't pay the cost for larger structures being copied, due to COW. When everything you own is uniquely yours, making all copies deep copies becomes not only trivial, but also the fastest way to copy. And there's no reason to talk about "shallow copy". The concept is meaningless, you're used to it, you have a habit for it, but if people could deep copy everything, they would. "Shallow" copies are a limitation of languages where everything is wired by handles, pointers, and references in an arbitrary graph. And then a deep copy may never terminate, among other things. It's a mess. Rust's graph is not (as) arbitrary, per se, it's stricter, but it keeps many of the same defects, in part to maintain C compatibility. Meanwhile, your operating system has a hierarchy-based memory management (single ownership segments, at the hardware level a hierarchy of pages etc.), and COW in select places where as I noted immutability is present for some content, some of the time (like Unix forking). So does your file system. Everything that's mature does this: DAG for immutables, TREE for mutables. Programming languages are clearly not as mature yet.
- dathinab 3y ago> Rust's graph is not (as) arbitrary, per se, it's stricter, but it keeps many of the same defects, in part to maintain C compatibility. I don't thing a core requirement for the design of an language needed for it's targeted use-cases is an defect, not having it would be an defect. I guess that's just a bad choice of wording. > It's not about deep vs non-deep (side not I noticed my wording choice wasn't grate, with flat copy I meant copy of a continuous memory region, it could still be a nested type as long as there are no indirections). It's all about that. Through more specifically the expected runtime cost associated with it. Copy only works for flat memory copies, is normally very cheap and in turn is automatic. Clone works for everything including deep-clones but might be very expensive and in turn is not ever done automatically and no doing any magic CoW tricks or similar was a very deliberate decision as for the use case rust was designed with in mind you do not want a Clone happen at a "unexpected" time. Sure it's not perfect e.g. there is are "cheap-Clone" types (like Rc) where I would prefer rust to automatically clone in at least some contexts. > In the simplified model I propose everything is copied (deep) by default, but you don't pay the cost for larger structures being copied, due to COW. But this approach has a wide variety of problem iff you want to use it for a system language. Like it being trivial to accidentally introduce potentially very expensive cloning or accidentally using a ton of unneeded memory. And yes you could come up with all kinds of mechanisms to reduce this. But defaults matter and rust by default motivates you to not do potentially expensive clones. This comes at a cost, sure. But depending on you use-case this is exactly what you want. > It's a mess. but a practically very useful mess which has a lot of success in difference to any of the research languages which in the past had tried to archive similar things as you stated. Perfection isn't and was never the goal. You always have to have a compromise between usability, compatibility, maintainability, debugability and performance. And for a "system programming language" this means C-compatibility and fine grained control over when or where copies happen. Doesn't mean it couldn't work for languages focused on other use-case, and could be suited for typical web server or app use-cases. But even through rust is used a lot for this it wasn't designed for this use case.
- kaba0 3y agoThere are many cases where having identity is paramount. Plus rust is a low-level language, why should it not expose low level hardware details?
- pjmlp 3y agoOther than inline Assembly, what low level hardware details does it actually expose?
- kaba0 3y agoI meant manual memory management, in reference to having only-COW semantics wouldn’t make the rust fit for its target niche.
- pjmlp 3y agoWell, even C# has manual memory management primitives, as one example out of many others.
- kaba0 3y agoYes I know, and I am a huge proponent of GCd languages as there really are only a tiny handful of cases where having a GC would be a problem. But this is out of context here, where we are talking about a language whose explicit design goal is to manually manage memory without a GC. In that respect, COW wouldn’t fit the rest of the design criteria of the language, which was my point.
- pjmlp 3y agoYeah, but that is another matter, and has nothing to do with the language level. Even C is considered high level nowadays, in regards to hardware exposure.
- kaba0 3y agoThere are two contending definitions for low/high level languages. The more objective one only considers assemblies low-level, but that is hardly a useful categorization. The other one is about what can be explicitly controlled by idiomatic code, in that vein C is lower level than, say, Java/C# due to pointer arithmetics (and yes, I do know that both can actually do pointer arithmetic just fine, but I wouldn’t call it idiomatic, especially not in case of Java). Rust/C++ is on the same level, if not lower than C due to having native access to SIMD.
- cmrdporcupine 3y agoAgreed on your general points. When I'm writing in Rust I'm really thinking far less about memory than I would have thought. Ownership, yes. Way more than other languages, obviously. And it can get very frustrating at time. But this also has payoffs for thinking about concurrency. I do think that there are places where this overhead gets extremely taxing. And complicated nested trees of objects, iterating through them, etc. like you'd have in a compiler or query planner etc. is definitely one of those places. The ownership and type system constraints in Rust make its Iterator pattern actually quite obnoxious for these kinds of things.
- ynik 3y agoFor compilers, arena allocation is king. IMHO: arenas > GC > RAII > malloc/free. With manual memory management (whether RAII or not), it's usually not difficult to add arenas into the mix. On the other hand, in GC languages: many GCs only allow heap pointers to the beginning of objects, not to their interior. This means arenas can't used in those GC languages; every full collection actually has to trace the millions of AST nodes individually.
- BulgarianIdiot 3y agoAny memory model that's not recursive (i.e. arenas within arenas ... within arenas) is deeply unserious to me.
- dathinab 3y agoYou can allocate a part of a arena to form a new arena, at least theoretically. And in many cases for compliance you want to 1) limit the memory a specific sub-component of your code can use 2) make sure it always can have this memory 3) make sure it doesn't leak any memory ever. Arenas are probably the most reliable way to archive this. The main problem is if the sub-component needs to allocate memory which outlives it's life time. A common way around this is that the component delegating work to it has to provide that memory (e.g. windows kernel API) but that doesn't always scale well. Another is to use another memory management for it (e.g. GC) but then it isn't enforcing compliance anymore. Etc. it's always some compromis.
- BulgarianIdiot 3y agoAllocating within the given arena is simply an option. You can still be a component owning an arena, but delegate the allocation to an ancestor/sibling context/arena if you want that allocation to outlive you.
- dathinab 3y agoOne of the various things I placed in the `Etc.` ;=) But like all the other things listed it's not perfect. For example if I just pass a unconstrained access to a other components arena to a component you can no longer "guarantee" that the component doesn't leak anything as it could use the arena to allocate things besides the thing for which it was passed it to the component. So now you need to have additional steps to assure this isn't happening if you want to make sure as much as possible that everything is compliant. Through depending on what you do this aren't necessary complicated additional steps through it really depends.