12 ms·
You can't Rust that
- joeconway 9y agoThank you Armin. Your rust work for sentry has been a great primer in the language for me.
- pcwalton 9y agoI really like the way you captured one of the fundamental differences between Rust and C++ as "Things Move". That's an interesting way to summarize it that I hadn't really considered before—and I designed a lot of that system :)
- jnordwick 9y agoIs is even remotely true though? C++ probably moves things around more than rust, and I thought rust would want to reduce cache churn. It isn't like GC were things magically change locations. I'm not really sure I understand what he's getting at with that description.
- the_mitsuhiko 9y ago> I'm not really sure I understand what he's getting at with that description. I'm getting at that an object in C++ is generally immovable. If it does move (which is new in C++11 i think?) a move ctor is invoked. Rust does not have a move ctor, it does not let you do things that would prevent moving an object in the first place (at least currently).
- gmueckl 9y agoA move in C++ is a special construct where a new object is allocated in a new memory location and then constructed from the existing object in a way that must guarantee that the existing object is invalidated in some way. This must be handled explicitly in the implementation. For example, for std::vector this means that the move-constructed container steals the internal pointer to the data buffer from the old vector and resets the old vector to empty. This avoids a big and usually superfluous copy operation.
- masklinn 9y ago> in a way that must guarantee that the existing object is invalidated in some way That statement is confusing as the spec very specifically notes that the moved-from object remains valid, so saying that it's invalidated is odd.
- gmueckl 9y agoYou are confusing the now invalid internal state with the continuing existence of an accessible object at the old location. So you can move from an object and then access it. The standard guarantees that the object is still around. But because its data got stolen, its internal state must now be something invalid or empty depending on the semantics of the object.
- cornstalks 9y agoYou’re using way too strong of language. A move in C++ doesn’t require the original object to be invalidated or even modified in any way. POD-ish types, for example, are “moveable” despite not being made empty or invalid after a move (their move is implemented as a copy, which is perfectly valid).
- MereInterest 9y agoAs an example of this, in `boost::asio`, a moved-from socket is in a valid state, and is ready to accept a new connection. https://www.boost.org/doc/libs/1_54_0/doc/html/boost_asio/reference/basic_stream_socket/basic_stream_socket/overload5.html https://www.boost.org/doc/libs/1_54_0/doc/html/boost_asio/re...
- masklinn 9y ago> You are confusing the now invalid internal state with the continuing existence of an accessible object at the old location. I'm not confusing anything, and the spec very explicitly notes that a moved-from object is in a "valid but unspecified state", not in an invalid state. > The standard guarantees that the object is still around. The standard's guarantees are significantly stronger than that, and furthermore > But because its data got stolen, its internal state must now be something invalid or empty depending on the semantics of the object. That is absolutely not a hard rule, a trivial move constructor is a copy constructor and does not affect the moved-from object in any way.
- staticassertion 9y agoThey're talking about move semantics, not actual memory moving around. In c++ you get copy semantics by default, requiring std::mov to get move semantics. In rust it is the opposite.
- masklinn 9y ago> In c++ you get copy semantics by default, requiring std::mov to get move semantics. And even then: 1. std::move only creates an rvalue reference, it doesn't necessarily move anything (not only does that depend on the presence of a move ctor & the receiver's arguments, there's still no requirement that anything actually moves) 2. a move in the C++ sense "empties" the object, the object is still there and accessible in "some valid but otherwise indeterminate state", objects can even fully specify the exact state they're in when moved-from (unique_ptr does)
- bergesenha 9y ago1. I thought the cast to an rvalue reference would make overloads taking rvalue references be chosen during overload resolution. Since this is deterministic I thought whether something is moved or not would be quite guaranteed
- jeremyjh 9y agoI think he just means the move constructor is not guaranteed to move anything. It is only expected to.
- snuxoll 9y agoThe rule of three/five is probably the most tedious thing to deal with in modern C++. Assuming all your member variables are C++ types it's not TERRIBLE, you should be able to handle everything in the initialization list (using std::move for the move constructor) - but it's still easy to forget to initialize a member. This is one area where Rust really makes life much easier.
- saagarjha 9y ago> C++ probably moves things around more than rust I don't think so, actually. As far as I'm aware moves are only done when explicitly asked for, because any implicit move would have the consequence of breaking references and being unsafe in general if you're not careful.
- wilun 9y ago> As far as I'm aware moves are only done when explicitly asked for Certainly not. Move are attempted (in the sense that the move constructor/operator= will be called if it exists) from rvalue references. Those are for example all temporaries. That has let accelerate all C++ programs using std objects by simply recompiling them (with the new versions of the objects supporting the move while it did not exist in the old versions)
- saagarjha 9y agoWell, it makes sense to move temporaries. There's no point in copying them again if the original won't be used.
- jcelerier 9y ago> As far as I'm aware moves are only done when explicitly asked for no, every time you have a function returning a temporary you can have a move, e.g. std::vector<int> myfun(); struct myclass { myclass(std::vector<int> v); }; here if you do myclass c{myfun()}; there can be a move
- Sharlin 9y agoAs far as I know Rust doesn't really do C++-style copy elision currently. So things passed by value get moved. Which is to say, bitwise copied (memmove) and the moved-from binding statically marked unusable unless it implements the Copy trait, basically promising that it doesn't contain owned references to external resources (in which case copying would violate the single-owner invariant). In C++ an object can be passed by reference, by copy (implicitly invoking the copy constructor), by move (implicitly invoking the move constructor), and additionally the compiler may choose to elide the copy/move. Whereas in Rust, you can only pass by reference (borrowing) or by bitwise copy (moving). There are no implicit copy or move constructors, and AFAIK at least currently the compiler doesn't hoist objects directly to the callee's stack frame.
- pcwalton 9y ago> AFAIK at least currently the compiler doesn't hoist objects directly to the callee's stack frame. Yes, it does. This is the "retptr" optimization.
- oconnor663 9y agoDoesn't the LLVM backend make tons of similar optimizations under the hood?
- pcwalton 9y agoYes, especially if callees are inlined.
- dbaupp 9y agoI don't think you're quite right with the consequences of C++ style copy elision. It fundamentally isn't nearly as important as in Rust because it doesn't implicitly copy. Especially pre-C++11, copy elision was critical to ensure that functions can return std::strings and std::vectors without doing expensive copies. This doesn't apply to Rust which started with move semantics built-in, and cheap move semantics at that (even cheaper than C++ in some respects, as objects don't need to be left in a valid-but-unspecified state). Stepping into the specifics of the options, you've classified things a little inconsistently: C++ allows arguments to be references or values, and the caller can choose for the latter to be by copy or by move. This the same as Rust, with references and values and the caller choosing by copy (with an explicit .clone()) or by move (default). You can argue that Clone isn't part of the language, or that being explicit doesn't count, but I don't think that's so interesting (certainly by move in C++ is often equally explicit: std::move). And, without move constructors, hoisting things into caller's stack frames is just a run-of-the-mill optimisations for returned values following the standard "as-if" rule, no need for explicit enabling in the standard. C++ needs it because part of the semantic model is running user-defined copy/move constructors on 'return', so the standard needs to make it okay to not do this in some cases, but without the user-defined code, it isn't necessary.
- oconnor663 9y ago> I thought rust would want to reduce cache churn I think the main way Rust reduces cache churn, is that it makes it safer to pass references/pointers around instead of making copies. For example, if I have f(char*) in C, I might be worried about what f is doing with that pointer. It might save it somewhere, or try to free it, or who knows what. In C++, I might prefer to write f(std::string) to avoid those worries, but that comes at the cost of unnecessarily copying the string. (You could pass it by reference to avoid the copy, but then you still have to worry about f saving the reference.) In Rust, I can safely write f(&str), and I know f can't do anything dirty with that reference.
- jstimpfle 9y ago> If I have f(char*) in C, I might be worried about what f is doing with that pointer. It might save it somewhere, or try to free it, or who knows what FWIW my assumption always is that it's "a borrow". I.e., the called function should not make any assumptions how the memory was allocated, what is its lifetime, or even try to free it or keep it in an associated datastructure that persists the lifetime of the function. The only thing it should normally assume is that it points to valid memory. In my experience the instances where ownership moves across function boundaries are far and far between, and can easily be documented, often simply by giving it an appropriate name.
- the_mitsuhiko 9y ago> I.e., the called function should not make any assumptions how the memory was allocated, what is its lifetime It needs to at least make the assumption that the memory will not be invalidated while the function did not return. In the presence of multi threading that in itself might already be surprisingly hard to guarantee sometimes. Also typically functions make assumptions about the immutability of things in such cases which are often just not true.
- jstimpfle 9y agoProblems with multithreading and "perceived immutability" are two more things that can largely be ruled out with a clear dataflow architecture. For example, I try to not have "cycles" such that a function call overwrites a memory location that was a (pointer) argument to the caller. I don't think that should happen - where did the called function get the pointer from without the caller knowing? An architecture with global data tables instead of OOP helps here, because it has much lower need for pointer arguments of unclear origin. I also agree very much with your approach "handles instead of pointers" (if by handles you mean integer array indices). Indices for example trivially survive reallocation of a dynamically growable array.
- nurettin 9y agoAt some point I stopped worrying and started passing by value. Compiler would sometimes optimize extra copies, the big objects that got copied would show up during profile. Works better than chasing move bugs if you are just writing ERP software and don't really care about cache misses. Edit: c++
- phkahler 9y agoIf pointers are useless, how do you create complex data structures. In a C++ program I have a struct that is nothing but 5 pointers (4 now since 2 can be stored as their XOR). I'm starting to wonder about this Rust thing that's been sounding so awesome...
- pcwalton 9y agoUnique pointers or reference counted pointers.
- gpm 9y agoPointers aren't useless... things don't move if you don't tell them to. The compiler can statically guarantee they don't move while you have a pointer to them (via references/lifetimes) or with unsafe code you can just tell the compiler that "I checked and I don't move this thing". I've had plenty of rust datastructures that are just a collection of pointers. The distinction I guess is that in rust you can move things assuming you have no references to them, while in c++ you can't really (I mean, move constructors, but that's really a fancy copy where the initial thing still exists in some form). And moving things is pretty normal - as is taking pointers to them - just not at the same time.
- jay-anderson 9y ago> things don't move if you don't tell them to This is something I've struggled a bit with when playing with Rust. I've had a hard time understanding whether I'm telling the compiler to move or copy. In other words I need to understand better what I'm telling it to do (I probably just need to buckle down and study the new book in this area better). It wasn't immediately obvious to me when I last played with it.
- steveklabnik 9y agoIt depends on the type, not what you say. Copy types copy, other types move.
- masklinn 9y ago
- jnordwick 9y agoI don't like the Things Move example. I'm not sure how true the general statement is (I'd never thought of it that way, but it isn't like how GC moves things around, and I'm not sure things are even more than in C++ -- I thought rust reduced unnecessary moves because would kill cache performance), but the example isn't entirely correct from my perspective. Return values that fit into a register will be returned in a register, and his example is an 8 byte struct, so that returns in a register. Return values larger than a register will add an implicit first argument that is a pointer to memory where the return value should be written to. In that sense, it is very similar to C++ in that you are initializing into an allocated buffer. As for "Refcounts are not Dirty", I would greatly disagree. Using refcounts to get around an overly aggressive borrow checker seems to be an ugly developing pattern in rust, and I feel they are giving away the performance many are fighting for by adding all these little inefficiencies to idiomatic rust. Add some refcounting here, add a Box or other indirection there, a chained monadic interface that can't short circuit and has to continually do error/null checks, etc... Soon it is death by a thousand papercuts. People fight hard for that extra 5% in performance only to have it taken away from them in interface and language issues. Edit: Forgot about handles. Ugh. Completely unacceptable when you want to grow your tree data structure and you have to do a realloc and basically copy every node. If your tree is complete, then you are copying the whole tree every time you start a new level. The conclusions sound more like ugly hacks, than what you would properly design.
- the_mitsuhiko 9y agoIf a value returned from a function actually moves or not is currently up for Rust to optimize. It's not something you can depend on. About the refcounts: since the counting is explicit (calls to clone()) they at least in my experience don't really show up. Most of the refcounted objects I deal with bump the refcounts once when some task spawns and decrements it when it ends. I have yet to see refcounts to change in hot code paths. //EDIT for your edit: > Edit: Forgot about handles. Ugh. Completely unacceptable when you want to grow your tree data structure and you have to do a realloc and basically copy every node. Sure, but that's not my point anyways. At any point you can fall down to writing unsafe code and building a safe abstraction on top of it. This is to help developers not run into walls. I don't think that handles are the best thing invented but I don't think "well you can't do that in Rust until we some time in the future" and not provide an alternative is a particularly good suggestion.
- mwcampbell 9y agoGiven the "things move" point, would it be feasible to use a compacting memory manager with Rust, e.g. for memory-constrained applications?
- kibwen 9y agoRust allows you to take interior pointers to things (important for performance), which precludes the ability to move objects within memory at random. But for "memory-constrained" applications like embedded devices/microcontrollers, fragmentation isn't a problem in the first place because you often don't have a heap. For long-running programs that do have heaps, picking a modern memory allocator (jemalloc, tcmalloc, et al) will go a long way towards reducing fragmentation. And if you really need compaction, you could probably design a Rust library to provide it for certain types (though the operations it could provide would likely be restricted).
- smaddox 9y agoIf you're memory constrained, why would you use a GC? Just use manual memory arenas. They're trivially simple once you've seen how to use them. Unfortunately, Rust currently requires breaking some conventions and using unsafe quite a bit to do this without overflowing the stack, but it's just an extra keyword or two compared to C, and the safety guarantees outside of the unsafe code make up for it.
- z3t4 9y agoI'm not a real programmer (as in someone who do not write low level code), but do real programmers actually rely on pointers - knowing that the data might move or change !? (I program in JavaScript where all values are immutable)
- tomsmeding 9y agoIf we have the same concept of "value" in Javascript, then values are certainly not immutable. E.g.: const a = {x: 1, y: 2}; console.log(a.x); // 1 a.x = 2; console.log(a.x); // 2 Note that 'a' remained constant indeed, but the object it points to can certainly take on different values.
- z3t4 9y agowith immutable I mean that a value (value 1 and 2 in your example) can never change, you can reassign the variable to a new value though. If I understand the article correctly values in Rust and C++ can change eg. the memory location of the bits representing the value. Making it possible to shoot yourself in the foot if you don't know how the internals work, more so then in JavaScript as it's much more to keep track of. Even though it seems to be Rust's motto to limit such cases. And I agree? that using const in JavaScript is like wearing a tin foil hat - it will rarely save you. If I have something that can change globally I make it upper case so it stands out and don't collide with other variables. And "use strict" should tell if you forgot to var(let/const) a variable. (unless there's a HTML id attribute with the same name) =)
- a_humean 9y agoIf that is what you mean, then yes you are right. Pointers point at bits of memory (its literally an address to a physical piece of memory or virtual memory allocated by something else - an OS for example) which store things like the 64bit floats '1' and '2' you are referring to in Javascript, and that is something someone writing C needs to think about pretty explicitly and is a source of errors. In Javascript you don't have to think about it very often in explicitly those terms as the runtime takes care of thinking about that for you. You create all of these string, numbers, objects, etc... that the runtime has to keep track of for you using pointers, and once it thinks you are done with that memory it frees it up. https://en.wikipedia.org/wiki/Garbage_collection_(computer_science) https://en.wikipedia.org/wiki/Garbage_collection_(computer_s... However, as I and other have said, while someone writing JavaScript doesn't have to think much about memory, you still have to think about references and values. Its also the case that the Garbage Collector isn't perfect, and sometimes as a JavaScript programmer you can accidentally create memory problems of your own: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Memory_Management https://developer.mozilla.org/en-US/docs/Web/JavaScript/Memo...
- LifeLiverTransp 9y agohttps://www.viva64.com/en/b/0324/ https://www.viva64.com/en/b/0324/
- skybrian 9y agoAs a non-Rust programmer, I'm finding the memory-mapped data example to be very opaque. Does anyone care to explain it?
- grayrest 9y agoIt's a contrived example and doesn't make a whole lot of sense aside from demonstrating what he means by handle. I'm not an expert but I'll have a go at explaining it. Start at the `Data` struct. It contains a Copy on Write (`CoW`) reference to a vector of bytes (`u8`) with a lifetime labeled `'a`. This is the Handle for the data. You get one by calling `Data::new` and passing in something that can be converted to the CoW. The example is hard coded to work with a vector of u32s (driven by the `Slice<u32>` in `Header`). To use it, you'd call `get_target` with an index and get a u32 back. The other methods on data are doing the pointer math (offset) and casting (`transmute`, `from_raw_parts`) the byte array into a slice of u32s in a safe way. I don't see anything verifying that the byte array passed in is, in fact, a bunch of u32s so I assume that's a given.
- the_mitsuhiko 9y agoIt’s actually a simplifiedversion of what we do. We deal with debug information files and write custom cache files thag look similar to that. https://github.com/getsentry/symbolic/tree/master/symcache/src https://github.com/getsentry/symbolic/tree/master/symcache/s...
- glenjamin 9y agoThe semantics of the final example sounds a lot like the concept of an Atom in clojure - https://clojure.org/reference/atoms https://clojure.org/reference/atoms Is this swap/deref pattern something that can or should be wrapped up into a create?