11 ms·
Move semantics in Rust, C++, and Hylo
- deleted 2y ago[deleted]
- einpoklum 2y agoQ: What's "Hylo"? Should I have heard of it? A: It's a niche programming language the author is involved with. It's not widely-used enough to get its own Wikipedia page. It used to be called "Val". See: https://www.hylo-lang.org/ https://www.hylo-lang.org/
- Gualdrapo 2y agoMaybe it's just me, but am no fan of they using the keyword ´fun´ to define a function. Nor Rust's ´fn´. Also is it a bit strange they wrote "rust" along all the article instead of "Rust"?
- amaurose 2y agoIts the brain child of Dave Abrahams, who is rather big in C++. https://www.youtube.com/watch?v=5lecIqUhEl4 https://www.youtube.com/watch?v=5lecIqUhEl4
- bluetomcat 2y ago> So apparently, move does not prevent generation of a copy, but the empty string instead of expected text “Dave” is very interesting. Apparently, after termination of show after the move, the object is invalidated. This does not affect the Person object, but only the string object. This is a shallow understanding of C++. It happens because the Person object is a POD type that doesn't define a move constructor, and the compiler creates a default one that calls the move constructors of the members. The string member has a well-defined move constructor, but the primitive uint8_t type doesn't.
- flohofwoe 2y agoA move constructor/operator for POD or primitive types doesn't make any sense in the first place though (also AFAIK an object that contains a std::string - like Person - is definitely not a POD?). Even if Person had a manually provided move-constructor and move-assignment-operator, a move would still perform a flat copy from the source to the destination object.
- gpderetta 2y agoCorrect on all accounts. It is definitely not a POD nor a standard layout type (the modern version of POD).
- mort96 2y agoPerson has an implicitly generated constructor and destructor which calls std::string's constructor and destructor. It's non-POD.
- bluetomcat 2y ago> It's non-POD. For a stricter definition of POD which requires that byte-by-byte copies are possible. More informally, it's a POD because it only defines members and all the constructors and destructors are implicitly generated.
- flohofwoe 2y ago
- bluescarni 2y ago> So apparently, move does not prevent generation of a copy, but the empty string instead of expected text “Dave” is very interesting. Apparently, after termination of show after the move, the object is invalidated. This does not affect the Person object, but only the string object. Recognize that I speak about a factual behavior on the hardware. I think we have undefined behavior here. And no compilation error. There is a lot of wrong in this paragraph: - a "copy" was not generated, at least not in the sense that the actual content of the string was copied anywhere; - there's no undefined behaviour here and no invalidation of the string. Standard library types are required to be left in an unspecified but valid state after move. "Valid" here means that you can go on and inspect the state of the string after move, so you can query whether it is empty or not, count the number of characters, etc. etc. "Unspecified" means that the implementation gets to decide what is the status of the string after move. For long enough strings, typical implementation strategy is to set the moved-from string in an empty state.
- flohofwoe 2y ago> at least not in the sense that the actual content of the string was copied anywhere ...unless it's a short string within the limits of the small-string-optimization capacity. I think what confuses many people is that a C++ move assignment still can copy a significant amount of bytes since it's just a flat copy plus 'giving up' ownership of dangling data in the source object. For a POD struct, 'move assignment' and 'copy assignment' are identical in terms of cost.
- mort96 2y agoI mean it'll copy 3 pointers worth of data in all cases. It's just that for short strings, those 3 pointers worth of data contains the text of the string.
- fluoridation 2y agoI feel like that's a pedantic detail. True, yes, but irrelevant. You may as well also point out that the return address is going to be copied to the instruction pointer when the constructor returns.
- einpoklum 2y agoNot sure why the author compares Rust's: println!("{} is {} years old", person.name, person.age); with C++: cout << person.name << " is " << unsigned(person.age) << " years old" << endl; ... while C++ actually has: println("{} is {} years old", person.name, person.age); essentially identical to Rust. See: https://en.cppreference.com/w/cpp/io/println https://en.cppreference.com/w/cpp/io/println
- glandium 2y agoProbably because it's very new (C++23)
- vlovich123 2y agoWell C++23 is fairly new so they probably just didn't know about it?
- cjfd 2y agoSome people are noticing that println is very new. But there already is https://github.com/fmtlib/fmt https://github.com/fmtlib/fmt and it has been there quite a long time.
- Philpax 2y agoThat would require introducing a dependency, which is a digression from the point of the article and would complicate reproduction for the reader.
- deleted 2y ago[deleted]
- eterevsky 2y agoIn C++ you can force the move of the parameter by wrapping it with std::move() this should take care of unnecessarily cloning the argument in the example.
- masklinn 2y agostd::move does not force anything , it is a cast to an rvalue reference (a movable-from). Whether the object is moved depends on whether the target / destination / sink cares.
- deleted 2y ago[deleted]
- fluoridation 2y ago>Apparently, after termination of show after the move, the object is invalidated. This does not affect the Person object, but only the string object. Recognize that I speak about a factual behavior on the hardware. I think we have undefined behavior here. And no compilation error. The std::string is not invalidated, it's reset to its empty state (i.e. null pointer and zero length). Standard classes are all in defined, valid states after being moved, such that using them again is safe. User-defined classes may be coded to be left in either valid or invalid states after being moved. It's the responsibility of the programmer to decide which is appropriate according to the situation. There are valid reasons to want to reuse a moved object. For example, you might want to force the release an object's internal memory: std::string() = std::move(s); It's somewhat unfortunate that there's no way to signal to the compiler than an object is not safe for reuse, though.
- account42 2y agoWhile the language doesn't forbid use after move, occurences of it are most likely a programmer error. Which is why clang-tidy has the bugprone-use-after-move check.
- alkonaut 2y agoThis sounds like an enormous footgun (but as I understand it there are warnings that will tell you). An object isn't "valid" in any reasonable business logic sense just because the fields are initialized to anything at all, such as their default state? If the valid state of a Person is "the name is not empty " and this is enforced by a constructor then I don't want the program to ever have Person object floating around with a blank name? I either want a compiler error (good) or an immediate crash at runtime (bad), but at least I don't want an invalid object in a still running program (worse). Maybe I misunderstand what the reset was or how big this risk is though.
- fluoridation 2y ago>An object isn't "valid" in any reasonable business logic sense just because the fields are initialized to anything at all, such as their default state That very much depends on your use case. >If the valid state of a Person is "the name is not empty " and this is enforced by a constructor then I don't want the program to ever have Person object floating around with a blank name If you have such strict requirements then you shouldn't be moving around Persons to begin with. You should just be using std::make_unique() and then moving the pointer. Person should not even have a move constructor defined. If you code your class such that it's possible to let it reach an invalid state, that's no one's fault but your own.
- saghm 2y ago> I think before rust, language designers mixed up the various properties these values can have. As a result, many incomprehensible designs were the result. rust models the most important memory-related properties through its two call conventions (passing or borrowing). And Hylo moves even more properties into the call conventions. Namely, Hylo uses the keywords let, set, sink, and inout. This way Hylo additionally represents e.g. initialization (rust models this with a separate type). Is anyone able to clarify what's meant by "initialization" here and what "separate type" Rust uses for this (e.g. something defined specifically for each type getting passed this way, or a generic warpper type in the standard library)? Offhand, my understanding is that three of the Hylo keywords listed correspond to passing by ownership, shared reference, or mutable reference in Rust, and whichever doesn't correspond to one of those is something that a separate type if used for in Rust, but I'm not confident that my understanding is correct because the only thing I can think of that might be related to "initialization" is constructors, which Rust notably does _not_ have any formal concept of in the language, since functions that return types are just like any other function implemented on a type without a self parameter. I'm also not completely sure what the intended distinction is being made between whatever separate type is and references in Rust, since a reference is also a separate type than the type of the value of references. I could imagine someone might think that references are different than user-defined types in a way that other standard library types like Box and Arc aren't, but I'd argue that the unique syntax that references have is actually not that significant, and semantically being located inside std makes them far closer to references in terms of potentially behaving in special ways due to them having access to certain unstable APIs around things like allocations and fact that std is developed in tandem with the compiler, which leaves the door open for those types to take advantage of any additional internal APIs that get added in the future.
- hmry 2y agoMy best guess is they're referring to writing functions that initialize something using an "out" parameter in Hylo, which would be equivalent to a "&mut MaybeUninit<...>" parameter in Rust.
- Measter 2y agoThey mean whether the value is properly initialized, as in all the bytes that make up that value have set values that are valid for that type. For example, in Rust the only valid values a boolean can have are 0 and 1, anything else is invalid. Notably, in the abstract machine, bytes actually have 257 values: 0-255 and uninitialized. Uninitialized means that an initialized value was never written to it. Reading a value that is not properly initialized is undefined behaviour, and optimization passes can result in unpredictable changes in behaviour of the code. The type they mentioned is MaybeUninit (https://doc.rust-lang.org/std/mem/union.MaybeUninit.html https://doc.rust-lang.org/std/mem/union.MaybeUninit.html), which is used to represent values that are not fully initialized. It's worth reading the documentation for that type.
- fuhsnn 2y agoCopy or move for C++ is just choosing which constructor/assignment overload to call. I believe it's possible to make C++ move-by-default if one go through the trouble of overloading every class you use with custom move procedures.
- starlite-5008 2y ago[dead]
- quietbritishjim 2y agoMost explanations of C++'s std::move fail because they don't focus on its actual effect: controlling function overloading. Most developers have no trouble getting the idea of C++'s function overloading for parameter types that are totally different, e.g. it's clear what foo("xyz") will call if you have: void foo(int x); void foo(std::string x); It's also not too hard to get the idea with const and mutable references: void foo(std::string& x); void foo(const std::string& x); Rvalue references allow another possibility: void foo(std::string&& x); void foo(const std::string& x); (Technically it's also possible to overload with rvalue and non-const regular references, or even all three, but this is rarely done in practice). In this pairing, the first option would be chosen for a temporary object (e.g. foo(std::string("xyz")) or just foo("xyz")), while the second would be chosen if passing in a named variable (std::string x; foo(x)). In practice, the reason you bother to do this is so the the first overload can pilfer memory resources from its argument (whereas, presumably, the second will need to do a copy). The point of std::move() is to choose the first overload. This has the consequence that its argument will probably end up being modified (by foo()) even though std::move() itself does not contain any substantial code. All of the above applies to constructors, since they are functions and they can also be overloaded. Therefore, the following function is very similar in most practical situations since std::string has overloaded copy and move constructors: void foo(std::string x);
- ajross 2y ago[flagged]
- quietbritishjim 2y ago> This language feature is best understood as an attempt to mitigate the complexity of this other language feature That doesn't really make any sense. Move semantics aren't meant to mitigate complexity of function overloading. It's more like, it uses function overloading as part of its implementation. I do strongly prefer Rust's strategy of move being the first class citizen (and being destructive), with copy (/clone) layered over the top. But of course C++ got where it is for historical reasons. And I've never used Rust properly so I don't really know if the grass is greener. > Rust isn't as far back on that road is it pretends and is likely catching up. This is just vacuous trolling.
- khold_stare 2y agoI see some confusion in the comments about C++ moves. I wrote an article in 2013 after it clicked for me: https://kholdstare.github.io/technical/2013/11/23/moves-demystified.html https://kholdstare.github.io/technical/2013/11/23/moves-demy... . It goes over motivation, how it works under the hood etc, has diagrams if you are a more visual learner.
- pjmlp 2y ago> We learned that working on pointers directly often leads to memory bugs. So we introduced references. Minor pedantic correction, references predate having pointers all over the place, in most systems languages. C adopting pointers for all use cases isn't as great as they thought.
- Thorrez 2y ago>I compiled the C++ examples with godbolt with “x86-64 gcc (trunk)” and “-Wall -Wextra -Wno-pessimizing-move -Wno-redundant-move”. Edit: everything below is incorrect. -Wno-pessimizing-move is automatically enabled by -Wall, so doesn't need to be specified manually. -Wno-redundant-move is automatically enabled by -Wextra, so doesn't need to be specified manually.
- quuxplusone 2y ago-Wno-foo is turning off those warnings, not turning them on.
- Thorrez 2y agoWow, thanks. The gcc documentation appears to have a problem. It lists -Wreorder as a warning, and says it's enabled by -Wall . It lists -Wno-pessimizing-move as a warning, and says it's enabled by -Wall . I think the documentation should be edited to not list -Wno-pessimizing-move , and instead list -Wpessimizing-move . https://gcc.gnu.org/onlinedocs/gcc-9.1.0/gcc/C_002b_002b-Dialect-Options.html https://gcc.gnu.org/onlinedocs/gcc-9.1.0/gcc/C_002b_002b-Dia...
- cpp_noob 2y agostruct Person { string name; uint8_t age; }; isn't this missing a move constructor? Person::Person(Person&& p) : name(std::move(p.name)), age(p.age) {} or is C++ able to make these implicitly now?
- Maxatar 2y agoThe move and copy constructors are implicit.
- nayuki 2y agoSome basic things in the article appear to be factually wrong. > Then we ask us the following questions: > 1. When we passed Dave to show, did we create a copy? > 2. If so, how do we avoid creating a copy? > C++ example > 1. Yes. You can insert cout << "Person record is at address " << &p << endl; before the call of show as well as the beginning of show. This reveals different memory addresses of the record. Judging copies by the object's address is incorrect methodology. In both C++ and Rust, "moving" an object will still copy the struct fields, but will avoid copying any of the pointees (such as the variable-size array that the string owns). > 2. Replace void show(Person person) with void show(Person& person). So only the function needs to change. The caller does not have to adapt to it. Passing by reference is a different concept to moving. While the author used this approach for C++, they did not use the same approach for Rust. This is comparing apples to oranges.
- ajross 2y ago> In both C++ and Rust, "moving" an object will still copy the struct fields, but Most people consider a shallow copy a "copy", certainly a shallow copy isn't a "reference"! One of the big problems in this space is in fact the divergence of terminology that leads to arguments like this. The introduction of move semantics to C++ was a terrible, terrible mistake; not because it doesn't solve a real problem but because the language is objectively much worse now as a routine tool for general developers. People used to hack on code to implement features, now they get confused over and argue about how many "&" characters they need in a function signature. It was a problem that was best left unsolved, basically.
- webnrrd2k 2y agoRe: "problem that was best left unsolved" This is a good example of a hard-won life lesson... There might be a solution to a problem, but the solution is worse than the original problem. I semi-jokingly call this "the healing power of apathy". The reality of it is that, sometimes, there are problems in life where benign neglect is the best response.
- otabdeveloper4 2y ago> now they get confused Sounds like a skill issue. Maybe they should go shopping. Jokes aside though, yeah, move semantics is taught bad. Once you start using it (say, with a unique_ptr in a container) it will quickly start making sense.
- Night_Thastus 2y agoI can't say examples like this sell me on Rust, coming from C++. I need to manually to_string(), every single time I want to use strings? And that bizarre scoping of Person p feels very un-intuitive. How would you work around that if you need to keep using it after show()? (Which is an extremely common use case)
- Slyfox33 2y ago"Dave" by itself is basically the same as in c++, just a pointer to a string literal. Dave.to_string() is like std::string {"Dave"}, it allocates a heap based string from said literal. So you can use "Dave" perfectly fine if you just want a string literal.
- winrid 2y agoto_string() gives you an owned string (like std::string) vs a borrowed string slice (kind of like char*). If you already have an owned string you don't need to do that obviously If you need to keep using Person after calling show() then don't pass ownership to show() - you can pass a reference or a mutable reference, or use Rc<> etc
- aseipp 2y agoA raw string literal gets embedded into the binary's data section at compile time, just like it would in C or C++. What this means is that the type of the string literal is actually a reference (to an underlying memory address). And so it has type '&str' which reflects the fact you are using a reference to a value that exists somewhere else. The type 'String' is instead an "owned" type, which means that it is not a reference, and instead a complete value and has a copy of the data. to_string() will create a String (owned value) from a &str (reference) by copying it. This is no different than if you had a global static compile-time string in C and you wanted to modify or update it: you would memcpy the global (statically allocated) string into a local buffer of the appropriate size and then modify it and pass it onward to other things that need it. You would not modify the static string in place. In short, no, you do not need to_string() every time you want to work with a string. You need it to convert a reference type to an owned type. Rust's type system is just used here to codify the more implicit parts of C or C++'s behavior that you are already familiar with, but the underlying bits and bytes behave as you would expect coming from C++. > And that bizarre scoping of Person p feels very un-intuitive. How would you work around that if you need to keep using it after show() You take a reference just like you would in C++. Possibly a mutable reference if you want to modify the thing and then use it afterwords. This is in the article as the "Advanced rust example" at the end, it's right there and not hidden or anything. It isn't really bizarre honestly; it's a matter of defaults. The difference is that Rust uses move-by-default, not copy-by-default or ref-by-default. Every time you write `x = y` for a given owned type, you are doing a move of `y` and into `x` and thus making `y` invalid. let g: &str = "Austin"; // statically allocated string let x: String = g.to_string(); // do a copy let y: String = x; // no copy, x is moved Once you internalize this a lot more stuff will make sense, or at least it did for me.
- w10-1 2y agoHasn't a language feature failed if even experts disagree on it? How would lay developers ever use it? This is not an algorithmic nicety; it's supposed to be second nature to write and automatic to read. And it seems weird to omit Swift from this comparison, since Swift seems to have the most user-friendly (but incomplete?) implementation of move-only types.
- Maxatar 2y agoNot even the people who implement C++ compilers can agree on how certain C++ features are supposed to work.
- enugu 2y agoIn this discussion of a specific point in the post, the promise of Hylo language and mutable value semantics can be overlooked. Namely, we get a lot of the convenience of functional programming (mutating one variable doesn't change any other variable) with the performance of imperative languages (purely functional data structures have higher costs relative to in-place mutation and are more gc-intensive). https://docs.hylo-lang.org/language-tour/bindings https://docs.hylo-lang.org/language-tour/bindings