13 ms·
A Taste of Rust
- steveklabnik 11y agoI am utterly thankful for new experience reports on Rust, especially for ones this well-written. Generally speaking, inaccuracies in such things are our fault, not the writers', due to a lack of documentation and or good examples. With that being said, a few notes: > It runs about five times slower than the equivalent program I'd be interested in hearing more about how these were benchmarked. On my machine, they both run in roughly the same time, with a degree of variance that makes them roughly equivalent. Some runs, the iterator version is faster. It's common to forget to turn on optimizations, which _seriously_ impact Rust's runtimes, LLVM can do wonders here. Generally speaking, if iterators are slower than a loop, that's a bug. > Rust does not have tail-call optimization, or any facilities for marking functions as pure, so the compiler can’t do the sort of functional optimization that Haskell programmers have come to expect out of Scotland. LLVM will sometimes turn on TCO, but messing with stack frames in a systems language is generally a no-no. We've reserved the 'become' keyword for the purpose of explicitly opting into TCO in the future, but we haven't been able to implement it because historically, LLVM had issues on some platforms. In the time since, it's gotten better, and the feature really just needs design to work. Purity isn't as big of a deal in Rust as it is in other languages. We used to have it, but it wasn't very useful. > But assignment in Rust is not a totally trivial topic. Move semantics can be strange from a not-systems background, but they're surprisingly important. We used to differ here, we required two operators for move vs copy, but that wasn't very good, and we used to infer Copy, but that ended up with surprising errors at a distance. Opting into copy semantics ends up the best option. > how that could ever be more useful than returning the newly-assigned rvalue. Returning the rvalue ends up in a universe of tricky errors; not returning the rvalue here ends up being nicer. Furthermore, given something like "let (x, y) = (1, 2)", what is that new rvalue? it's not as clear. > I’ve always thought it should be up to the caller to say which functions they’d like inlined, This is, in fact, the default. You can use the attributes to inform the optimizer of your wishes, if you want more control. > It’s a perfectly valid code, In this case it is, but generally speaking, aliasing &muts leads to problems like iterator invalidation, even in a single-threaded context. > but the online documentation only lists the specific types at their five-layers-deep locations. We have a bug open for this. Turns out, relevant search results is a Hard Problem, in a sense, but also the kind of papercut you can clean up after the language has stable semantics. Lots of work to do in this area, of course. > Rust won’t read C header files, so you have to manually declare each function you want The bindgen tool can help here. > My initial belief was that a function that does something unsafe must, itself, be unsafe This is true for unsafe functions, but not unsafe blocks. If unsafe were truly infectious in this way, all Rust code would be unsafe, and so it wouldn't be a useful feature. Unsafe blocks are intended to be safe to use, you're just verifying the invariants manually, rather than letting the compiler do it. > but until a few days ago, Cargo didn’t understand linker flags, This is not actually true, see http://doc.crates.io/build-script.html http://doc.crates.io/build-script.html for more. > the designers got rid of it (@T) in the interest of simplifying the language This is sort of true, and sort of not. @T and ~T were removed to simplify the language, we didn't want language-support for these two types. @T's replacement type, Gc<T>, was deemed not actually useful in practice, and so was removed, like all non-useful features should be. In the future, we may still end up with a garbage collected type, but Gc<T> was not it. > Rust’s memory is essentially reference-counted at compile-time, rather than run-time, with a constraint that the refcount cannot exceed 1. This is not strictly true, though it's a pretty decent starting point. You may have either 1 -> N references, OR 1 mutable reference at a given time, strictly speaking, at the language level. Library types which use `unsafe` internally can provide more complex structures that give you more complex options. That's at least my initial thoughts. Once again, these kinds of reports are invaluable to us, as it helps us know how we can help people understand Rust better.
- leoc 11y ago> but messing with stack frames in a systems language is generally a no-no. Could you expand on this? Optimising away a stack frame that lies on the border of some security barrier would obviously be Bad News, but what other specific problems are there? Conversely, it seems there are some possible benefits to TCO in a systems language: I'm thinking of those secure-C coding standards which (apparently) tend to ban recursion for fear of stack overflow.
- pcwalton 11y agoLLVM does do sibling call optimization, which allows for TCO in many common cases, including all cases of a function tail calling itself (but note that RAII makes the definition of tail position subtler than it may seem at first glance).
- leoc 11y ago> (but note that RAII makes the definition of tail position subtler than it may seem at first glance) Like http://www.nhplace.com/kent/PFAQ/unwind-protect-vs-continuations-original.html http://www.nhplace.com/kent/PFAQ/unwind-protect-vs-continuat... this?
- masklinn 11y ago> This is, in fact, the default. You can use the attributes to inform the optimizer of your wishes, if you want more control. Isn't the default that the optimiser will do whatever the hell it wants, and the attributes simply skew the optimiser's factors in one direction or another? I think what the author means here is that the caller function should be able to define whether the callee should be inlined or not. > The bindgen tool can help here. Would be really useful to have an implicit bindgen thing. Maybe a compiler plugin using e.g. Clang's C parser? That way there's no need to maintain the binding. I'd say I'd like a header generator more than a reader though.
- steveklabnik 11y agoMaybe I misunderstood what the parent wants, but you're right that the optimizer can do as it pleases, and you can use annotations to help it make the right decision. An 'implicit' tool may in fact be cool. It's not perfect, and so needs tweaking in many cases, so the current state is pretty good, but for easier cases and/or when you don't care, I can see such a thing being useful.
- kibwen 11y agoThe author's observed speed discrepancy between iterators and while loops makes me think that they forgot to compile with optimizations, as the difference between those programs is almost negligible on my end.
- Manishearth 11y agoIIRC there _are_ some possible optimizations that we don't yet do for iterators; but these are something that can easily be added in the future, and not something that really impact design per se.
- Veedrac 11y agoTo add to this, I found that when using i64 both generated exactly the same code.
- GolDDranks 11y agoI always thought that iterators would be _easier_ to optimize (assuming that the compiler can statically access the implementation) because they are more constrained in form at the use site. Loop unrolling etc.
- saosebastiao 11y agoI too have found the assignment semantics to be a little baffling, and the errors to be ungoogleable (which may have changed in the last 4 months since I used it last). A pragma determining semantics seems quite brittle as well. I wish that there were some sort of distinguishing operators for copy vs move, much like how F# has different operators for initial assignment vs mutation.
- steveklabnik 11y agoWe used to have two operators, but it wasn't actually helpful. The only difference between a move and a copy in Rust is that you can use a copy value afterward. Why does this matter? Okay, imagine this code: let v = vec![1, 2, 3]; let v2 = v; Since Vec does not implement Copy, it's a move. Moves memcpy the value on the stack, which, in a Vec's case, is a triple: pointer to the data, a length, and the capacity. You haven't actually copied the data on the heap, just the three pointers on the stack. If we let you use v after the assignment, there'd be two pointers to the same data. No bueno. Compare that with this code: let i = 5; let i2 = i; In this case, 5 is an i32, which implements Copy. The same thing happens: a memcpy. But now, the entire data was copied. There's nothing on the heap. In this case, it's totally okay to keep using i. Does that make sense?
- saosebastiao 11y agoIt makes sense...but only because you explained that Vec doesn't implement copy and i32 does. Without scouring the source/docs, or possibly having a mythical IDE that can discover this for me, I have to rely on error messages. I just checked, and the error messages are definitely better since January, but I think it would be helpful to have an extra operator just for a visual understanding of the code, rather than enhancing its procedural semantics.
- steveklabnik 11y agoThis sort of gets at the root of why using two operators wasn't great: you're basically forcing a human to annotate what a compiler already knows. It doesn't improve the situation significantly, and just adds noise. It's sort of similar to why we don't force you to write out every single type. I'm trying to find the blog post that had an example of the old syntax, but I can't at the moment :/ As you code in Rust, you sort of figure out what should be Copy and what shouldn't: can it be completely copied with a memcpy? Then it's Copy, or it should be. Does it contain a pointer to some sort of external resource? Than it shouldn't be. And, once you've written some Rust, you know what 'use of moved value' is: it means the type moves, not copies. So the error _does_ tell you what's up, although it may not be 100% awesome for completely new people. The error from my first program, trying to print `v` after: hello.rs:5:22: 5:23 error: use of moved value: `v` hello.rs:5 println!("{:?}", v); ^ note: in expansion of format_args! <std macros>:2:25: 2:56 note: expansion site <std macros>:1:1: 2:62 note: in expansion of print! <std macros>:3:1: 3:54 note: expansion site <std macros>:1:1: 3:58 note: in expansion of println! hello.rs:5:5: 5:25 note: expansion site hello.rs:3:9: 3:11 note: `v` moved here because it has type `collections::vec::Vec<i32>`, which is moved by default hello.rs:3 let v2 = v; ^~ hello.rs:3:9: 3:11 help: use `ref` to override `ref` is sort of an awkward suggestion, but it does tell you, it's moved by default.
- pcwalton 11y ago> I’m not sure if the ownership rule is actually helpful in single-threaded contexts, but it at least makes sense in light of Rust’s green-threaded heritage. It's necessary for prevention of use-after-free. Here's a simple example (which can be translated into the equivalent C++): let mut vector = vec![ "1".to_owned(), "2".to_owned(), "3".to_owned(), ]; for element in vector.iter() { vector.clear(); println!("{}", element); }
- eridius 11y agoIt's always interesting to see the experiences of new people to Rust. There's a few curious misconceptions in here that I hadn't seen before: > Iterators are a great and reusable way to encapsulate your blah-blah-blah-blah-blah, but of course they’re impossible to optimize. I'm very curious to know where you got this idea from. Iterators actually optimize very well, in most cases being indistinguishable from a manual imperative for-loop. If you forget to turn on compiler optimizations then you'll see a significant performance difference, but that's true for a lot of different things you might want to do. Compiler optimizations are important whenever you're measuring performance. > pragma It's not a pragma. It uses the same # sigil that C compilers use to introduce pragmas, but in Rust, it actually denotes an attribute which modifies the next item (or in the enclosing item with the #! syntax). This is used for a number of things. As you've seen, it can be used to automatically derive implementations of traits, and it can be used to mark a function as being a candidate for inlining, among other things. > Rust rather inelegantly overloads the assignment operator to mean either binding or copying. This is indeed a very curious misconception. Assignment actually isn't overloaded like that at all. In your code example, when you say let y = x; You're not binding y to the value of x, you're actually moving the value of x into y. There is no implicit bind-by-reference in Rust. If you want a reference, you need to use the & sigil. The confusion here stems from the fact that some values can be copied and some values cannot. When you move a value in Rust, if the value can be copied (if it conforms to the Copy trait), then the original value is still usable after the move. This means you can say let x = 1; let y = x; let z = x; The line `let y = x` moves the value of x into y, but since the value is copyable, it's really just moving a copy of the value. In this context, "copyable" basically means memcpy() can be used to produce a valid copy. There's a different trait called Clone which has an explicit .clone() method that is used for values that require additional work beyond memcpy() to copy. In your code example, your struct is not Copy, so moving its value makes the original value inaccessible. Basically, it's considered garbage memory and cannot be read from again. This is why your code let x = MyValue::Digit(10); let y = x; let z = x; throws an error. It isn't because `y` and `z` would be referring to the same value, it's because after the line `let y = x;` the original value `x` is garbage. So really, any time you do assignment like that (or any time you pass a value as an argument to a function) without taking an explicit reference (using &), you're doing a move. Values that confirm to Copy will move a copy of the value, and all other values will leave the original value inaccessible. This should actually be familiar to people coming from C++, where values that aren't Copy are basically like std::unique_ptr, except that instead of leaving the original value with a known-"zero" state, the compiler prevents you from accessing the original value at all.
- Manishearth 11y agoThis is a great post! :D Some nits on various errors or misrepresentations: > pragma You have some complaints about attributes (what you call pragmas) -- I suspect that many of them are due to you looking at them as if they were C++ preprocessor directives or pragmas. They aren't, even if the syntax may be reminiscent :) In some cases they're like decorators in python (but much more powerful), in others, Java annotations. They're a different concept. > Sadly, Rust is not a target for my favorite parser generator, and the lexers in Servo don’t look much better than C-style state machines, with lots of matching (switching) on character literals. https://github.com/servo/html5ever https://github.com/servo/html5ever is the largest parsing library we use in Servo, and there are a bunch of Rust tricks done there. HTML parsing is hard (the spec is insanely complex), and this library does it well with much less code. > it would be nice if the Rust compiler got rid of the split_at_mut secret password and could reason sanely about slice literals and array indexes. `split_at_mut` is just a library function that uses `unsafe` internally, much like many other core data structures and methods. It has nothing to do with the compiler. > It is also worth noting that processing non-overlapping slices in parallel is destined to come into mortal conflict with The Iterator, which is by its nature sequential. I believe the thread::scoped API can be used to process nonoverlapping slices in parallel with a regular iterator. Not sure, but I recall seeing an example where this was done. > ...seems like an excessive amount of ceremony, at least for a language that keeps using the word “systems” on its website. I don't see what verbosity has to do with systems programming. Most of those are zero or low cost (so verbosity doesn't correspond to more steps in the generated assembly); I believe there's a utf8 check at one point and that's it. Rust is verbose wherever errors or footguns are possible. > Rust won’t read C header files, so you have to manually declare each function you want to call by wrapping it in an extern block, like this: https://github.com/crabtw/rust-bindgen https://github.com/crabtw/rust-bindgen > but until a few days ago, Cargo didn’t understand linker flags You can specify -L and -l flags in the .cargo/config file under the rust-flags key (or output the same in a build script); this has been allowed for a while now. More complex linker args are now possible with the cargo rustc command.
- Jweb_Guru 11y ago> I believe the thread::scoped API can be used to process nonoverlapping slices in parallel with a regular iterator. Not sure, but I recall seeing an example where this was done. Yes, it can. And while the original API had to be reworked a bit, this capability will not change (and someone can probably implemented the proposed solution in a library today, if they so choose, so people can try it in the beta). I really enjoyed reading the article, and I think it presented its argument well, but I fear that some of the conclusions it came to were based on technical misconceptions about the language. It's also making me feel more and more like turning optimizations on by default may be a good idea... for every person who writes a blog post or asks a question on StackOverflow about why code isn't running fast enough, there are probably ten others who just assumed Rust wasn't ready and moved on to other things.
- chaoky 11y agothe article brings up a good point with 'systems language'. What is a systems language anyways? I guess C is, but whats the definition? Is common lisp a 'systems language'? After all, a good number of operating systems have been written in common lisp, but is it too freewheeling and high level to be considered a 'systems language'? Is java a systems language?
- parley 11y agoIt is an interesting question and one that's quite debatable. I enjoyed the discussion in [0] the panel "Systems Programming in 2014 and Beyond" from 2014 with panel members Bjarne Stroustrup (C++), Niko Matsakis (Rust), Andrei Alexandrescu (D) and Rob Pike (Go). Have to say I agree with Bjarne and Niko on most points discussed. [0] https://channel9.msdn.com/Events/Lang-NEXT/Lang-NEXT-2014/Panel-Systems-Programming-Languages-in-2014-and-Beyond https://channel9.msdn.com/Events/Lang-NEXT/Lang-NEXT-2014/Pa...
- acomjean 11y agoMy understanding is that 'systems languages' run without a runtime and are therefore good for writing low level programs for embedded and operating systems. I could be wrong...
- Manishearth 11y agoPedantry: Even C++ has a runtime (so does Rust); they're just tiny and I think can be disabled.
- Retra 11y agoImplicitly, the most important feature of a 'systems language' is predictability: how easily can you predict how much memory or time a program will use when running?
- dllthomas 11y agoIf we take that seriously, I wonder if we should look at something that isn't turing complete.
- ufo 11y agoThe blog author mentions at one point that Algebraic Data Types cannot determine pattern exaustiveness for things like match 100 { y if y > 0 => “positive”, y if y == 0 => “zero”, y if y < 0 => “negative”, }; and wonders if there is some kind of "algebraic data values" to notice that the previous case is exaustive. In Haskell they solve this particular problem with an "Ordering" ADT: data Ordering = LT | EQ | GT case (compare 0 100) of LT -> "Positive" EQ -> "Equal" GT -> "Negative" Using a richer datatype instead of boolean predicates solves most problems. For some more advanced things you need extensions to the type systems. For example, to be able to say "this function, when applied to non-empty lists returns non-empty lists", you need Generalized Algebraic Data Types (gadts).
- pcwalton 11y agoSame in Rust: http://doc.rust-lang.org/std/cmp/enum.Ordering.html http://doc.rust-lang.org/std/cmp/enum.Ordering.html
- Veedrac 11y agomatch 100.cmp(&0) { Ordering::Greater => "positive", Ordering::Equal => "zero", Ordering::Less => "negative", }
- steveklabnik 11y agoTo elaborate a bit further, the reason that the original isn't exhaustive is that to know that, you would have to execute arbitrary code at compile time, the 'if'. The exhaustiveness check is only on the `y` part of the pattern.
- viraptor 11y agoIsn't this possible for some limited cases? I mean, if you only match on the integer value, this case is trivial. Similar cases are already analysed by gcc to tell you "this comparison is always true/false". I know it doesn't generalise and it "is it worth doing" may be a question, but it seems like most match blocks would fall easily under a very limited set of known cases.
- Veedrac 11y ago> but five function calls (from_ptr, from_utf8, to_bytes, unwrap, to_string) just to convert a C string to a Rust string seems like an excessive amount of ceremony Well, only the first four are really needed. The last one turns it into a local, growable copy. The rest should probably get a convenience wrapper, though. > the libxls API has an array-of-structs in a few places [...]; because Rust doesn’t believe in pointer arithmetic, I found myself manually writing the pointer arithmetic logic Forgive me if I'm being stupid, but can't you just use slice.from_raw_parts? https://doc.rust-lang.org/std/slice/fn.from_raw_parts.html https://doc.rust-lang.org/std/slice/fn.from_raw_parts.html > writing things in a pseudo-functional, lots-of-chained-method-calls style, for which Rust is not all that well-designed If this is about speed, then see the other comments. But if this is for another reason, what reason is it? Personally this style of programming suits Rust beautifully.
- Dewie3 11y ago> The pattern list looks pretty exhaustive to me, but Rust wouldn’t know it. I’m that sure someone who is versed in type theory will send me an email explain how what I want is impossible unless P=NP, or something like that, but all I’m saying is, it’d be a nice feature to have. I Am Not A Type Theorist, but that looks like it could be very hard for a compiler to deduce in general. You might have to bust out a proof yourself for things like this. In which case you might not feel it is worth it for a "nice to have".
- imron 11y ago> But then, this won’t compile: let x = if something > 0 { 2 }; Which makes perfect sense. After all, what should the value of x be if something is <= 0?
- erkl 11y agoNow, I'm not sure this is useful at all, but I think it would make sense for the value of an if-expression to be Option<T>.
- Jweb_Guru 11y agoThat might actually make sense... and I have to say that it would be useful, too. Quite often I find myself doing if condition { Some(foo) } else { None }; being able to just write if condition { foo } could be neat syntactic sugar for that (though it might also be confusing, since in Rust generally types don't form magically like that). The solution I'd come up with was just to give booleans a .then method (maybe they already have one that i missed).
- noelwelsh 11y agoFrom reading the article I came to the conclusion the author is rather inexperienced with modern functional languages. Their claims about tension in the design of Rust between functional and imperative features for me mostly come down to them not understanding natural consequences of the modern statically typed paradigm, or Rust's memory model. For instance, take the discussion about if expressions. They claim an if expression with a single arm returning unit is unintuitive. Firstly, we should always recognise that claims about "intuitive" behaviour solely depend on one's background. This behaviour is completely intuitive to me, coming from a Racket background (which behaves the same way). It's also a natural consequence of having to give a type to an if expression with one arm. You have two choices: the else case either returns the bottom type (which operationally means it raises some kind of error), or it returns no interesting value -- which is exactly what unit is. Since a single arm if can only be used for effect, unit is the best choice here. Rust is certainly a very different language to OO-ish imperative languages that most people are familiar with. I think the author made the common mistake of expecting Rust to behave like one of these languages, and then blaming the discrepancies between their mental model and actual behaviour on the language.
- GolDDranks 11y agoOn the other hand, an alternative design would be that the single-arm if wrapped the return type R to "Some(R)", and None for the "phantom else" arm. I haven't considered the ramifications of this, but I'd expect it to work well if the Option<R> type is defined flexibly and compose-ably enough.