4 ms·
Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster
- gigatexal 23d agoBut isn’t the enum far more readable and maintainable than having to do bit operations on things?
- tom_ 23d agoThe computer's the one running the code, and it'll be running it a lot (or so its author hopes), so it's probably worth bearing its limitations in mind in the interests of making its life easier (so to speak) rather than prioritising the people who will modify the interpreter - a far less common occurrence. The bit operations involved are pretty simple and won't take you long to figure out even if you've never done them before.
- steveklabnik 23d ago> As you can see, it has many convenience methods to make it easy to work with, compensating for the loss of the Rust enum. Also, Rust does try to do some of these optimizations itself. These aren't exposed in the stable language to let you do some more advanced things, but it wouldn't be impossible for you to get the best of both worlds by letting you communicate this stuff more directly to the compiler. Right now those things are more like "this value is where you should put the tag" than the more advanced stuff here, though. Would be cool to see someday!
- lowbloodsugar 23d agoquibble: unsafe is stable. you can't do this in safe rust, but you can do it in unsafe rust. just isolate all the unsafe code in a single type, ideally a tiny crate.
- tialaramex 23d agoSteve wasn't talking about unsafe. He was talking about being able to mint your own non-enum types with user defined niches. Rust provides for example NonZeroU8 which is an 8-bit unsigned integer that's never zero, leaving it with 255 possible values and a convenient niche. You cannot make one of these yourself directly, because the mechanism used by Rust itself is a deliberately perma-unstable compiler-only proc macro which says "Hey compiler, I promise I only ever use bit patterns 0x01 through 0xFF inclusive". Today you can either - hide a NonZero type inside your type and use that to get the niche, or, use an enum itself which automatically knows ever pattern it didn't use is a niche. In the future a hypothetical "Pattern Types" feature would let you make such types yourself as easily as Rust does Personally I would like to make a Balanced set of types, like BalanacedI8 (the 8-bit integers except the most negative, so -127 to +127 inclusive) because I think lots of people have a use for types like i8 or i32 but don't need their unbalanaced most-negative value and could re-purpose it this way. And you can make such types... indeed I have... but it's only really practical in unstable Rust.
- fpoling 23d agoArbitrary subranges of int types were available in Ada for over 40 years via static enforcement via Spark and compilers were able to optimize them nicely. For a system language I wish Rust would support such things rather than coming with NonZero hacks.
- steveklabnik 23d agoThis is called "pattern types" in Rust land, as your parent mentioned, and is exactly the kind of work being talked about. NonZero isn't a hack: it's an example of a common pattern. If pattern types were available today, you'd still want NonZero, as an example of a pretty standard pattern. The idea is, as always: prove out the specific version, then generalize.
- BoingBoomTschak 23d agoCan it do holes like this? (typep 3 '(or (integer 0 10) (integer 50 100)));; => T
- compiler-guy 23d agoThe reason the enum is so nice is that it works as a terrific language-supplied abstraction that covers up those bit manipulations. It's very nice to get those abstractions for free like you do in Rust, but it can't be optimal for every specialized use case. This new code also supplies similar abstractions. That actual specific code is much harder to reason about, but most users--and even the next person who works on the interpreter--simply won't care, or even know what is going on underneath the hood. The abstractions provided by the author do that work and apparently do it cleanly. For most use-cases, that extra hand-written code isn't worth it. But in specific cases it can be, and the author has actually measured the value and determined that it is.
- mwkaufma 23d agoIf you add "fits in a register" to your list of correctness requirements, then it's no-go even if the source has less cognitive overhead.
- diath 23d agoHighly optimized code in hot paths is rarely readable.
- locknitpicker 23d ago> Highly optimized code in hot paths is rarely readable. ...unless it's supported by the language as a first class feature. See for example C++ and RVO.
- diath 23d agoNot necessarily, in C++ you'd still drop from smart pointers to raw pointers, from virtual dispatch to switches/computed gotos, from std::function to function pointers and so on. These abstractions all come at a cost.
- dzaima 23d agoUnless the highly-optimized parts are wrapped by an interface that looks similar to the non-optimized version. In the case of a tagged object in Rust, depending on how well the compiler can wrangle through it, you might even be able to add a `.unpack()` method that returns a pretty enum from a packed value, that you can pattern-match on or whatever, and let the compiler remove all the code of unpacking unused cases. (using that directly for the addition example would end up less efficient of course, but still most likely beneficial. It's after this when there's a potential true readability vs performance tradeoff)
- wat10000 23d agoIt can be worth sacrificing readability and maintainability for better performance in hot code.
- lowbloodsugar 23d agoTake a look at triomphe's ArcUnion and extrapolate from there. Basically make a crate for just your 64bit union type, do it unsafe there, test with miri, and now you have a safe 64bit type you can use with match. You're happy digging around assembly so this is well within your wheelhouse. The only challenge will be if you do use miri to verify then you need to use the 'provenance-preserving' pointer adjusting functions. Worth the learning experience in my opinion. I did one for my system and it was super fun and had the performance impact you describe.
- krick 23d agoThat's very unpleasant to hear. It's sad to be reminded that Rust compiler is not magic and cannot just... do these things somehow. Sure, all abstractions do have some cost, but, man, 17% performance gain by virtue of replacing enum with this monstrosity? That's very annoying.
- maplant 23d agoIt can't do these things because it's not wanted. Say you have the following: enum Value { Float(f64), Ptr(*const T), } Do you want the compiler to disallow certain bit patterns in the Float variant simply so that it can implement NanBoxing?
- vlovich123 23d agoProbably with an annotation around a NanBoxable(f64) type that tells it to do that. That being said, the optimization is complex that may be insufficient: > For my boxing scheme, I picked a bias value such that the lowest two bits end up being 10. That 1 in bit index 1 indicates that doubles can't be directly compared for equality. Amazingly, we only lose two bits of exponent, and we keep the full precision of the mantissa, meaning we lose no significant digits in the flonum representation. This suggests the optimization needs more information about specifically how you want to box the float. There probably is some primitives worth considering standardizing to make this kind of optimization possible so that the tunable parameters are passed as const generic values.
- gpm 23d agoOr actually doing what they did here where it changes it to a enum Value { SimpleFloat(64bit value), ComplexNan(Heap pointer) Ptr(*const T) } You lose out on performance if you use the bit patterns in the float that most people don't use very much, but you keep the correctness. I'd be very unhappy if a compiler silently did this to me - it would make performance extremely hard to reason about. But it's not quite as bad as changing the semantics.
- speedstyle 23d agoA 64-bit sum type can't magically combine an i64, f64, and several raw pointers, each of which carry a full 64 bits themselves. You have to change the semantics of the code. Some semantics could be expressed more easily with compiler improvements, allowing eg `Aligned<T>` like `NonNull<T>`, or `FiniteF64` like `NonZeroU64`, or even `#[range(0..1<<60)] u64`, but you still couldn't overlap two `Aligned`s in one enum, because only one can be stored unchanged, the others need masking off before usage. Even if the enum semantics allowed this, I'm not sure the compiler should do this kind of compute/memory tradeoff automagically. Which doesn't mean you can't write nice abstractions over it, there's a few tagged ptr crates which aim to do it for you
- fpoling 23d agoThe article title is misleading. It is not that Rust compiler was not able to optimize some low-level operations. Rather the author came up with encoding schema that fit most things the interpreter dealt with into 64 bit. This replaced the previous schema that used 128 bit for everything but that can be directly mapped into Rust enums. The catch was that it was necessary to allocate some things on the heap and use pointer indirection but that was used for rare values so on average the new schema provided nice win. One cannot expect a compiler to come up with such encoding.
- win311fwg 23d agoWhat is misleading about the title? A custom encoding scheme is exactly what it suggests. Maybe it has been edited since your comment was posted?
- dymk 23d agoIt wasn’t replacing one rust enum, it was replacing what are effectively multiple enums
- dzaima 23d agoHow so? It's replacing multiple enum variants, but just one enum, "enum Value". (also; if anything, the title is implying the exact opposite of "Rust compiler was able to optimize ...", "Replacing a Rust [...] with [...]" is clearly moving away from Rust-magic to something else)
- Brian_K_White 23d agoProbably in the sense that you can remove the word rust and nothing changes. It's not about some failure of rust to be efficient at enums, but the title says it is.
- win311fwg 23d ago'Enum' is ill-defined so the addition of Rust is significant as it indicates what one can expect with how data is structured. There is nothing in that speaks to the Rust compiler or Rust being inefficient or anything of the sort. It remains unclear where this idea is coming from. There is nothing in title that would send you there. Unless, again, the title was edited at some point?
- nwhitehead 23d ago"the smart thing to do is to give integers zeros as their tag bits, because then, adding or subtracting two shifted integers remains a plain add or sub machine instruction" this is brilliant, love it. stealing this idea immediately.
- Ozzie-D 22d ago[flagged]
- MindSpunk 22d agoI'm not convinced the performance benefits are entirely the result of the more compact object representation. It definitely would help, but looking at the code snippets the author provides for the add instruction there's an important structural change that would be making a huge difference. The old, enum based value type used a single big match statement to dispatch between all possible type combinations. Their assembler output looks like the match gets compiled to something like a big stack of nested if statements. The new code uses an explicit fast path check with a dispatch into a tagged 'cold' path when the common case isn't hit. The generated code is a single upfront branch for the fast path that exits immediately, with a dispatch into the slow path in a separate function. This would be contributing significantly to the performance improvements. The old path requires taking several branches even on the hot path. The new code has a single, highly predictable branch that skips all the messy dispatch for the other types. This could have been implemented for the enum based value type, and I would expect to see a jump in performance there too even without the new compact value type. There will be a much higher branch predictor hit rate with the explicit fast path.
- orielhaim 22d ago[dead]
- vbezhenar 22d agoBut CPU branch predictor should have figured out hot paths in the original implementation?
- maxime_cb 22d agoAuthor here. The disassembly for the old enum handling had many spills, simply because the old value enum can't fit in a single register. If you have an instruction that two operands with two of those big value enums, it needs 4 registers instead of 2. That, coupled with better cache-friendliness, explains a lot.
- whizzter 22d ago100% agreed (see my sibling comment), this all accumulates for in-language function calls,etc since code at runtime often spends a surprising amount of time just moving around values instead of doing useful work, having them just as singular register values really helps a ton.