4 ms·
is monomorphization a more concise term for what is called "implicit template instantiation" in C++? If so, is Rust worse than C++ when it comes to compile tim
by vmarsy 8y ago
is monomorphization a more concise term for what is called "implicit template instantiation" in C++?
If so, is Rust worse than C++ when it comes to compile times of those? Heavily templated C++ code is known to be slow to compile due to this. Is it much worse in Rust? Are you aware whether C++ compilers do some special optimizations that Rust compiler doesn't have yet?
I don't see what can be done to make this better, the whole point of templates is to "pay" at compile time for some gains at runtime. (or in this case for gains in the amount of code a developer has to write)
- steveklabnik 8y ago“Monomorphization” is what happens when a template is instantiated. It’s essentially the same thing in both languages, and has roughly the same compiler performance implications. There’s at least one thing that could be done, in Rust at least. But it’s only in some limited circumstances. See here: http://www.suspectsemantics.com/blog/2016/12/03/monomorphization-bloat/ http://www.suspectsemantics.com/blog/2016/12/03/monomorphiza...
- vmarsy 8y agoThanks, that's an interesting read. indeed a developer can use that trick "in the case of conversion traits". Monomorphization seems to be the source of 2 major problems: compile time, and binary size. An optimization that could reduce the compile time by caching the generated code (at least when compiling over-and-over the same code base. i.e. "Incremental compilation") - and it seems that Rust already is doing something with it [1]. I wonder if something specific is done for template instantiations there. Optimizations aimed at reducing binary size seems much more tricky, if not impossible, except as you pointed out in the limited cases described above. In the regular cases, templates are kind of working as intended: the developer has to think if he/she would have written the same amount of code N times if templates were non existent. [1] https://blog.rust-lang.org/2016/09/08/incremental.html https://blog.rust-lang.org/2016/09/08/incremental.html
- strmpnk 8y agoThe translation units in C++ tend to pay this cost a number of times over because of the way header files are processed. With modules, this theoretically could start getting closer to Rust, but for large projects, C++ templates can still be orders of magnitude worse in terms of instantiation cost. This is partly why there are hacks like unity builds (not related to the engine, where all source is bundled into a single translation unit). These have plenty of drawbacks too so it's not a clear win. Adding to all of this, there are fancier mechanisms for template meta-programming like SFINAE rules + computed template values. Sure, it's "turing complete" but this is why we see such clever libraries with huge explosions in code generation size. I'm far from an expert in modern C++ features but it's clear that there is an entire interpreted programming language of templates bolted on the to the rest of C++. It reminds me of this post on a similar take on Haskell type level programming: https://aphyr.com/posts/342-typing-the-technical-interview https://aphyr.com/posts/342-typing-the-technical-interview (or similar feats by Oleg Kiselyov).
- jules 8y agoThe compiler could do that automatically. Suppose you're operating on the control flow graph and you are to specialise a basic block in some type environment. You can cache the result so that the next time the same basic block is specialised in the same type environment you look up the result in the cache and create a call to that existing basic block. pub fn big_function<T: Into<i32>>(x: T) { let y: i32 = x.into(); ... code that uses y but not x .... } In the first line of the function you have the environment {T: Into<i32>, x: T}. After the let you have the environment {T: Into<i32>, x: T, y: i32}. If you keyed the cache on the full type environment you wouldn't solve the issue because the code that only uses y would still get specialised to the {T: Into<i32>, x: T} too. However, you could detect that that code doesn't actually use T and x, so that you can specialise it to the type environment {y: i32} only. That detection can happen as a side effect of specialising that code to some particular {T: Into<i32>, x: T, y: i32} for the first time. As you specialise that code you record which parts of the type environment actually got used, and you use only those parts as a key in the cache. The type environment object itself could take care of recording what the compiler looked up in it. Another advantage of doing it this way, rather than analysing the code ahead of time, is that it can handle cases where a particular type variable does or doesn't get used depending on what type some other type variable is instantiated to. You could also use the same system to avoid duplicating code that only relies on particular aspects of a type. For instance, a function that permutes the values in a &mut[T] might only care about the size of T and not about the precise type T, so that all its specialisations to T of 4 byte size can call into the same code. Another thing you'd probably want to do is integrate a basic form of constant propagation & dead code elimination during monomorphisation, so that you don't spend a lot of time monomorphising code that ends up dead for particular type instantiations.
- steveklabnik 8y ago> The compiler could do that automatically Yes, that's what I'm saying. > Another thing you'd probably want to do This is a very interesting idea!