3 ms·
10 years ago, I commented on the Rust issue for "Incremental recompilation", where it was suggested that Rust could at least adopt Haskell GHC's model of increm
by nh2 2mo ago
10 years ago, I commented on the Rust issue for "Incremental recompilation", where it was suggested that Rust could at least adopt Haskell GHC's model of incrementality, which is currently file-level:
https://github.com/rust-lang/rust/issues/2369#issuecomment-128857193 https://github.com/rust-lang/rust/issues/2369#issuecomment-1...
This would already help a lot.
I recommend anybody who's interested in incremental recompilation to read what GHC does, because the effort to achieve that is relatively low.
Of course there's always desire for more:
GHC currently needs to parse+typecheck+codegen a file before it can process other files that import it. Codegen is slow. Thus, there's currently demand split compilation into "stages", so that the next file can be typechecked after its imports have been just typechecked (not codegenned).
I would also enjoy if recompilation avoidance were to happen at the function level, not the file level.
Macro systems are a key language feature that can destroy incremental recompilation. In theory, Haskell is well set up for that, as its macro system (TemplateHaskell) is fully AST based and _theoretically_ could distinguish "fully pure" macros from side-effectful macros (such as splicing the current git commit in as a string literal). But the recompilation avoidance system does not currently exploit such differences.
- steveklabnik 2mo agoJust to be clear about it, Rust today does do some amount of incremental compilation, and there is more work being done to continue to make it moreso. It's just very difficult to re-architect such a large and heavily used codebase. People are putting in heroic amounts of effort to improve things. An example that's being funded right now: https://rust-lang.github.io/rust-project-goals/2026/expansion-time-evaluation.html https://rust-lang.github.io/rust-project-goals/2026/expansio...
- amelius 2mo agoCan't we have a system where we trade some performance for quick incremental compilation? We can always compile with full optimization just before shipping?
- pjmlp 2mo agoOf course we can, C++ even REPL and hot reloading tools. The main issue is that so far such tools haven't been a priority for Rust.
- steveklabnik 2mo agoThe article gestures at (and the author has made a comment in this thread about) how this is the case for Zig. You are right that there is tension here, and so that's exactly what you do: accept less performance for the gains in incremental, and then don't do incremental for final builds. It's a fine way to go about it, assuming that the lack of performance doesn't make the program unusuable. (Some people add some basic optimizations to their Rust debug builds, for example, because no optimizations is too painful to actually use.)
- SkiFire13 2mo ago> Haskell GHC's model of incrementality, which is currently file-level From what I see in Haskell files are the unit of compilation, and that's what allows incremental compilation to be file based (because it's really unit-of-compilation based) I can see you can have circular dependencies between files with the `{-# SOURCE #-}` pragma, but I don't see documentation about how that affects incremental compilation. A couple of issues I see with doing this in Rust are: - in Rust the unit of compilaion is a crate, which can contain many fils/modules with circular imports, which is much more coarser than what can be done in Haskell. - in Rust downstream crates can depend on function bodies upstream for running compile time functions; as such the crate/module interface is not enough to gate recompilation, but at the same time including all function bodies will also not give the wanted benefits. This is solvable but likely requires more work than what was done in Haskell. In general you cannot take a language approach and blanket applying it to another one without considering their different quirks, which is likely why your proposal didn't get much attention. Or am I missing something that would make it easier to apply Haskell approach here?
- panstromek 2mo ago> GHC currently needs to parse+typecheck+codegen a file before it can process other files that import it. Codegen is slow. Thus, there's currently demand split compilation into "stages", so that the next file can be typechecked after its imports have been just typechecked (not codegenned). > I would also enjoy if recompilation avoidance were to happen at the function level, not the file level. This sounds like Rust is already doing a lot more incremental then GHC then. Rustc only needs to parse, expand macros and do name resolution. Everything else is incremental after that, on a very granular level.
- nh2 2mo agoCan you point at what you mean? If you changed a function implementation in the libc crate, not changing that function's signature, how much codegen would happen in downstream packages? The maximally recompilation-avoiding effect would be: Only that one function gets codegenned. Everything else just gets relinked into their final executable or .so.
- panstromek 2mo agoThis is difficult to answer, because these systems exhibit somewhat chaotic behaviour, and it gets even more complicated cross-crate. I was mostly talking about single crate scenario. Since you mentioned libc, the likely answer is that nothing gets codegened in downstream crates. But this is only because libc functions are usually not generic or `#[inline]`. Changes to generic or inline functions can dirty downstream codegen units where the function was called. Inside `libc`, the change will trigger recompilation of at least one codegen unit, depending on how the function is used inside libc itself. Single crate is split into 256 units in incremental mode. Nevertheless, even if the codegen is needed just for the `libc` crate, `libc` will dirty its metadata, which means that downstream crates will still need to recompile the initial steps before incremental kicks in (which is roughly parsing, macro expansion and name resolution). After that, the query system just returns cached results for all the subsequent steps. There's some work going towards skipping the rustc invocation altogether in those cases (usually referred to as "Relink don't Rebuild" proposal), because even just loading the dependency graph and figuring out that you don't have to do anything can take quite bit of time for larger programs.