3 ms·
When you ("you" in this comment is probably most usefully read as "a compiler" :) ) monomorphize a generic function or data structure with the various type para
by mcronce 4y ago
When you ("you" in this comment is probably most usefully read as "a compiler" :) ) monomorphize a generic function or data structure with the various type parameters that it's been called with, you make a copy of the code for every set of type parameters it's been called with.
Each of these copies requires space for the instructions...this means a larger binary on disk, more memory required when it runs, and more pressure on the CPU's caches, especially precious L1I.
The cache pressure issue is typically the big one worth thinking about, although binary size itself can be an issue for some cases - I've primarily got embedded use cases in mind, but I'm sure there are others I'm not thinking of.
Anyway, back to cache pressure. If there are a whole bunch of different monomorphizations of a given function, and they're all called frequently, that could mean that the CPU will frequently need to refer to slower L2$, or much slower L3$ (or much much slower RAM) to load instructions. That's no bueno from a performance standpoint.
Because of this, there are cases when dynamic dispatch can outperform monomorphization - it's definitely not as simple as "vtable slower, monomorphization faster" across the board, even if that is an OK rule of thumb.
- ayende 4y agoThe linker can usually merge identical functions, and the compiler does this too. You'll only get different copies if they are relevant, and usually the branch elimination is more than enough to compensate. Another aspect about this is that the cache will hold the _hot_ stuff. In most systems, even if you have three copies of a function, it will likely only have one of those that is hot.