4 ms·
On the other hand, function calls in hot code are expensive.
by cactusface 11y ago
On the other hand, function calls in hot code are expensive.
- michaelfeathers 11y agoPareto Principle: don't treat all of your code as if it is "hot." It isn't.
- cactusface 11y agoOr the law of opposites: you can't have hot without cold. Then again, you can also judge something as cold when it's hot, for instance a big loop inside a function that gets called once, and you're only profiling at the level of function call counts.
- eru 11y agoOnly if your compiler ain't smart enough. Yes, for some compilers and interpreters performance considerations will force you to write less than readable code.
- cactusface 11y agoInlining is great but it's generally finnicky and hard to control so when it comes to hot loops that matter you may as well do it by hand.
- pubby 11y agoMost compilers have a way to explicitly force functions to be inlined. "Just inline it by hand" is terrible advice.
- cactusface 11y agoI'd be very interested to see a compiler where you can force inlining. How does it handle cycles in the call graph? How does it handle code bloat because the inline keyword was used too often? In GCC at least, you can only give suggestions. The following options control the inliner: -- max-inline-insns-single: Several parameters control the tree inliner used in gcc. This number sets the maximum number of instructions (counted in gcc's internal representation) in a single function that the tree inliner will consider for inlining. This only affects functions declared inline and methods implemented in a class declaration (C++). The default value is 300. max-inline-insns-auto: When you use -finline-functions (included in -O3), a lot of functions that would otherwise not be considered for inlining by the compiler will be investigated. To those functions, a different (more restrictive) limit compared to functions declared inline can be applied. The default value is 300. max-inline-insns: The tree inliner does decrease the allowable size for single functions to be inlined after we already inlined the number of instructions given here by repeated inlining. This number should be a factor of two or more larger than the single function limit. Higher numbers result in better runtime performance, but incur higher compile-time resource (CPU time, memory) requirements and result in larger binaries. Very high values are not advisable, as too large binaries may adversely affect runtime performance. The default value is 600. max-inline-slope: After exceeding the maximum number of inlined instructions by repeated inlining, a linear function is used to decrease the allowable size for single functions. The slope of that function is the negative reciprocal of the number specified here. The default value is 32. min-inline-insns: The repeated inlining is throttled more and more by the linear function after exceeding the limit. To avoid too much throttling, a minimum for this function is specified here to allow repeated inlining for very small functions even when a lot of repeated inlining already has been done. The default value is 130. max-inline-insns-rtl: For languages that use the RTL inliner (this happens at a later stage than tree inlining), you can set the maximum allowable size (counted in RTL instructions) for the RTL inliner with this parameter. The default value is 600. -- So, while the inliner is great, if performance really matters for a hot loop, and you definitely want that code to be inlined, you might end up having to do it by hand. Because otherwise, if you change something somewhere else in your program, the inliner might suddenly decide not to inline your function anymore. You could say the same thing about virtual functions: devirtualization is great but it's not guaranteed. Or loop unrolling. Or codegen. Or whatever.
- plorkyeran 11y agoGCC and clang have __attribute__((always_inline)). VC++ has __forceinline. All of these will always inline unless it hits a recursive call. They do not attempt to avoid code bloat from the programmer telling it to do stupid things.
- cactusface 11y agoOk, I forgot about always_inline. That is a good point. Anyway, it doesn't work for polymorphic (technically non-static) or recursive code. If you've split your program up into lots of pretty little functions, which is fine, you probably have a bunch of polymorphic and recursive code. Typically you get rid of those things by inlining and/or cloning. (Recursive algorithms use the call stack as an implicit data structure, which you can get rid of by introducing an explicit stack locally.) So, to force the compiler to inline some stuff, you have to inline some other stuff. Which is also fine, just that compilers are not magic. For really important things, you need to look at the generated output and decide if you can do better.
- JadeNB 11y agoOn the third hand, premature optimisation is the root of all evil. :-) It's a lot easier to convert well chunked, modular functions into monolithic blocks than it is to go the other way around.
- taeric 11y agoThere was just recently an article about how there is merit to the idea that it is easier to chunk up a monolith than it is to reason about poorly done microservices. Functions, when used in excess, seem to fit the same pattern. Now... good luck getting me to define excess. :)
- JadeNB 11y agoThe HN discussion of the article (MonolithFirst) was https://news.ycombinator.com/item?id=9652893 https://news.ycombinator.com/item?id=9652893 , and it's a good point to bring up. I think that that article describes a reaction against a too-hasty move towards microservices, and (as someone who's not a businessperson myself!) as such it seems well advised; but I suspect that there is, among programmers at large, no too-hasty move towards 'microfunctions', and so no corresponding need to urge people to make the kind of monolithic monstrosities that novice (and even sometimes experienced) programmers are all too willing to make anyway.