4 ms·
Indeed, it is a fundamentally flawed principle. Code needs to be considered as a graph. All indirections within the code create an edge. Linear blocks of atomic
by robert_tweed 7y ago
Indeed, it is a fundamentally flawed principle. Code needs to be considered as a graph. All indirections within the code create an edge. Linear blocks of atomic instructions make up a node.
The first fallacy is only considering the size of the nodes and ignoring the number of edges. This is like micro-optimising without paying attention to the Big-O factor.
The second fallacy is failing to consider the difference between what I will call static branches and runtime branches. An if statement or a for loop are static branches. Other forms of indirection happen at runtime (with "goto" being famously difficult to analyse). Static indirection is easy to analyse, runtime indirection is hard.
The code graph (I use this term to mean something slightly simpler than the call-graph, which includes things like recursion) tends to get considerably more complex at runtime. One of the best ways to make your runtime code graph more complex is to add more functions, especially ones that are not private to the current file.
As soon as this happens, you have no immediate way of knowing how many runtime edges there will be going in and out of each one of those functions. So hypothetically, let's say you take a 21-line function and break it into 3, 7-line functions. While each of those is easy to "read", you still have to understand 21 lines of interdependent code, but now the total complexity has potentially increased exponentially.
How bad this problem is depends on the scope of the functions (can dramatically increase how many call-sites can potentially form new runtime edges - something you cannot tell from looking at the code in front of you) and whether that code has any side-effects. Certain types of runtime indirection like DI frameworks can completely thwart attempts to statically analyse this runtime graph.
If you are only dealing with pure functions then it can seem like this isn't a problem because of referential transparency. However, this ignores two problems: the name of the function must precisely describe what it does, otherwise it will mislead "distant code readers" (people reading the call but not the source). That must also never change in the future, otherwise you will still have to deal with all that extra complexity if you ever want to refactor safely.
This is not to say that abstractions are bad or that long functions are good. Premature abstraction however tends to produce bad abstractions, which are worse by far than no abstraction.