9 ms·
Measuring Software Complexity: What Metrics to Use?
- exdsq 5y agoMy first job as a dev was for a consultancy and I had to review a code base for a client to support their argument that they require a rewrite. I had no idea about any of this so googled metrics, found cyclomatic complexity, and wrote a load of bullshit about the results of analyzing the code base showing it was complex. It served its purpose - they got corporate to accept a rewrite - but I’ve never used those metrics again.
- cjfd 5y agoThis talks about code complexity a lot. This, however, is not the chief source of complexity in many code bases. The number of tools needed to build things is the bane of the modern software developer. Also, the use of microservices where none are needed results in and enormous increase in complexity. Code complexity is relatively easy to fix compared to all of this.
- deleted 5y ago[deleted]
- brodo 5y agoExactly. Every dependency is a liability.
- defanor 5y agoIndeed; I tried to compose yet another list of factors contributing to (or approaches to estimating) software project complexity in the past, with code complexity being just one out of 18, and not seeing it as particularly outstanding. Perhaps just a bad title.
- nine_zeros 5y agoYup. Long gone are the days of concise packaging in single executables. Quite literally, software engineers today spend most of their time fighting dependencies and poorly built delivery machines.
- kello 5y agosecond this. toolchain complexity is way more of a PITA than code complexity for me these days.
- kitd 5y agoToolchain simplicity is a key reason for why I like Go. One binary to do it all.
- epylar 5y agoI have seen a team be forced to use microservices for things still running on the same machine years later, with absolutely no benefit other than the powerpoint architecture slides looking fancier.
- worik 5y agoThat may have well got them extra sales, which gave the team pay rises....
- fivea 5y ago> I have seen a team be forced to use microservices for things still running on the same machine years later, with absolutely no benefit other than the powerpoint architecture slides looking fancier. The main driving force for microservices is not technical but organizational. Therefore, there are plenty of non-technical issues, such as for example budget to grow and split a team, that make or break the adoption of this style of architecture.
- epylar 5y agoMaybe. These particular ones were never properly decoupled.
- throwawaybutwhy 5y agoNo mention of function points.
- yetanother-1 5y agoI agree. Certains functions are way more complicated and hard to show value for than others, which is really hard to comprehend for some managers. Functionalities like undo/redo take a lot of planning, coordination and integration efforts than others like a simple export function, but good luck selling that to any marketing or product owner for xx manndays. I still think that this subject is way too specific to be generalized like this, but general thumb of rules still apply like good estimation and technical planning.
- marcosdumay 5y agoFunction points measure the complexity of the problem you are solving, not of the software you are writing.
- sam_bristow 5y agoThe last section on coupling reminded me of the concept of connascence[1] which I've found really helpful when talking about code. [1] https://connascence.io https://connascence.io
- vinodkd 5y agoThanks for this link. TIL. Also, the videos linked from the site were useful
- unbanned 5y agoIf it looks ugly, or you have to navigate too much (long files or a lot of external dependencies/dependency chains)... it's complex. That's all you need to know really
- flaratt_ljos 5y agoThere are a lot of things you can do to lower these sort of metrics without addressing actual complexity, sweeping the problem under the rug. Perhaps a better metric for software complexity would be the amount of work the computer has to do, or the number of instructions it has to execute.
- nmehner 5y agoIsn't this the true for all metrics? They are useful for pointing out problematic areas, but as soon as you start optimizing only based on the metric things will get worse. Number of instructions is probably not a good metric. Without any loops/jumps you might have a lot of instructions, but a very low complexity. An endless loop executes a lot of instructions, but does not have to be complex.
- magicalhippo 5y agoComplexity isn't a single thing though. The problem can be complex, a specific implementation can be complex, the codebase can be complex. I recall implementing some linear algebra numerical code. The problem was a bit complex, resulting in a bit of code complexity. However I realized I had some extra information I hadn't used, and I spent half a day going over the math again. After a couple of pages of derivations I could narrow down the result to a couple of dot products. So, I ended up with a commit where I had 100 or so lines of comments including equations to justify my two lines of code. The implementation became super-simple, but why it worked was suddenly not so simple. I had effectively moved complexity from code-space to problem-space.
- drewcoo 5y agoThe problem is not computer processing speed. The problem is human cognitive load.
- hutzlibu 5y ago"Perhaps a better metric for software complexity would be the amount of work the computer has to do" By that definition a loop doing 1 billion times a simple calculation would count as very complex, even though it is very easy to understand. LOC would be better, even though code can be dense and complivated, or very verbose and simple.
- hacoo 5y agoUgh, no. I’ve worked in a codebase where CI would reject changes that had too much ‘code complexity’. You’d constantly have to find clever ways to split up your code, when doing so did not make sense, to appease the complexity checker. Oh yeah, and if you ever make a one-liner change, you might end up being forced to do a full refactor because that one line pushed the complexity threshold over the edge. The results: PITA for developers and worse code. What a crock of shit.
- juancb 5y agoDid you also have the same feedback from your IDE? In other words would it have been less painful if you didn't have to suffer the long iteration times required to get feedback from some remote CI job?
- nikita2206 5y agoThe thing is that cyclomatic complexity for example (most popular complexity measure that linters use), it doesn’t make sense. Most of the time high cyclomatic complexity of a method is indicative of high business logic complexity… which is fine. And dogmatically saying that methods shouldn’t have that many lines and branches and that you should just come up with better abstractions doesn’t help anyone, whereas having closely related functionality concentrated in a single place, rather than synthetically exploded into N different files, well this does help.
- rhn_mk1 5y agoI'm sorry about your experience, but from what it seems, your problem was in your team. Any measure can be turned into a bad policy, code complexity is a red herring here.
- pydry 5y agoCouldnt yout just raise the threshold by tweaking the config with your PR instead? I do this all the time with other automated checkers (linters, etc.). I don't see why this should be different. If another human agrees it shouldnt be a problem.
- throwaway81523 5y agoSome blog linked from here (maybe jvns.ca?) made the case that depth of your project's software dependency tree is an important metric. The more crap you have to pull in, the more things can go wrong. You're better off with a large program with no dependencies, than a somewhat smaller program with a ton of dependencies. Language features on the other hand can let you develop complex programs quickly and reliably, by catching errors before the code gets deployed and so on.
- juancb 5y agoWhat about the depth of dependencies of any part of the standard libraries for a given language?
- throwaway81523 5y agoI think that doesn't count, as long as the install is in one piece. That's how Python got its popularity. Its "batteries included" approach meant you got lots of stuff in the stdlib instead of having to chase it all over the interweb. Unfortunately they seemed to have abandoned that approach in more recent times.
- closeparen 5y agoPython has a lot of batteries included, but nearly everyone finds it necessary to bring in “requests.” And then once you have one pip package, what’s a few more?
- throwaway81523 5y agoI use urllib instead of requests just to get rid of a dependency. It does all the same stuff with slightly uglier calls. More to the point, Python Central seems to now favor shovelling off stuff to external dependencies.
- preseinger 5y agoIsn't the standard library definitionally depth=1?
- jillesvangurp 5y agoI always liked the notions of coupling and cohesion because they are simple to understand and you can see at the glance of an eye if a particular bit of code would have good metrics for those, without actually bothering with the metrics. E.g. long list of parameters or imports == high coupling, large number of fuctions in a module == low cohesion. Specifying exactly how much isn't that useful. It's easier to think in terms of "relative to the rest of the code". Debating what is too high or too low even less so. But if you are struggling with working with a particular bit of code, being able to identify why it is hard to deal with is useful; especially if you know how to fix it. But mostly metrics should not be telling you things you can't already know just looking at the code; if it looks complicated, it probably is. Metrics only become useful when you need to tell without looking. Sometimes that's useful.
- streamofdigits 5y agoThe Einstein quote is that "something should be a simple as needed but no simpler", but how to be pin down "required" in an objective / quantifiable way? Somehow it is more important to measure "gratuitous" complexity, redundant complexity that is not justified by present or plausible future requirements... The problem is that the code itself does not capture requirements, so code analysis can give you absolute indicators but never an "efficiency" measure (how efficient and justified the measured complexity)
- okl 5y agoIsn't the general goal rather to avoid/eliminate overly complex cruft (that every programmer should be able to recognize) than to find the maximally simple solution to a problem?
- streamofdigits 5y agothat seems like a more tractable goal if confined to local analysis (something like lines of code in a function or file unit) but people generally try to come up also with overall complexity that seems harder
- cryptica 5y agoNumber of lines of code is a good metric for complexity. The hard part is estimating how many lines of code is reasonable for a specific feature. Some complexity is necessary. The problem is unnecessary complexity.
- okl 5y ago> Number of lines of code is a good metric for complexity. I don't agree with that premise. LOC are a metric for code size, not for complexity. I've found that in practice the number of statements is a more reliable indicator for code size than lines of code. (For typical imperative languages anyways.)
- cryptica 5y agoLines of code is a very good approximation so long as each line is compiled into a similar number of bits across different systems and languages. The total size of the compiled source code in bits is literally the entropy of the system so it's the definition of complexity.
- okl 5y agoI think you applying that definition in your way to the issue of source code complexity is outlandish.
- cryptica 5y agoWhy is it outlandish? You're confusing the reliability of using source lines of code as a metric for measuring the productivity of developers with measuring the complexity of a system. It's a bad metric for measuring productivity but a good metric for measuring complexity. The problem with using lines of code as a metric for developer productivity is precisely that it leads to developers introducing unnecessary complexity into the system since they try to add as many lines of code as possible for implementing any feature. There is no drawback in keeping the source lines of code to the minimum amount necessary to get the job done.
- mikro2nd 5y agoI didn't see anything on Function Point counts in the article. No advocating it per se but it was, at one time, considered one of the more useful ways to evaluate codebase complexity.
- okl 5y agoFP analysis is a measure for the "amount of functionality" not complexity. Maybe in relation to, e.g., the number of statements (statements/FP) it could make sense.
- kqr 5y agoNumber of function points is, given implementation language, strongly correlated with lines of code. Lines of code is, in turn, strongly correlated with all other complexity measures. In other words, function point count is a measure of complexity.
- lincpa 5y agoIt is a simple, systematic, math-based method that `The Math-based Grand Unified Programming Theory: The Pure Function Pipeline Data Flow with Principle-based Warehouse/Workshop Model`, it makes development a simple task of serial and parallel functional pipelined "CRUD". ### Mathematical prototype - Its mathematical prototype is the simple, classic, vivid, and widely used in social production practice, elementary school mathematics "water input/output of the pool". ### Basic quality control - The code must meet the following three basic quality requirements before you can talk about other things. These simple and reliable evaluation criteria are enough to eliminate most unqualified codes. - Function evaluation: Just look at the shape of the code (pipeline structure weight), and whether the function is a pure function. - Functional pipelined dataflow evaluation: A data flow has at most two functions with side effects and only at the beginning and the end. - System evaluation: Just look at the circuit diagram, you can treat the function as a black box like an electronic component. - Code Quality Visualization: - For Lisp languages, S expression is contour graph, can be very simple transformation into contour map, or 3D mountain map. - If the height of the mountains is not high, and the altitude value is similar, it means that the quality of the code is good. - For non-Lisp languages, you can convert the source code into an abstract syntax tree (AST), and then into a contour map, or a 3D mountain map. ### Programming Aesthetics Simplicity, Unity, order, symmetry and definiteness. ---- Lin Pengcheng, Programming aesthetics The chief forms of beauty are order and symmetry and definiteness, which the mathematical sciences demonstrate in a special degree. ---- Aristotle, "Metaphysica" My programming aesthetic standards are derived from the basic principles of science. Newton, Einstein, Heisenberg, Aristotle and other major scientists hold this view. The aesthetics of non-art subjects are often complicated and mysterious, making it difficult to understand and learn. The pure function pipeline data flow provides a simple, clear, scientific and operable demonstration. Simplicity and Unity are the two guiding principles of scientific research and industrial production. - Unification of theories is the long-standing goal of the natural sciences; and modern physics offers a spectacular paradigm of its achievement. It can be found from the knowledge of various disciplines: the more universally applicable a unified theory, the simpler it is, and the more basic it is, the greater it is. - The more simple and unified things, the more suitable for large-scale industrial production. - Only simple can unity, only unity can be truly simple. In the IT field, only two systems fully comply with these 5 programming aesthetics: - Binary system The biggest advantage is that it makes the calculations reach the ultimate simplicity and unity, so digital logic circuits are produced, and then the large-scale industrial production methods of computer hardware are produced. - The Math-based Grand Unified Programming Theory: The Pure Function Pipeline Data Flow with Principle-based Warehouse/Workshop Model ### Others - Software and hardware are factories that manufacture data, so they have the same "warehouse/workshop model" and management methods as the manufacturing industry. - From the perspective of system architecture, it is a warehouse/workshop model fractal system. It abstracts every system architecture into a warehouse/workshop model . - From the perspective of component, it is a pure function pipeline fractal system. It abstracts everything into a pipeline. - It adheres strictly to 10 principles and 5 aesthetics, and it consists of 5 basic components. - It uses the "operational research" method to schedule the workshop to complete tasks in optimal order and maximum efficiency. ### Reference The Math-based Grand Unified Programming Theory: The Pure Function Pipeline Data Flow with Principle-based Warehouse/Workshop Model https://github.com/linpengcheng/PurefunctionPipelineDataflow https://github.com/linpengcheng/PurefunctionPipelineDataflow
- lngnmn2 5y agoVerbosity. Number of unnecessary abstractions.
- abacadaba 5y agotldr, blood pressure
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- SKILNER 5y agoThe real answers about complexity come from thinking about why we even care. It's because our feeble minds have to build internal models of the code so we can work with it. The cognitive aspects of building those models is why complexity matters. What things make it more difficult to build those models? A partial list, mostly as others have mentioned: - tool and library dependencies - nested conditions - loops and especially nested loops - asynchronous processing, callbacks, etc - non-descriptively named variables and functions - using non-standard code patterns for standard functionality - delocalized code, as in, you have to navigate somewhere else to see it (throws off your working memory) By one study, developers using Eclipse for Java spent 27% of their time just doing code navigation. The starting point for code complexity is about how our minds work. As many people have said, "it's easier to write a program than read it."
- deleted 5y ago[deleted]
- ip26 5y agoI would argue complexity is also the wellspring of corner cases. The more corner cases, the more mercurial- and that isn’t just due to the limitations of our minds.
- kristov 5y agoNot just build internal models of the code, but also an internal model of the execution of the code. For example, a mental model of the scoping rules. As a developer you have to read a function and build a model of when variables are assigned values - you have to execute the code in your head to understand what the state will be at a particular point in time. I think a big part of complexity is having to imagine what the current in-flight state is - the larger the current in-flight state, the harder to reason about how the code will interact with it.
- exabrial 5y agoThe number one metric: Consistency. If an app is similar to itself in all places, it's very easy to understand. Better yet, if it's similar and consistent with how other things have been built, we can call it "clean code". Back in the day we called these things architecture, but I'm old and salty.
- oblak 5y agoWhat if the code is consistently crappy and complex all around. Would that make it simple? Trick question, old man. It would not. I am not exactly young but damn I am sweet
- jackblemming 5y agoYou're not wrong. "Consistency" is a poorly defined principle and basically used to justify "what I'm doing is good (consistent). What you're doing is bad (inconsistent)" And it was called Uniformity back in the day, not "architecture". A better name for "consistency" is "following project conventions" which IS well defined.
- tharkun__ 5y agoI agree that consistency is good in general. I don't think "project conventions" is well defined. In most places project conventions are either not actually defined anywhere or if they are defined in writing, they're usually either very very old and outdated vs. the actual conventions that everyone is currently using or it's just one guy updating the text and hitting everyone else over the head with the document to push his opinion through. I personally like to be 'locally consistent'. I don't care how old and crusty the code base is. If the file that I have to change or add to calls everything a "giraffe", I will call my stuff "giraffe" as well, even if it really is a "gorilla". If I start calling it a gorilla, nobody will understand that the gorilla is the same as the giraffe if they don't have the same background knowledge I have. Unless I do a refactoring and I am changing the giraffes to gorillas. Which might either be a first PR to "clean up" or a follow up PR. Unfortunately I see so many people not doing that and it wreaks havoc with the code base. Especially if we're now outside of the place that defines the giraffes and gorillas. It's really hard for the caller to figure out that they're one and the same thing.
- pfdietz 5y agoThere were attempts to predict bugs by looking at complexity metrics. As I recall, the research found that when you adjusted for code size, none of the metrics mattered. In other words, just use LOCs as your metric.
- kqr 5y agoThis is correct. https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=code+complexity+lines&btnG=#d=gs_qabs&u=%23p%3DYrhiA1YB9RgJ https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=code...
- marcosdumay 5y agoAFAIK the one unambiguously relevant metric for how many bugs you will find in a codebase is how many its "clients" define as acceptable. Any kind of internal code quality affects the productivity of the debugging procedure, not the final number of bugs.
- ok_dad 5y agoOne thing I hate is having like 20 different Python repos for one small company. At most places I've worked at, you have basically one thing you do business-wise, but it's split into what I believe are arbitrary repo delineations. This causes trouble with dependency resolution in your build systems and IDEs and increases cognitive load and merge complexity for changes across repos. Just put the whole folder structure into one repo! This would reduce build complexity, and you can set up the configs for your tools and CI one time rather than 20. You can still have several services out of one repo, if you want to, but it's easier to reason about and easier to change those service delineations later, where in 20 repos you're having to "clone and cut code" to separate things. In one repo, you just move code around as you split or merge services. I routinely create a "super repo" for myself at these companies using submodules, so that I can actually work with the code more easily, but that still requires me to check in maybe 5 or more PRs for one feature, so it's not ideal. This only solves the developer's problems with local tools and still requires more complex debugging since the services are not actually in one repo under one config for deployment.
- quantified 5y agoI propose a challenge, similar to the Obfuscated C challenge, to devise code that is impenetrable to the human mind and yet fantastically clean to all metrics.
- jpswade 5y agoWhat if I told you that most software complexity doesn’t come from the code but from the software requirements?
- lightbendover 5y agoThe number of confused comments from PMs when engineering timelines are provided for “easy tasks.” Why does moving an ad above the fold require 3 weeks? Because it implicates 6 different teams.
- hulitu 5y ago> Our work, as developers, pushes us to take many decisions, from the architectural design to the code implementation. How do we make these decisions? Most of the time, we follow what “feel right”, that is, we rely on our intuition. So no engineering best practices. That explains very good the quality of SW.