3 ms·
I actually think raw byte-count is a pretty good metric. Documentation size and variable name length are also symptoms of complexity. Specifically I think wc -c
by anon_d 15y ago
I actually think raw byte-count is a pretty good metric. Documentation size and variable name length are also symptoms of complexity. Specifically I think wc -c $(find -type f) (how many total bytes) and find -type f | wc -c (how many files and directories) make good metrics. Obviously you need to filter out data files and such.
This has some problems (e.g. spaces vs. tabs, utf8, etc.), but all of these size metrics will be pretty loose.
- jacques_chester 15y agoOne metric I've seen is gzip-compressed size, which has the nice property that it identifies the size of the incompressible elements -- ie it discounts repetitive boilerplate. Another interesting set of metrics is Halstead's "software science" metrics[1]. They fell out of favour because initially they were hard to count and didn't seem to correlate with anything else. [1] http://en.wikipedia.org/wiki/Halstead_complexity_measures http://en.wikipedia.org/wiki/Halstead_complexity_measures
- gruseom 15y agoBut repetitive boilerplate is exactly the last thing that should get away scot-free in a measurement of code size.
- jacques_chester 15y agoIt depends on what you want to know. Pure physical lines is one thing, "size" is a another.
- gruseom 15y agoI want a way to measure how complicated a program is that's independent of language and obviously extraneous things like line length.
- jacques_chester 15y agoYou may find that the Halstead metrics I mentioned are closer to what you're after.
- gruseom 15y agoI've changed my mind. I'm interested in what I originally said: what's the best way to measure code size, and what are those studies (if they exist). Otherwise we get into debates about size vs. complexity, which is actually less interesting IMO. Size as a proxy for complexity is good enough for me.
- anon_d 15y agoI never understood the gzip one. Repetitive boilerplate is bad; why hide it?
- jacques_chester 15y agoYou're trying to understand the "true" size of the software in spite of the idiosyncrasies of a given language. As I noted somewhere above, "size" is an abstract, dimensionless quality. It can only be approached through proxies. The more the merrier, I reckon, especially if they turn out to correlate with different things.
- jcromartie 15y agoIn the case of most projects, copy/paste code is not just because of the language. It's because of lousy programmers. I've seen large codebases which are made up of a full 40% duplicate code. There's no way to blame that on the language.
- pyre 15y agoYou're missing the point of the parent: There is no 'one true' metric. If you use different metrics (actual lines, logical lines, gzip'd size, etc) you may well find different correlations.
- gruseom 15y agoBut this penalizes programmers who like to use long readable names. I'm not one of them (though I used to be), but they have a strong case here. Take any program. Replace all the names with the smallest possible character sequences. Have you made the program simpler? Or smaller in any meaningful way? Surely not. I'd say what you've done is left its logical structure precisely intact (another way of saying that token count is a good metric) while reducing its readability.
- anon_d 15y agoThis metric relies on the assumption that people are trying to produce readable code. IMHO long variable names are much more helpful in complex codes than simple ones.
- gruseom 15y agoOk, but now I'm wondering if we have opposite views of code size. In my view, code size is bad bad bad. More code means more complexity. Any time you add code, you're subtracting value; it's just that (if it's good code) you're adding more value than you're subtracting. So a higher score in a code size metric is a bad thing to aspire to, and we should greatly favor approaches to writing software that -- all other things being equal -- lead to smaller programs. I don't think that programmers who use long names for readability should have their programs discounted as longer (and thus more complex). Just because their names are longer doesn't mean their programs are.
- anon_d 15y agoNo no no. My logic is this: Take tight, readable code with short names a replace them with long names, and you'll have worse code. The converse isn't true because complex (bad) codes are more readable with long variable names. Complexity -> Code Size Code Size -> Long Variable names (win for big codes) Complexity is bad Therefore long variable names are a symptom of a problem, but not the problem themselves. Long variable names aren't bad, but they are still a good predictor of badness. Since size metrics are meant to predict badness, long identifiers should increase size metrics.