4 ms·
I feel that these popularity indices are pretty flawed and most often used as arguments under a confirmation bias. Defining "Popularity" is no good measure in m
by cessor 12y ago
I feel that these popularity indices are pretty flawed and most often used as arguments under a confirmation bias. Defining "Popularity" is no good measure in my opinion. This graph just shows that with many lines changed many questions arise, it does not say whether people enjoy doing so.
I would be much more interested in spots like "Logos" in the bottom right - 160 Mio Lines changed, only 38 questions of StackOverflow. So what does that mean? Is everything about the language crystal clear? Beat that, JavaScript.
Github is a reliable source but I am not sure whether it is representative. At least it appears to be very popular with web folk. Other sources are missing such as bitbucket, especially since it allows for free of charge private repositories. Like this, the only thing this says is that github use correlates with the lines of JavaScript. Not sure whether this is a good measure for popularity. Much like a kid in kindergarten, who shows up every day and talks to everyone. Yes, his presence correlates with the presence of others, but what if he is just a jerk and nobody really likes him?
- mjw 12y agoFair points. And of course basing the github measure on lines of code gives quite an unfair advantage to verbose languages like Java. Perhaps they could look at compression ratios when (say) gzipping a decent sample of each language, and use these to weight the LOC metric?
- Sandman 12y agoOr they could do away with LOC as a metric entirely and use number of repos instead?
- dtech 12y agoThat would probably overvalue scripting languages, as they are more likely to be used for smaller languages. Also a lot of repos contain multiple languages (web languages more often than not do)
- cessor 12y agoYou are absolutely right, I didn't think about this when I wrote my comment. LOC is in itself a very poor metric, therefore "Changed LOC" doesn't improve any interpretation. Including the number and sizes of repos would be better, since LOC pushes verbose languages. Java and C# code usually features a lot of empty or hollow lines (in C# for example, one opens the curly brace for a function in a new line). I mean, yeah, most of the comments indicate that we agree that this kind of graph or "false statistic" is flawed. But can we find any value in it? How do we interpret this graph, despite its basic problems? What good stuff can we do with it?
- mjw 12y agoA scale for cross-language comparisons seems a hard ask because everyone is implicitly interested in answering different questions. Which languages do people enjoy using the most? Which have the most code "out there" in some setting or other (open source, commerce, in deployment, scripting, ...)? Which have the most engaged and vocal communities? Which have the biggest pool of skilled workers? which will I get the most career benefit from learning? Perhaps the useful information is more in the correlations between different related metrics, than in drawing up ranked lists (surprise surprise, Java > Haskell!). This graph helps visualise the correlation between these two, which seems significant but far from perfect. Outliers like SQL can then be identified, which point to problems with the metrics (e.g. with SQL, presumably github fails to spot lines of SQL embedded in other langauges).