5 ms·
I've always had mixed feelings about this graph. If you look at the actual number of responses for Julia, it's only 786 loved, 298 dreaded. Compare this to Pyth
by ducharmdev 4y ago
I've always had mixed feelings about this graph. If you look at the actual number of responses for Julia, it's only 786 loved, 298 dreaded. Compare this to Python, which is 22,999 loved, 11,156 dreaded - is this really an accurate presentation of the data?
Sure, taking the Julia responses alone, you could argue that a high proportion of language's users love it. But when placed alongside other languages, it leads some to conclude that Julia is loved more than Python; if 22,999 people love Python and 786 love Julia, is that really true?
- wbsss4412 4y agoEven with only a bit over 1,000 total responses, yes you can make such statements. More isn’t necessarily magically better when it comes to sample sizes. In both cases, we are making inferences about what the total population thinks about the languages based on a subset of survey responses. Now, is there reason to think that there is bias in these survey results? That is certainly a possibility. To a degree I agree with you in that the framing of “most loved” could be approached better with niche vs widespread languages. There are a lot of people who self select into small languages, vs languages in widespread use, so perhaps noting the relative usage would be helpful in contextualizing what people think.
- nonameiguess 4y agoI believe what the parent is getting at isn't that the sample size for Julia is too small to conclude anything, it's that the differences in sample size between Python and Julia reflect differences in sizes of the total user base. That means Python is more likely to be used by people who have no other choice, whereas more Julia users are using it because they like it. If Julia grew to be as popular as Python, it might no longer be loved by as large a proportion of its users.
- sgt101 4y agowell - also self selection. People who respond Julia have (almost certainly) not been made to use it. Python... definitely not so much. Disclaimer: I don't really like Python, but I don't dread it. Disclaimer 2: I quite like Julia but every time I have tried to drop Python to use it I have had to go back to Python.
- nequo 4y ago> I quite like Julia but every time I have tried to drop Python to use it I have had to go back to Python. Is that because of library availability or something to do with the interpreter?
- nemetroid 4y agoThe issue is not with sample size but population size.
- time_to_smile 4y agoThe issue isn't sample size (nor many of the other things people are pointing out). First off we're ultimately talking about ordinal data here and even the Evan Miller post in another comment misses that ordinal data is fundamentally tricky. Everyone is talking about the p("most loved") but that makes assumptions that "most loved" has a consistent meaning. The problem with ordinal data is that the only thing that's known about it is an ordering of the values, but the distance between values is undefined. That is the difference between a 5 star and a 4 star review is not necessarily that same a between a 2 and 1 star review. This means that you can't meaningfully average these. However there is an even bigger problem, which is what I think parent is feeling here, and that is selection bias. It's very hard to compare the average rating of a film such as Star Wars and a film like Cannibal Holocaust. The first is a general audience film and the later is an extremely niche subset of gore horror. If you show Cannibal Holocaust to the general population it will receive wildly lower rating than Star Wars. However if you were to do a survey of all people who have a DVD of Star Wars and a DVD of Cannibal Holocaust, I wouldn't be at all surprised if Cannibal Holocaust was more loved. There are plenty of people that have Star Wars on their shelf and feel 'meh' about it, but almost no one who owns Cannibal Holocaust on DVD that doesn't feel like it is an essential horror classic. The annoying truth is, there's really no statistical way to solve this. There are some solutions to better modeling ordinal data (you could for example, transform this problem into an ordinal regression problem). But ultimately "loved" is not really a measurable thing, so we're always looking for imperfect proxies.
- ducharmdev 4y agoThanks for the thoughtful reply - I don't have a stats background, but it didn't quite sit right with me. The movie example you give here really puts it in perspective.
- marrone12 4y agoThis article gives a handy method to account for exactly what you're asking https://www.evanmiller.org/how-not-to-sort-by-average-rating.html https://www.evanmiller.org/how-not-to-sort-by-average-rating...
- ducharmdev 4y agoThat's really cool, thanks for sharing! I particularly appreciate the SQL snippet.
- spywaregorilla 4y agoI find the conclusion to a really silly solution if you need to explain it to anyone. My preferred answer is a basyesian average. That is, everything gets n votes for the average score by default, and you just sort by the average rating.
- SleekEagle 4y agoWe demand confidence intervals!!
- samch93 4y agotry the following R code: fisher.test(matrix(c(786, 298, 22999, 11156), nrow = 2)) #> Fisher's Exact Test for Count Data #> #>data: matrix(c(786, 298, 22999, 11156), nrow = 2) #>p-value = 0.0003266 #>alternative hypothesis: true odds ratio is not equal to 1 #>95 percent confidence interval: #> 1.116046 1.469612 #>sample estimates: #>odds ratio #> 1.279435
- mushbino 4y agoIf 11,156 dreaded Python, but only 298 dreaded Julia, doesn't that make Julia look a whole lot better?
- leephillips 4y agoI’m an incorrigible Julia booster, but I don’t see much real value in these statistics or rankings. Always consider the data source, bias due to self-selection, etc. Results from surveys of any kind are rarely informative.
- throwaway894345 4y agoMoreover, these kinds of rankings often conflate user satisfaction with hype (which isn't to say a hyped language isn't a good one, but just that "I love it" can mean "I love what I've heard about it, but I haven't used it"). I'm not sure about SO's love/dread in particular.
- raid615 4y agoI'd say that since way less people use Julia than Python, there is likely to be more people among the sample who are Julia enthusiasts. While Python on the other hand is used extensively because it is a previously established and much more popular language, there are likely to be more people that use it just for that reason. The perception is that there's no reason to use Julia if you don't like it, but you may have to use Python even if you don't like it.
- sieste 4y agoWhen YouTube still showed dislikes, some justin bieber video was simultaneously the most liked and most disliked video of all. i quite like to think of python as the justin bieber of programming languages.
- ryandvm 4y agoI'd also imagine that people using Julia in 2022 are doing so because they wanted to use Julia. They're probably senior and they have a lot of flexibility in tooling selection. Whereas I know A LOT of people using Python and JavaScript against their will.