5 ms·
I'm not trying to be facetious, but isn't this something you learn in junior-level stats? I had this drilled in in both undergrad math courses and grad machine
by tyrankh 8y ago
I'm not trying to be facetious, but isn't this something you learn in junior-level stats? I had this drilled in in both undergrad math courses and grad machine learning courses; I'm confused to see it warrant an article.
- whatshisface 8y agoThe harsh reality is that most scientists are not sieving through every statistics book they can get their hands on in order to find out all the reasons they might be wrong. The "individual motivation" to become statistics experts is only present in a few fields, and in the others it is ousted into applied courses taught by other departments. Statistics is directly necessary in ML, so it's a "profit center" and emphasized. In many sciences it's treated like a cost center (something that you need, like IT, but that lies outside of your central expertise.)
- tyrankh 8y agoTIL thanks for the explanation, I guess I never thought about the fact that other STEM fields would not emphasize its meaning.
- pmyteh 8y agoIt's well known what p-values show. But they are, in practice, used as a gatekeeping mechanism in academic journals in many fields (including mine). Worse, getting p<0.05 is informally seen as a measure of practical significance, rather than simply as one statistical test amongst many passed. So yes, it is something you learn in introductory quantitative methods classes. But I don't think most researchers understand just how much it matters. Also, a key R package for producing regression tables of coefficients for journal articles is called 'stargazer'. Given the unwarranted focus of many readers on those indicia of 'significant' results, I think it's well named. I currently have the opposite problem. Given that I work with very large online datasets (N=1M or so) everything, including the random noise, is statistically significant to p<0.05. It really is effect sizes or busy at that point.
- AstralStorm 8y agoReal measures of practical significance are OR (odds ratio), effect size and dose response curve. Response histogram for statistical effects. (Or the 2D component analysis island histogram.)
- danieltillett 8y agoTo misquote Upton Sinclair you can’t get a scientist to understand statistics when their job depends on misunderstanding statistics. The basic problem is under the current funding environment it is far better to pump out a dozen wrong papers than one carefully researched paper.
- rossdavidh 8y agoThe literal answer to your question is, "no, generally not". That a greater emphasis on statistics should be included in science is certainly the case, but then there is a school of thought that know a little bit of frequentist statistics is better than knowing none at all. But regardless, I am fairly confident that most scientists (or engineers) do not actually learn this as juniors (or seniors) (or Ph.D's)
- Fomite 8y agoYou can end up with a Ph.D. in some fields being exposed to almost no statistics, or only statistics which work in very confined settings (certain experimental sciences where "Just do an ANOVA... is genuinely the answer to almost every question). That often works...right up until the moment when a scientist has to step outside that context. This often cuts both ways though. I have seen beautiful math and statistics around problems that don't make any sense if you've taken more than one semester of microbiology.