8 ms·
Maybe I'm just being jaded, and I'm certainly not a researcher or statistician, but I don't see how removing "statistical significance" from scientific nomencla
by krisrm 8y ago
Maybe I'm just being jaded, and I'm certainly not a researcher or statistician, but I don't see how removing "statistical significance" from scientific nomenclature is going to prevent lazy readers (or science reporters) from trying to distill a "yes/no" or "proven/unproven" answer from P values listed in a complex research paper.
- deleted 8y ago[deleted]
- bunderbunder 8y agoWell, that's why the article doesn't propose simply ditching P-values, it proposes reporting confidence intervals instead. Not only to they provide more information (by simultaneously conveying both statistical and practical significance) they're also easier to interpret correctly without special training.
- curiousgal 8y ago> they're also easier to interpret correctly without special training. Heh, you'd be surprised! Most people I met would interpret a 95% CI by saying that there is a 95% chance that it contains the true mean.
- taktoa 8y agoIs this not the definition of "confidence interval"? The first few Google results all define it this way...
- aidenn0 8y agoThat cannot be the result of an experiment without some sense of the prior probability, and I don't think CIs as suggested in the article account for that. For example, if I perform an experiment on a light source and find a 95% CI of 400-1200 lumens for the brightness, the actual probability of this being true is much higher if the light source is a 60W incandescent bulb than if the light source is the sun.
- mamon 8y agoexcept in real experimentation you control for all the variables except the one you're investigating. So it is not a very usefull thing to compare a bulb with the Sun. Instead, if you state that "luminescence OF THE 60W INCANDESCENT LIGHTBULB is 400-1200 lumens with 95% CI", then that's the usefull information that let's you set the right expectations when designing lightning for your new house, for example.
- aidenn0 8y agoI think you missed the point of my example. I was suggesting that an experiment performed on the sun showing a 95% CI of the brightness being 400-1200 lumens should result in a reasonable person believing that the probability of the Sun's brightness falling in that range is approximately zero, while the same result for a 60W light-bulb should result in a reasonable person being more than 95% certain that the bulb's brightness falls in that range. Just like a large number of people misinterpret a P value of .01 to mean a 1% chance of the results being due to chance[1], CIs can be similarly misinterpreted. 1: A .01 P value actually means that if the null hypothesis is true, then you would get the result 1% of the time. The analogy to my above example would be that if I run an experiment and get a result that "the sun is less bright than a 60W light-bulb" with a P value of .01, it's almost certainly not true that the sun is less bright than the light-bulb, since the prior probability of the sun being less bright than a 60W light bulb is many orders of magnitude smaller than 1%.
- khr 8y agoA 95% confidence interval will contain the true mean 95% of the time (across an infinite number of replications of the experiment/study). For a single confidence interval, you have either captured the mean in your confidence interval, or you've not -- there's no probability about it.
- justinpombrio 8y agoI believe this is the correct frequentist interpretation. To quote wikipedia: > A 95% confidence level does not mean that for a given realized interval there is a 95% probability that the population parameter lies within the interval (i.e., a 95% probability that the interval covers the population parameter).[10] According to the strict frequentist interpretation, once an interval is calculated, this interval either covers the parameter value or it does not; it is no longer a matter of probability.
- EForEndeavour 8y agoThis is where I get lost: > For a single confidence interval, you have either captured the mean in your confidence interval, or you've not -- there's no probability about it. Isn't there? The underlying truth is that you either definitely have or have not captured the population mean in any specific confidence interval. But you can't know this truth. In the long run, if "a 95% confidence interval contains the true mean 95% of the time across an infinite number of replications of the experiment/study," then isn't it true that any single specific experiment's CI has a 95% probability of containing the true value?! In my untrained mind, this is exactly equivalent to flipping an unfair coin with a 95% chance of heads. Sure, before flipping, the outcome of heads has a 95% probability. After flipping, you either get heads or tails. But if you flip a coin and hide the outcome without looking at it, doesn't it still have a 95% chance of being heads as far as the experimenter can tell?
- dxbydt 8y agoTo see how absurd that definition is, think about this - CI itself is random! So if you conduct a 100 experiments, you'll get a 100 (non-overlapping) CI's! So in which of those 100 CI's does the "true population mean" lie ? All 100 of them ? 95 of them ?! You tell me.
- BeetleB 8y ago>So if you conduct a 100 experiments, you'll get a 100 (non-overlapping) CI's! So in which of those 100 CI's does the "true population mean" lie ? All 100 of them ? 95 of them ?! You tell me. I don't know why you get the idea that all 100 will be non-overlapping. That's simply false. And yes, if your assumptions were correct, regular (i.e. frequentist) statistics will state that roughly 95% of the CIs will contain the true mean. There is nothing absurd about it.
- dxbydt 8y agoSpeaking as a Stat TA, literally over 90% of the students taking the class will conduct 1 experiment, not 100, which gives you 1 CI, not 100, & then say that particular CI has a 95% chance of containing the population mean! Then when I tell them the mean is either in that CI or not ( either 100% in, or 0% in ), they google the CI definition & point me to that. That's why I said that definition doesn't work for the masses. It can be interpreted as - "if you conduct a 100 experiments & get 100 CI's then roughly 95 of those will contain true mean", but then nobody ever conducts 100 experiments, so from their pov its an absurd definition. Its imperative to understand that these definitions are written not for the average user of statistics, but for a trained statistician. Unfortunately, the average stat consumer vastly outnumbers the professional. Papers are littered with statements like p value proves H0, or proves H1. I have had numerous conversations with scientists ( not statisticians, but pharma/epidem/engg people who show up to the stat lab for consult ) that their p value doesn't prove H0 or H1. "What do you mean you can't prove H1 ? Oh you mean it only rejects H0 ? Ok but isn't that same as prove H1 ? It isn't ?! Well in my field if I just state it rejects H1 it won't be well understood so I am going to instead say H1 has been proved!" So there's little the statistician can do. Regards overlap, I meant total/exact overlap, as in no two CIs will be identical on any conti dist.
- lutorm 8y agoUnless I've forgotten more than I hope, I believe the formal definition of a 95% confidence interval is that "if the model is true, 95% of the experiments would result in a point estimate within the interval." This is distinctly different from "a 95% probability that the true value is contained within the confidence interval", but that is typically what is loosely inferred.
- mr_toad 8y agoConfidence intervals can be applied to point estimates, estimates of means, and estimates of other things, including higher order moments. The difference is hugely important, the central limit theorem and the implication of a normal distribution commonly applies to sample means. You can calculate confidence intervals for most (not all) other statistics, like point estimates, but the distributions might not be normal.
- mike_ivanov 8y ago"for most (not all)" - yes, if analytically. If you can afford bootstrapping, then it is just "for all".
- BeetleB 8y agoNope. It is (from a frequentist's model of statistics), exactly what the article is claiming it isn't: If the model is true, and we repeat the experiment several times, 95% of the intervals we calculate will contain the true value. The actual CI you get in each experiment will differ. Another discrepancy between frequentists statistics and the article is that yes, the values at the boundary of your interval are as credible as in the center.
- ronjobber 8y agoThis is the first post I've seen that states the definition of a CI correctly. Another note: 'confidence interval' typically refers to the frequentist meaning, whereas 'credibility interval' is used in the Bayesian setting, when describing an interval of the posterior with 95% probability (which is arguably more interpretable). The usages of the two terms do not seem to generally be strict, however.
- pdkl95 8y agohttp://jakevdp.github.io/blog/2014/06/12/frequentism-and-bayesianism-3-confidence-credibility/ http://jakevdp.github.io/blog/2014/06/12/frequentism-and-bay...
- lucienlecam 8y agoThis one of many Bayesian vs. frequentist blog posts where the frequentist example is presented in such bad faith or is so wrong that it's impossible to take seriously. Why is the sample mean used for the frequentist CI when it is not a sufficient statistic, and especially since it appears after the section discussing a "common sense approach" in which the author does mention a sufficient statistic: min(D)? All this blog post shows is reasonable bayesian approaches are better than frequentist approaches where common sense isn't allowed.
- fenomas 8y agoUnless I'm missing something the author answers that here: > Edit, November 2014: ... Had we used, say, the Maximum Likelihood estimator or a sufficient estimator like min(x), our initial misinterpretation of the confidence interval would not have been as obviously wrong, and may even have fooled us into thinking we were right. But this does not change our central argument, which involves the question frequentism asks. Regardless of the estimator, if we try to use frequentism to ask about parameter values given observed data, we are making a mistake.
- lucienlecam 8y agoIn other words, the author has no mathematical examples to support the argument, and the objection is purely philosophical...
- fenomas 8y agoNot as I understand it. Note that the author's argument there isn't "frequentism is bad because it gives an unreasonable answer here", it's "the fact that frequentism gives a different answer here demonstrates that it really is answering a different question".
- alexgmcm 8y agoTechnically no for the standard frequentist confidence intervals, but if they use the Bayesian Credible Interval then I believe that would be the correct interpretation.
- deleted 8y ago[deleted]
- darkpuma 8y ago> "Well, that's why the article doesn't propose simply ditching P-values, it proposes reporting confidence intervals instead." You could possibly enforce that when publishing in journals, but there is no way in hell "science journalism" would follow suite, so the general public would still be just as hopelessly mislead, if not moreso if we account for the general statistical incompetence of the typical science journalist.
- ekianjo 8y agoThats a ridiculous assumption because I have seen over and over again people not understanding what confidence intervals actually mean. They look intuitive but they really are not as simple as they seem.
- vanderZwan 8y agoSure, but people are not necessarily scientists.
- stewbrew 8y agoWell, no. Hardly any non-statistician knows what a CI really means. It's also just a statement about significance in disguise.
- nabla9 8y ago"Lazy reader" is not the audience for scientific papers. Protecting against misinterpretations by outsides is not something that scientific research papers should worry about.
- bunderbunder 8y ago"Lazy reader" is precisely the audience for scientific papers. That's the gist of the crisis science has been having of late: The popular meme for a long time was that it was just non-scientists who would misinterpret the research. Recently, it has become clear that misinterpreting papers also rampant among trained scientists, to the extent that entire fields are being shown to be houses of cards.
- everdev 8y ago> "Lazy reader" is precisely the audience for scientific papers Exactly. There are plenty of personalities and pundits that love to say "look at the data" or "the data is clear". It bugs me when data is weaponized as truth to prove a conjecture, especially in the social sciences where studies are routinely difficult to replicate with consistent results.
- deleted 8y ago[deleted]
- MRD85 8y agoIt's not just misunderstanding papers, the entire social science field is in a replication crisis. A lot of people don't really want objective science and would rather push their agenda. Either that or they're simply incompetent.
- mar77i 8y agoIf anything, getting rid of some term because we currently find it insufficient to convey what we mean by it will, in all likelyhood, open the race for far more confusing, rosy and well-meaning, yet more meaningless nomenclature. By that measure, I find this entire idea rather idiotic.
- antidesitter 8y ago> getting rid of some term because we currently find it insufficient to convey what we mean by it will, in all likelyhood, open the race for far... more meaningless nomenclature You have given no reason to believe this.
- mateo1 8y agoThere's every reason to believe this. The article makes it very clear that the problem is the need by journals, funding agencies and researchers to have a boilerplate value so they can categorize the results as "true" or "false". You can change the value or the phrasing, but without solving the problem of "laziness" you won't fix anything.
- colechristensen 8y agoIf you eliminate p-value then you can't have authors that search for anything with p<0.05 and then publish, there will simply have to be some other justification. If p-value is gone it will have to be replaced with something and that something, the supposition is, will result in better science. Writing a paper, you need to support your conclusion, removing p-value doesn't remove that need for support, it will just find something different, hopefully better.
- seizethecheese 8y agoThis could make it worse. Instead of hacking for 5%, maybe they’ll hack for the lowest possible.
- colechristensen 8y agoTo be clear I mean rejecting p-values all together (or at least requiring additional evidence) not the specific <0.05 requirement.
- afiori 8y agoFor some context in the article they do not call for abolishing p-values but for a stop of the "significant" false dichotomy. An example would be to explain the consequences of all the values in the confidence interval or even to simply reformulate a sentence from "no significant effect was found" to "our data neither prove or disprove the presence of a significant effect"
- rejschaap 8y agoDescribed pretty accurately by this XKCD https://xkcd.com/882/ https://xkcd.com/882/
- stewbrew 8y agoThe proposal isn't about eliminating p.
- Fomite 8y agoNote that there are already journals that effectively do this (it is very hard to get a p-value put into the journal Epidemiology for example) and as far as I can tell, there are few if any negative repercussions evident.