11 ms·
P values are not as reliable as many scientists assume (2014)
- deleted 11y ago[deleted]
- omginternets 11y agoAll of us? Scientific experiments always involve some degree of assumption. I'm most skeptical of those who don't state their assumptions.
- deleted 11y ago[deleted]
- geofft 11y agoHow about the assumption that the fundamental constants of the universe are not slowly changing day-to-day?
- unabst 11y agoA better word for that would be "premise" or "axiom". Premises and axioms are objective exceptions with objective merits. Assumptions are too personal because they infer belief which is purely subjective. Nature doesn't care about what anyone believes and science should never be a democracy. Assumptions also imply some independent existential entity as valid and are self-validating, whereas premises and axioms are highly self-deprecating. Hence, assumptions are dogmatic self-fulfilling prophesies that are an end unto itself, whereas premises are unfortunate unavoidable constraints as a means to an end. They are what couldn't be eliminated despite towering doubt and cynicism. Premises lead to science, axioms to logic, and assumptions to religion (and the like -- not saying it is good or bad).
- Retra 11y agoYou're making a distinction that I've never heard anyone make before, and I don't think you're making a convincing argument for it now.
- unabst 11y agoThe distinction between axiom and assumption is quite clear, so I'll assume you are referring to premise. The distinction already exists, which is the beauty of words. I am not making this distinction up. I am merely enforcing them as we do as speakers by the words we choose, based on the accuracy of our expressions. From Popper: > A theoretical system may be said to be axiomatized if ... (d) necessary, for the same purpose; which means that they should contain no superfluous assumptions [0]. So either Popper is wrong, or theories should not include unnecessary assumptions. But how do we know if an assumption is necessary without doing the science? And after we do it, are we still going to call necessary assumptions assumptions, even with its subjective implications? Only a person is capable of assuming. A theory with assumptions is still subjective. A premise on the other hand is objective and specific. It's "a proposition supporting or helping to support a conclusion" [1]. It's a simple device in logic that asserts a dependency. Or in other words, a "necessary assumption". So then would it not be safer to say axiomatized theories have premises, not assumptions? And in Popper's words, yet to be axiomatized theories have assumptions. That is what makes them hypotheses. And so all the words and their distinctions fall cleanly into place (I did not make anything up). The original article in Nature was written on the premise that science is based on assumptions and that scientists doing the science rely on assumptions. This is not the premise of science, and is incorrect. Refining our word selection to reflect this understanding would be of great service, particularly to the students. -- [0] https://books.google.com/books?id=cAKCAgAAQBAJ&pg=PA51&lpg=PA51&dq=%22no+superfluous+assumptions%22+popper&source=bl&ots=2SpoXnlniS&sig=hocFKDmxKzHeOFlKUoQ0mhOr1lE&hl=en&sa=X&ved=0CC0Q6AEwA2oVChMImc3VrtHSxwIVFCuICh3eywzT#v=onepage&q=%22no%20superfluous%20assumptions%22%20popper&f=false https://books.google.com/books?id=cAKCAgAAQBAJ&pg=PA51&lpg=P... [1] Just from the dictionary, as to avoid my own words. http://dictionary.reference.com/browse/premise http://dictionary.reference.com/browse/premise
- sciguy77 11y agoFirst off, there's no need for all-caps. This isn't 4chan. > The fact that assumptions are considered as some unavoidable, forgivable, intricate part of science is part of what fuels anti-science and politics. No one here, as far as I can tell, is saying, 'oh well, science is full of assumptions therefore science is invalid.' The problem is not with science in general being valid or invalid, but rather with the sorts of experiments being conducted right now. Studies that are not replicable most of the time are bad science, which is different than science itself being 'bad.' I'm not saying such faulty experiments and studies are unavoidable and forgivable. Quite the opposite, they're flawed and need to be scrutinized more, not less. Finally, this is not an incurable problem. Bayesian math is one potential solution. There are others.
- unabst 11y ago> 'oh well, science is full of assumptions therefore science is invalid.' No, they are saying "science is based on assumption, but in this case 'many scientists' were making the wrong assumptions." The title should have said "scientists find errors in p-values as premises." Assumptions are avoided, not depended upon. Bad assumptions would certainly lead to irreplicable experiments also. The distinction you make between science and its experiments is not a common distinction. If science = experiments, then you totally agree that this kind of science is bad science. Which was precisely my point.
- omginternets 11y ago>First off, there's no need for all-caps. This isn't 4chan. Oh man, it looks like I missed something priceless.
- qudat 11y agoI agree that science strives to remove bias from its body of knowledge, but it's absolutely unavoidable for humans to paint their subjective experiences onto it. Humans have to bring their own preconceptions to any scientific experiment, even the medium by which we convey the knowledge is bathed in assumptions about what those words mean. Every facet of human knowledge is premised on how humans experience the universe. Given your definition, I don't think anything we know would be considered a fact. https://www.youtube.com/watch?v=cG3sfrK5B4E https://www.youtube.com/watch?v=cG3sfrK5B4E
- unabst 11y ago> Every facet of human knowledge is premised on how humans experience the universe Absolutely. But that is why this is where we start. Before science, we had no way to invalidate illusions and validate what was real because assumption on their own are neither. They are naked intuitions pending validation. For the longest time we were unable to validate them and we ended up with the mess we had before science. Basically, no one would ever have made it to Mars. But with scientific validation knowledge becomes more than just an assumption or an intuition based on an experience, or a theory we came up with that we find ingenious because, well, we came up with it. By overcoming our assumptions we achieve objectivity, universality, and factuality. We discover knowledge that has rigid practical persistence. In this process something transcends from our subjective personal ideas to becoming objective impersonal facts. There is no self in science. And it is from this arduous feat that technology is born. There is nothing in this monitor or the components of this phone that are based on assumptions. These devices are selfless. "Assumption" is as evil a word as "metaphysics" and "subjective" in science. Yet, there are still people who use the word as a synonym for axiom. This is simply bad word-choice. The correct term here would be "premise" and you used it yourself. Theories can have premises, but not assumptions. Are the premises assumed? No. They are granted. Since this is HN, here is an analogy to software. A program that assumes certain behavior code or of external APIs will be rigged with bugs. Every aspect of it's execution must be tested, and the assumptions of the programmer must be eliminated by production. Of course, being human, we start with assumptions - such as "this would be the perfect library for this project". The "assumptions" that we being with however, eventually manifest themselves into "premises". And in software, these are the dependencies of a program. It is only natural for software to be dependent on other software. What is unnatural and anti-software would be to make assumptions about other software especially within its own execution. The path from the assumptions of subjective raw experience to the subjective consumption of reliable technology is paved with the work of competent scientists (and analogously, by competent programmers). > Given your definition, I don't think anything we know would be considered a fact. If fact is to mean truth, then sure. But there is an abundance of statements that have been backed by evidence. And all these statements are truer than most. Measurably truthier, rather, and that is what counts because that leave room for progress. This is a better definition of "fact".
- BorisVSchmid 11y agoFrom my experience, scientists, -at least in biology, where like in sociology you might have a lot of noise to deal with-, have an internal intuition that a single paper with a significant result does not mean that we have found the truth. The recent study which reported a reproducibility in sociology of about 36% strikes me as pretty accurate. I think the scientific system can work with that. It means that if you build follow-up experiments based on a single paper there is a good chance that the experiment fails. In some way, the scientific system of publishing is self-correcting in this regard, because you can then cast doubt on the previous paper, which is easier to publish than if you only have a fresh negative result (p-value > threshold).
- pekk 11y agoThere is no way to know how many people tried to build a follow-up experiment which failed and was not published because the failure to replicate will usually be assumed to be due to some mistake, and even carefully finding p > threshold is not very publishable.
- csirac2 11y agoLuckily some disciplines have a journal of negative results :-) Eg. http://www.jnr-eeb.org/index.php/jnr http://www.jnr-eeb.org/index.php/jnr
- stdbrouw 11y agoA large amount of published results that are wrong is definitely something science can live with: we have to trade off Type I against Type II. But we should value accuracy: if we report something as being very, very unlikely if chance was at play, and it turns out that in fact (1) it'd be very likely even if the null hypothesis holds and (2) in fact even if P(D|H0) is low, P(H0|D) might be high... then what's the point in writing up all those fancy statistical analyses anyway? At that point significance testing becomes more of religious ritual and should either be discarded entirely or be amended.
- jsprogrammer 11y agoIf the p-values were accurate and averaged around 0.05, ~95% of results should be reproducible. That only 36% were points to deep, fundamental errors.
- danharaj 11y agoScientists have to do their work in a system that incentivizes bad science. How many people actually get to do their work in an environment that isn't hostile to them?
- themodelplumber 11y agoa) Are you serious about that second question and b) if so, can we discount thermodynamics in our answer? Otherwise it's kind of boring.
- danharaj 11y agoWe can restrict ourselves to social factors. Nature isn't hostile, it doesn't have human intentionality like that. Seems to me we make work unpleasant for everyone in the misguided belief that people work harder for it.
- kazinator 11y agoPrevious post with discussion, 563 days ago: https://news.ycombinator.com/item?id=7225739 https://news.ycombinator.com/item?id=7225739 PDF via same nature.com: https://news.ycombinator.com/item?id=8404620 https://news.ycombinator.com/item?id=8404620 Related, dupes of each other: https://news.ycombinator.com/item?id=9463806 https://news.ycombinator.com/item?id=9463806 https://news.ycombinator.com/item?id=9486059 https://news.ycombinator.com/item?id=9486059 Related: https://news.ycombinator.com/item?id=9119228 https://news.ycombinator.com/item?id=9119228
- eruditely 11y agoRelevant, from Deborah Mayo. http://errorstatistics.com/2015/03/16/stephen-senn-the-pathetic-p-value-guest-post/ http://errorstatistics.com/2015/03/16/stephen-senn-the-pathe...
- themodelplumber 11y ago"Essentially, all models are wrong, but some are useful." --George E.P. Box
- danparsonson 11y agoThe p-value test isn't a model, it's a measure of the significance of an effect in data against random noise.
- tel 11y agoWhich arises from a model (!) of random noise and of your effect.
- danparsonson 11y agoI see - my mistake. That's a very broad definition of 'model' though isn't it? Including 'random numbers'? You might as well say everything is a model in which case the original quote says nothing :-)
- tel 11y agoIt's perhaps a bit like "everything is a model" in the sense that all of these tests, even the model-free ones, arise from a coherent choice of assumptions and, if you for a moment take the Bayesian perspective very seriously, prior distributions over conditionals. The original quote should be taken to mean that any particular choice of assumptions is limiting, but making interesting choices can drive interesting questions which are thought provoking and meaningful even if they are wrong.
- RA_Fisher 11y agoGreat article. I'm not sure that replication itself will solve the problem since Type 1 error rate requires asymptotics. We'd have to run many replications and then show convergence. That'll be broadly cost-prohibitive for all but the most important conclusions. Lower thresholds probably won't do it either. Right now, the only solutions I see are: a) Baysian methods b) Fisher's single H hypothesis method c) Tukey's Exploratory Data Analysis method. d) All of the above.
- Fomite 11y agoRelying on effect measures and their confidence intervals, rather than relying on p-values as a "Yes/No" threshold should likely be on your list, especially as an interim step that should be easy for those who still want p-values to swallow.
- sgerrish 11y agoI don't see why (e) teaching scientists to be statistically literate so they don't abuse or misunderstand these tests, and/or (f) focusing on reproducible results and shaming researchers with sloppy methodology, wouldn't work. The hypothesis test has known limitations, but it's not clear that we should blame null hypothesis tests for people mis-using them, when researchers untrained in stats are just as likely to mis-use any method you give them.
- rndn 11y agoIsn’t a main problem with p-values that you don’t know whether significance (low p-value) is a result of big effect and small sample or big sample and small effect. This is why you also need a measure for the effect, for example the distance of the two measurements in terms of standard derivations.
- james1071 11y agoThat is a separate issue. The main problem with p-values is that, without further information, one cannot infer from them how likely it is that a result is genuine.
- jjoonathan 11y agoI agree with TFA that p-hacking is a bigger problem. Low p-value <=> null hypothesis is unlikely. Choose a shitty null hypothesis ("aliens did it!", "everything is Gaussian", etc) and you trivially get low p. Peer review checks this to some extent (you won't get away with "aliens did it") but there's a large gray area of null hypotheses shitty enough to give low p but not shitty enough to be rejected by peer review. Choosing the hypothesis after-the-fact is the most common strategy because it's undetectable except by repeating the experiment, which is hard.
- haddr 11y agoIt is not that P-values are now bad by definition. It's only that they are many times wrongly intepreted. Putting too much confidence in P-values only might result in some wrong conclusions. And this is what some meta analyses discover. Many scientists try hard only to reach the "golden" <0.05 in order to claim discovery and publish it. This is why there is so many papers that misteriously cluster around 0.05...
- tel 11y agoThere's also the systemic effect of prioritizing particular p values in that negative results are omitted leading to replication bias across the community.
- marvy 11y agoI'm probably commenting too late to get my question answered, but here goes: the article has a pretty picture where they show how likely your p-values will mislead you depending on how likely the null hypothesis is. For instance, they say if you think that the null hypothesis has a 50% probability of being right, and you get p=5%, then there's still a 29% chance the null hypothesis is true. But according to my calculations, the right number should be 1/21 = 4.8%. What am I missing here? Or are they wrong? My calculations are below: Curious George has 200 fascinating phenomena he wishes to investigate. In reality, 100 of those are real, and the other hundred are mere coincidences. The experiments for the 100 real phenomena all show that "yes, this is for real". (I'm assuming no false negatives.) Most of the 100 experiments that test bogus phenomena show that "this is bogus", but 5 of them achieve a significance of p=5%, as expected. George then runs of to tell the Man in the Yellow Hat about his 105 amazing discoveries. If Yellow Hat Man knows that half of the phenomena that capture George's attention are bogus, he knows that 5/105 = 1/21 = 4.8% of George's discoveries are likely bogus, even though he doesn't know which ones.
- kgwgk 11y agoAssume that you're sampling from a normal distribution with known standard deviation sigma (1 for simplicity) and unknown mean mu. To test if the mean is larger than (the null hypothesis) mu=0 you can check if the observed value is larger than 1.64 sigma (for the 95% confidence test). So if your observation is larger than 1.64 you reject the null hypothesis. Your calculation would be correct only if the assumption "no false negatives" is approximately valid. This is the case when the true value is large in terms of sigma (say mu=6). Then for the 100 cases with mu=0 you'll reject the null 5 times on average, and for each one of the 100 cases with mu=100 you will reject the null (unless you're unlucky: there will be a false negative around once in 150000 trials). But you're conditioning on p<0.05, not on p~0.05. It's easy to see that it's much easier to get p=0.05 if mu=0 (this is a 1.64 sigma event) than if mu=6 (it's a 4.36 sigma event). If mu=0, you will get on average 1 (out of 100) observation with 0.04<p<0.05 (i.e. 1.64<x<1.75). The probability of obtaining an observation on that range when mu=6 is very small (0.0004 out of 100). Almost 100% of the "discoveries" with p~0.05 will be false (when mu=6 you will get p-values around 1e-9). When the true value of mu gets closer to 0, you cannot ignore the false negatives. For example if mu=0.1 the rejection rate will be quite similar to the mu=0 case (the probability of getting 0.4<p<0.5 is 1.2% and 1% respectively) and almost 50% of the "discoveries" with p~0.05 will be false. Somewhere between the two extreme cases, there is a lower bound for this "false discovery rate". See http://faculty.washington.edu/jonno/SISG-2011/lectures/sellkeetal01.pdf http://faculty.washington.edu/jonno/SISG-2011/lectures/sellk... and in particular figure 2.