4 ms·
Related, about p-values: > Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p <
by rom1v 7y ago
Related, about p-values:
> Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p < .05) in 95 of those experiments, you might expect that 5% of these 95 significant effects would be false positives. However, as an example shown later in this blog will show, the actual false positive rate may be 47%.
> […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, and we want to know whether we should accept or reject the null hypothesis. The probability of a false positive in that situation is not the same as the probability of a false positive when the null hypothesis is true. It can be way higher.
https://lucklab.ucdavis.edu/blog/2018/4/19/why-i-lost-faith-in-p-values https://lucklab.ucdavis.edu/blog/2018/4/19/why-i-lost-faith-...
> Here's a more simple thought experiment that gets across the point of why p(null | significant effect) /= p(significant effect | null), and why p-values are flawed as stated in the post.
> Imagine a society where scientists are really, really bad at hypothesis generation. In fact, they're so bad that they only test null hypothesis that are true. So in this hypothetical society, the null hypothesis in any scientific experiment ever done is true. But statistically using a p value of 0.05, we'll still reject the null in 5% of experiments. And those experiments will then end up being published in scientific literature. But then this society's scientific literature now only contains false results - literally all published scientific results are false.
> Of course, in real life, we hope that our scientists have better intuition for what is in fact true - that is, we hope that the "prior" probability in Bayes' theorem, p(null), is not 1.
https://news.ycombinator.com/item?id=16917158 https://news.ycombinator.com/item?id=16917158
- swsieber 7y agoFor about a year or so now Ive been wanting to make a game about science. It'd basically be a research and discovery simulator, and there would be a free-play mode. Some of the knobs would be # of required replication, required p-value and how good people are at generating hypothesis. I think it'd be eye opening.
- marcosdumay 7y agoThere is a "people will mostly replicate/extend articles about X, and ignore articles about Y" (groupthink) effect that I imagine is also very relevant.
- swsieber 7y agoOooh, that's a great one to add to the list!
- jeffhuys 7y agoPlease do. We're getting closer and closer to the tipping point (although it still might be far away) where people go "Oooohhhh... Well, shit.". Where we all realise. Your game would be one of the many catalysers for this. This is my opinion of course, it's not like I have any scientific basis for this.
- heymijo 7y agoI hope you do! I think simulations can be a powerful way to teach/learn and I would like to see this one. I put a reminder to check back with you about this in a few months. Is Keybase your preferred contact method? I know it in name only, but I imagine I can figure it out.
- SilasX 7y agoYes! Riffing off that, I'd want to do a version where all the data in the game is pure randomness, but where your experimental configurations can influence the outcome. (e.g. you can clean the equipment, which alters the reading, or redo experiments that "went wrong") You'd be given a prize/goal for publishable findings. Then you'd gradually introduce enough bias into the experiments to get something publishable, and then get hit with the reveal that "oh you were generating effects from random data, jerk". (Okay, maybe that's overcomplicating it.)
- kazinator 7y agoThat's not what the p-value means. It means that if you run 1000 of the experiments in a universe in which the hypothesis is false, around 50 of them will confirm the hypothesis anyway. If the hypothesis is true, then there are no false positives; all positives confirm the hypothesis. In a universe in which the hypothesis is true, there can only be false negatives. "False positive" means that the effect or condition we're looking for is not true, but the experiment yields a true answer: the positive answer of the experiment is a falsehood. If the condition we're looking for is true, then there can't be a false positive. Even if the experiment yields a positive due to some flawed step, it's still a true positive.
- rom1v 7y ago> The false positive rate (Type I error rate) as defined by NHST is the probability that you will falsely reject the null hypothesis when the null hypothesis is true. In other words, if you reject the null hypothesis when p < .05, this guarantees that you will get a significant (but bogus) effect in only 5% of experiments in which the null hypothesis is true. This is just a language issue: a false positive of the rejection of the null hypothesis.
- BeetleB 7y ago>If you run 1000 experiments over the course of your career, and you get a significant effect (p < .05) in 95 of those experiments I'm all for criticism of p-values, but when I read a lot of critiques, I get to this point and simply stop reading. No statistics text book that I've read assigns the magic value of p=0.05 and labels it as significant. All the ones I've read tell you to pick a p-value appropriate to your experiment. Yes, I get it that many social scientists don't have much of a clue and use 0.05 as some special threshold, but let's direct the criticism to the guilty parties, instead of blaming a statistical methodology. I mean, we all know people who misuse the mean ("the average number of breasts a person has is 1") and ignore the shape of the distribution and the standard deviation. Yet we don't say "Let's stop using the mean!"
- fastaguy88 7y agoThis blog post is extremely misleading -- it appears to confuse the "null hypothesis" with an actual scientific hypothesis that is being tested. For example: "However, there are many cases where I am testing bold, risky hypotheses—that is, hypotheses that are unlikely to be true." Those may be the hypotheses being tested, but they are not NULL hypotheses. NHST is not about testing the NULL hypothesis, it is about testing a non-Null hypothesis (the null hypothesis should always be incredibly boring and expected). There may be problems, but this blog post does not describe one.