10 ms·
P < 0.05 Considered Harmful
- deleted 3y ago[deleted]
- belter 3y agoThere is a joke between Mathematicians: The Confidence Interval is nothing more than, the interval between first learning about it, thinking you understood it, and the time it takes until you realize what it really means.
- kneebonian 3y agoThe thing that made me realize how ineffective P < 0.05 was, was playing D&D, because the odds of something happening 5% of the time was the same odds as rolling a 1 on a 20 sided dice, which happens surprisingly frequently once you roll it more than a couple of times. Also XCOM taught me that 98% != 100%
- galkk 3y agoXcom isn’t a good example because the game actively lies to you with displayed probabilities https://youtu.be/l0KEDYFWbVc https://youtu.be/l0KEDYFWbVc
- kneebonian 3y ago"That's XCOM baby!"
- pphysch 3y agoMost big games implement "randomness" with "pseudorandomness" in the name of controlling variance of outcome, chopping off the long tail
- C-x_C-f 3y agoDo pseudorandom distributions have chopped tails? Wouldn't that go against the definition of pseudorandom?
- pphysch 3y agoTo be clear it's a mix of pseudorandomness and "procedural randomness" to increase "fairness"
- the_af 3y agoIs there any game that does NOT use pseudorandom generators? And does this significantly change probabilities?
- pphysch 3y agoSome games I follow talked about "switching to pseudorandomness" but that may be misuse of terminology. There's an argument for competitive games to use true randomness to eliminate any possibility of abuse, but I'm not aware of specific examples.
- burnished 3y agoThink there is confusion here - if I understand correctly you are asking about the number generator, they are talking about the process of determining success. Like in league of legends you have a listed crit chance but the way they determine success isnt to generate a number and compare it to your chance, you start with a smaller base number that gets incremented each time you fail and reset when you succeed - the end result is that your overall chance remains the same but the likelihood oh a streak (of fails or successes) goes down. Doesnt change the overall probability, drastically reduces variance.
- PeterisP 3y agoSome games (both computer games and physical board games) intentionally use "shuffled randomness" where e.g. for percentile fail/success rolls you'd take numbers from 1-100 and use true randomness to shuffle that list; in this way the overall probability is the same, but has a substantially different feel as it's impossible for someone to have bad/good luck throughout the whole game and things like "gambler's fallacy" which are false for actual randomness become true.
- teddyh 3y ago> the overall probability is the same Only for the very first roll. After that, the outcome becomes more and more predictable for each roll.
- 3y ago
- the_af 3y agoCan you elaborate on how XCOM lies? I often suspected this (but you can never be sure, since human intuition is bad at probabilities). Is there hard evidence?
- fauxpause_ 3y agoI believe it fudges the numbers to give you better than expected results on lower difficulties.
- galkk 3y agohttps://youtu.be/l0KEDYFWbVc https://youtu.be/l0KEDYFWbVc
- kibwen 3y agoI've seen no source that shows that Xcom fudges its displayed hit chances. You may be thinking of Fire Emblem, whose games use a variety of well-documented approaches to fudging their rolls: https://fireemblemwiki.org/wiki/True_hit https://fireemblemwiki.org/wiki/True_hit
- reibitto 3y agoI know at least XCOM 2 does on certain difficulties. The aim assist values are directly in the INI files (it fudges the numbers in your favor for lower difficulties). Here are instructions on how to remove the aim assists: https://steamcommunity.com/sharedfiles/filedetails/?id=617993180 https://steamcommunity.com/sharedfiles/filedetails/?id=61799...
- yunruse 3y ago2σ is fine, but the benefit of modern technology is that we can tell exactly what standard deviation would be needed for the null hypothesis to randomly generate our results. Particle physics holds itself to an "industry standard" of 5 sigma, for example. The real conversation to be had is -- what standard deviation will we tolerate? Is this something we'll keep doing, and thus A/B test ourselves into a (perhaps quite horrifying) local minimum based on random noise? Is this a single great experiment? Are lives on the line? Will these results be taken Quite Seriously? Is this a test to say "look, this is worth further investigation"? I'm no statistician but in my opinion this is the first conversation that needs to be had when doing an experiment: what p / σ levels are satisfactory to claim confidence? p<0.05 is a decent heuristic for some experiments, but not all.
- revision17 3y agoAmerican Statistical Association also released a statement on pvalues: https://www.tandfonline.com/doi/epdf/10.1080/00031305.2016.1154108?needAccess=true&role=button https://www.tandfonline.com/doi/epdf/10.1080/00031305.2016.1...
- clircle 3y agoMaybe tech industry insiders can tell me this ... but do real people actually make product decisions based solely on p < 0.05? Seems like the author is writing about a contrived problem.
- candiddevmike 3y agoProduct decisions are made based on someone's gut instinct. p values are mostly used when (or abused until) they align with that instinct, if they are being considered at all.
- bookish 3y agoThe "stasis" and "arbitrarily adjustments" regimes that I wrote about are certainly ones that I've seen, which don't rely solely on p < 0.05 but are still pretty suboptimal. Furthermore, it's not only about whether 0.05 is the sole criteria, but also about whether it's a useful criteria for us to highlight at all, depending on whether the anchoring effect of it is damaging relative to alternatives. But let me turn that around and ask: what product decision regime do you see most often or think would be the most relevant to use as an example? I'd be happy to hear your perspective and make sure I keep it in mind for future blogs.
- ftxbro 3y agoHi I wrote some other response in this thread where I called you a techbro, so sorry about that I probably count as one too I wasn't trying to insult you too much. Anyway when you say "Furthermore, it's not only about whether 0.05 is the sole criteria, but also about whether it's a useful criteria for us to highlight at all, depending on whether the anchoring effect of it is damaging relative to alternatives." I love this analogy of the null hypothesis vs. alternative hypotheses in frequentist statistics to the 'anchoring effect' cognitive bias that you try to work to your advantage in marketing and sales and negotiation or management. https://en.wikipedia.org/wiki/Anchoring_(cognitive_bias) https://en.wikipedia.org/wiki/Anchoring_(cognitive_bias) If you don't want to be tied to p-values and you only care about downstream decisions rather than quantifying beliefs, you can use some ideas in decision theory https://en.wikipedia.org/wiki/Decision_theory https://en.wikipedia.org/wiki/Decision_theory For example, maybe you are deciding between two alternative ways of doing something. You don't know which one is better, and you are confronted with not only the decision of which one to use, but also with the decision of whether to experiment (with A/B testing for example) to be more sure of which one is right, versus whether to exploit the one that you currently think is better. This is the multi-arm bandit problem and it doesn't necessarily use p-values so your intuition is right! https://en.wikipedia.org/wiki/Multi-armed_bandit https://en.wikipedia.org/wiki/Multi-armed_bandit Maybe that's not your situation. Maybe your situation is that you have an existing business process and you want to know whether to switch to one that might be better. Someone might say it's a p-value problem, but again I agree with your intuition that it really isn't best to think about it that way, especially when there is a cost to switching. Instead, it's a more complicated decision that depends on what is the switching cost, how much better you think the new process would be (including uncertainty of it), and what kind of business horizon you care about. There might even be a multi-armed bandit effect again even in this situation, where you also have to weigh the costs of reducing your uncertainty of the switching improvement or even of reducing the uncertainty of the switching cost itself. Anyway, these problems do involve concepts from probability and statistics but it's for sure true that the decisions don't always reduce to P < 0.05 at the end! Good luck best wishes living your best techbro life!
- gjm11 3y agoThe article is all about why "0.05" might be a bad value to choose. But, more fundamentally, p is often the wrong thing to be looking at in the first place. 1. Effect sizes. Suppose you are a doctor or a patient and you are interested in two drugs. Both are known to be safe (maybe they've been used for decades for some problem other than the one you're now facing). As for efficacy against the problem you have, one has been tried on 10000 people, and it gave an average benefit of 0.02 units with a standard deviation of 1 unit, on some 5-point scale. So the standard deviation of the average over 10k people is about 0.01 units, the average benefit is about 2 sigma, and p is about 0.05. Very nice. The other drug has only been tested on 100 people. It gave an average benefit of 0.1 unit with a standard deviation of 0.5 units. Standard deviation of average is about 0.05, average benefit is about 2 sigma, p is again about 0.05. Are these two interventions equally promising? Heck no. The first one almost certainly does very little on average, and does substantially more harm than good about half the time. The second one is probably about 5x better on average, and seems to be less likely to harm you. It's more uncertain because the sample size is smaller, and for sure we should do a study with more patients to nail it down better, but I would definitely prefer the second drug. (With those very large standard deviations, if the second drug didn't help me I would want to try the first one, in case I'm one of the lucky people it gives > 1 unit of benefit to. But it might well be > 1 unit of harm instead.) Looking only at p-values means only caring about effect size in so far as it affects how confident you are that there's any effect at all. (Or, e.g., any improvement on the previous best.) But usually you do, in fact, care about the effect size too. Here's another way to think about this. When computing a p-value, you are asking "if the null hypothesis is true, how likely are results like the ones we actually got?". That's a reasonable question. But you will notice that it makes no reference at all to any not-null hypothesis. p < 0.05 means something a bit like "the null hypothesis is probably wrong" (though that is not in fact quite what it means) but usually you also care how wrong it is, and the p-value won't tell you that. 2. Prior probability. The parenthetical remark in the last paragraph indicates another way in which the p-value is fundamentally the Wrong Thing. Suppose your null hypothesis is "people cannot psychically foretell the future by looking at tea leaves". If you test this and get a p=0.05 "positive" result, then indeed you should probably think it a little more likely than you previously did that this sort of clairvoyance is possible. But if you are a reasonable person, your previous opinion was a much-much-less-than-5% chance that tasseomancy actually works[1], and when someone gets a 1-in-20 positive result you should be thinking "oh, they got lucky", not "oh, it seems tasseomancy works after all". [1] By psychic powers, anyway. Some people might be good at predicting the future and just pretend to be doing it by reading tea leaves, or imagine that that's how they're doing it. 3. Model errors. And, of course, if someone purporting to read the future in their tea leaves does really well -- maybe they get p=0.0000001 -- this still doesn't oblige a reasonable person to start believing in tasseomancy. That p-value comes from assuming a particular model of what's going on and, again, the test makes no reference to any specific alternative hypothesis. If you see p=0.0000001 then you can be pretty confident that the null hypothesis's model is wrong, but it could be wrong in lots of ways. For instance, maybe the test subject cheated; maybe that probability comes from assuming a normal distribution but the actual distribution is much heavier-tailed; maybe you're measuring something and your measurement process is biased, and your model assumes all the errors are independent; maybe there's a way for the test subject to get good results that doesn't require either cheating or psychic powers. None of these things is helped much by replacing p=0.05 with p=0.001 or p=0.25. They're fundamental problems with the whole idea that p-values are what we should care about in the first place. (I am not claiming that p-values are worthless. It is sometimes useful to know that your test got results that are unlikely-to-such-and-such-a-degree to be the result of such-and-such a particular sort of random chance. Just so long as you are capable of distinguishing that from "X is effective" or "Y is good" or "Z is real", which it seems many people are not.)
- uniqueuid 3y agoSo much has been said about p-values and null hypothesis significance testing (NHST) that one blog post probably won't change anybody's opinion. But I want to recommend the wonderful paper "the null ritual" [1] by Gerd Gigerenzer et al. It shows that a precise understanding of what a p-value means is extremely rare even among statistics lecturers, and, more interestingly, that there have always been fundamentally different understandings even when p-values were invented, e.g. between Fisher and Pearson. Beyond that, there was a special issue recently in the american statistician discussing at length the issues of p-values in general and .05 in particular [2]. Personally, I of course feel that null hypothesis testing is often stupid and harmful, because null effects never exist in reality, effects without magnitude are useless, and because the "statelessness" of NHST creates or exacerbates problems such as publication bias and lack of power. [1] http://library.mpib-berlin.mpg.de/ft/gg/GG_Null_2004.pdf http://library.mpib-berlin.mpg.de/ft/gg/GG_Null_2004.pdf [2] https://www.tandfonline.com/toc/utas20/73/sup1 https://www.tandfonline.com/toc/utas20/73/sup1
- bookish 3y agoThanks for that. I'll give [1] a read. I'm familiar with [2], and cited one of those papers in the blog. About the stupid or harmful nature of null hypothesis testing in general, what do you recommend instead for decision making and for summarization of uncertainty? In the scenario of large (yet fast moving) organizations where most people will have little stats background.
- uniqueuid 3y agoThanks, I hope you find Gigerenzer useful. The paper is a bit academic, but he also wrote a couple of nice popular science books on the (mis-)perception of numbers and statistics, those might be useful in a business environment. For real-world applications outside engineering and academia, I would rely heavily on confidence intervals and/or confidence bands. For example, the packages from easystats [1] in R have quite a few very useful visualization functions, which make it very easy to interpret results of statistical tests. You can even get a textual precise description, but then again, that's intended for papers and not a wider audience. Apart from that, I would mainly echo recommendations from people like Andrew Gelman, John Tukey, Edward Tufte etc.: Visuals are extremely useful and contain a lot of data. Use e.g. scatterplots with jittered points to show raw data and the goodness of fit. People will intuitively make more of it than of a single p-value. [1] https://easystats.github.io/easystats/ https://easystats.github.io/easystats/
- pasc1878 3y agoxkcd has commented on that different pvalues https://xkcd.com/1478/ https://xkcd.com/1478/ and shown an example of ,marketing using p values https://xkcd.com/882/ https://xkcd.com/882/
- bookish 3y agoClassics!
- ftxbro 3y agoThis post is amazing. When I see a post "P < 0.05 Considered Harmful" I think OK now they are going to talk about maybe Bonferroni vs. other multiple hypothesis corrections if they are a frequentist or otherwise they are going to try to explain Bayesian things. But no, this one isn't from a frequentist, or from a Bayesian. It's from a techbro whose solution isn't any kind of multiple hypothesis correction or getting Bayes-pilled, it's to say "Why not just admit you want something akin to p = 0.25 in the first place?" for 'ship criteria' in the only stats context he appears to know which is A/B testing, talking about Maslow hierarchy and namedropping Hula and Netflix. It's seriously like some Silicon Valley parody. Wait is it actually a satire blog?
- whimsicalism 3y ago> "Why not just admit you want something akin to p = 0.25 in the first place?" for 'ship criteria' That's culture shock for me - I guess this is why I don't work at startups.
- bookish 3y agoIt all depends on what you're experimenting on. I do think there are many teams out there who are (in effect) making decisions with less certainty than this, but they wouldn't want to actually quantify it.
- mlyle 3y agop=0.2 doesn't work too well for ship criteria for medicine. p=0.2 for "this reordering of the landing text improves the rate conversion events" is fine. People make changes based on less information all the time. Waiting for certainty has its own expenses.
- mike_hearn 3y agoNot that fine. The author claims that neutral changes are not costly for users so the only reason to avoid them is to avoid wasting time/money. But that's not really true. Pointless churn annoys users. If your P threshold isn't low enough then you can get stuck in an endless treadmill of making changes that you thought would have benefit but which don't actually do anything because your threshold for something being considered significant is too loose.
- Joel_Mckay 3y agoStill not significant: https://mchankins.wordpress.com/2013/04/21/still-not-significant-2/ https://mchankins.wordpress.com/2013/04/21/still-not-signifi... Nothing like spending 10 minutes reading a paper to see results which are likely nonsense. However, it pales in comparison to spending 3 weeks trying to replicate popular works... only to find it doesn't generalize... you know that ROC was likely from cooked data-sets confounded with systematic compression artifact errors... likely not harmful, but certainly irritating. lol =)
- ttpphd 3y agoWhat a bananas list! Thanks for sharing it.
- hsjqllzlfkf 3y agoP < 0.05 considered harmful 5% of the times.
- begemotz 3y agoBesides the Gigerenzer article mentioned below (there are others by the same author worth reading e.g. 'Mindless Statistics'), I would recommend: 'Moving to a world beyond "p <0.05" by Wasserstein et al in the American Statistician https://www.tandfonline.com/doi/full/10.1080/00031305.2019.1583913 https://www.tandfonline.com/doi/full/10.1080/00031305.2019.1... as well as the classic article by Jacob Cohen "The Earth is Round (p <.05). YMMV, but in certain disciplines still, "statistical analysis" is little more than checking for p-values and applying a binary decision rule. That is without recognizing the shaky theoretical ground of NHST as practiced.
- staunton 3y agoThe fundamental error that causes misuse of p-values (and statistics in general) is misunderstanding what statistics is in the first place. Statistics is applied epistemology. There is no algorithm that fits all situations where you're trying to think about and learn new things. It's just hard and we have to deal with that. Arguably, the main point of p-values is trying to prevent people who really know better from "cheating" by reporting "interesting" results from very low sample sizes. Having a very rigid framework as a rejection criteria helps with this. However, the scientific community and system are not capable of dealing with "real cheating" that includes fabrication of data. Also, any such rigid metric is going to be gamed. Of course, there are some people who would also cheat themselves and maybe learning about p-values makes this less frequent. But such people using them don't understand what those values are telling them, beyond "yes, I can publish this". This is counterproductive because it prevents deep thinking and thorough investigation. Most scientists realize that in practice there is no simple list of objective criteria that tells you what experiments to perform and how exactly to interpret the results. This takes a lot of work, careful thinking, trying different things and very much benefits from collaboration. But who's got time for that? There's papers to publish and grant applications to write. So p-values it is. Or maybe some other thing eventually, that also won't solve the fundamental problem.
- Kalanos 3y agoI don't like p-val as a cutoff, but do like to know if the p is low. if I see 0.003 I can quickly get a sense of legitimacy
- tqi 3y ago> Neutral changes aren’t costly on our users, so while we should be somewhat averse to wasting time and adding tech debt for neutral changes, it isn’t the end of the world. While I think most of the points in the post are reasonable, I strongly disagree with the idea that neutral changes aren't costly. Some neutral changes are good because they are a part of a larger strategy or vision, but in my experience most neutral features are simply change for changes sake. Any reasonablely large company that ships all/most neutral tests is going to end up feeling like a bloated, unfocused mess.
- bookish 3y agoDoes the sentence imply that they aren't costly? Would you write "strongly averse" instead of "somewhat averse" or "it is the end of the world" instead of "it isn't the end of the world"? If so, what language would you use to convey that strongly negative changes are that much worse than neutral ones?
- beckhamc 3y agoA relevant quote here seems to be: "if a measure becomes a target, it ceases to become a good measure". Also true of academia in general because of the obsession with citations.