3 ms·
Where was I wrong? >Do you agree now that if the null hypothesis is true the p-value is uniformly distributed between 0 and 1? If the null hypothesis is true,
by jsprogrammer 11y ago
Where was I wrong?
>Do you agree now that if the null hypothesis is true the p-value is uniformly distributed between 0 and 1?
If the null hypothesis is true, the p-value will converge to 1. (This makes complete intuitive sense as well, if you hypothesis is true, every observation you make should/will agree with it, making the ratio of agreeable_observations:total_observations = 1.)
You haven't shown anything where a true null hypothesis uniformly generates p-values between 0 and 1. Perhaps for a single observation you can only get a 0 or 1 value, and, perhaps for a over/under the average test (where every experiment will give 0 or 1) your sequence of p-values for each observation would be uniformly distributed, but p-values are generally computed as a normalized sum over many observations, in which case the value should converge to 1 if the null hypothesis is true.
One unintuitive aspect is that you might expect p-value to converge to 0 if the null hypothesis is false, however p-value is undefined when the null hypothesis is false.
- kgwgk 11y agoOk, I see you still have your very own definition of "p-value". I just thought it was useful to make clear to the audience that your "p-values" are not the same "p-values" being discussed here. Edit: Probably you don't care, but for the record: "Since the value of x that defines the left tail or right tail event is a random variable, this makes the p-value a function of x and a random variable in itself defined uniformly over [0,1] interval, assuming x is continuous." [ https://en.wikipedia.org/wiki/P-value https://en.wikipedia.org/wiki/P-value ] "In statistics, when a p-value is used as a test statistic for a simple null hypothesis, and the distribution of the test statistic is continuous, then the p-value is uniformly distributed between 0 and 1 if the null hypothesis is true." [ https://en.wikipedia.org/wiki/Uniform_distribution_(continuous) https://en.wikipedia.org/wiki/Uniform_distribution_(continuo... ]
- jsprogrammer 11y agoMaybe you should pull out your full derivation. Random Wikipedia text doesn't count (particularly when it is just bald assertion, as in this case). Most null hypotheses hypothesize a normally distributed measurement error. For any given distribution (continuous uniform), of course, the p-value means something different. That is what I have been arguing this whole time.
- kgwgk 11y agoYou complained I didn't show anything, so there you have something. Do you have anything, even if it is a bald random Wikipedia assertion, supporting your position? (Rhetorical question, I couldn't care less.) Here is a full derivation from a random blog, which I guess doesn't count either: https://shihho.wordpress.com/2012/11/27/pvalue_distribution/ https://shihho.wordpress.com/2012/11/27/pvalue_distribution/ Another one: https://joyeuserrance.wordpress.com/2011/04/22/proof-that-p-values-under-the-null-are-uniformly-distributed/ https://joyeuserrance.wordpress.com/2011/04/22/proof-that-p-...
- jsprogrammer 11y agoYou seem awfully invested in a discussion that you don't care about. >https://shihho.wordpress.com/2012/11/27/pvalue_distribution/ https://shihho.wordpress.com/2012/11/27/pvalue_distribution/ This treatment commits the same error as the other poster (evanpw; I believe nonbel commits the same). The image at the bottom is only showing p-values for single observations. That is near useless (as the chart shows). An actual experiment requires repeated observation and analysis, not one-off observations. npval <- pnorm( rnorm(nSim, mean, std.dev), mean, std.dev ) hist(npval, breaks=40, xlab=xlabel, main=title, col='gray') R is probably the worst choice of language here, but my understanding of `pnorm` is that, when passed a list of numbers as the first argument (as `rnorm` returns), it will return a list of p-values (assuming normal distribution with given parameters), where the index in the list corresponds to the p-value, taking into account only that single observation (ie. it ignores all other observations in the list [ie. not how p-value hypothesis testing works {at least, in theory; real world practices may differ}]). Your second link appears to be the same argument (at least, they appear to use the same equations). The charts are a nice visualization of how you can get a uniform distribution from a normal one (or vice versa) though.
- jsprogrammer 11y agoYou might also want to take a look at the source of the function you are relying on: * DESCRIPTION * * The main computation evaluates near-minimax approximations derived * from those in "Rational Chebyshev approximations for the error * function" by W. J. Cody, Math. Comp., 1969, 631-637. This * transportable program uses rational functions that theoretically * approximate the normal distribution function to at least 18 * significant decimal digits. The accuracy achieved depends on the * arithmetic system, the compiler, the intrinsic functions, and * proper selection of the machine-dependent constants. * So...yeah. I hope the constants the R authors picked work for your machine! const static double a[5] = { 2.2352520354606839287, 161.02823106855587881, 1067.6894854603709582, 18154.981253343561249, 0.065682337918207449113 }; const static double b[4] = { 47.20258190468824187, 976.09855173777669322, 10260.932208618978205, 45507.789335026729956 }; const static double c[9] = { 0.39894151208813466764, 8.8831497943883759412, 93.506656132177855979, 597.27027639480026226, 2494.5375852903726711, 6848.1904505362823326, 11602.651437647350124, 9842.7148383839780218, 1.0765576773720192317e-8 }; const static double d[8] = { 22.266688044328115691, 235.38790178262499861, 1519.377599407554805, 6485.558298266760755, 18615.571640885098091, 34900.952721145977266, 38912.003286093271411, 19685.429676859990727 }; const static double p[6] = { 0.21589853405795699, 0.1274011611602473639, 0.022235277870649807, 0.001421619193227893466, 2.9112874951168792e-5, 0.02307344176494017303 }; const static double q[5] = { 1.28426009614491121, 0.468238212480865118, 0.0659881378689285515, 0.00378239633202758244, 7.29751555083966205e-5 }; https://svn.r-project.org/R/trunk/src/nmath/pnorm.c https://svn.r-project.org/R/trunk/src/nmath/pnorm.c Edit: Ah, a drive-by downmodder. How nice.
- nonbel 11y ago>"You haven't shown anything where a true null hypothesis uniformly generates p-values between 0 and 1." Try this in R. It calculates p-values for testing the hypothesis that two groups (n=10) come from the same distribution (normal with mean=0, sd=1). It does 100,000 such comparisons, then makes a histogram. This results in approximately a uniform(0,1) distribution of p-values: hist(replicate(10^5,t.test(rnorm(10),rnorm(10),var.equal=T)$p.value)) If you increase to 10^6, 10^7, etc it will converge upon the uniform distribution. That is what people mean. Edit: Or something like this. Keep adding to the sample, the p-value will random walk between 0 and 1. The mean will converge onto 0.5: n=10; Nsim=10000 a=rnorm(n); b=rnorm(n) p=matrix(nrow=Nsim,ncol=2) for(i in 1:Nsim){ p[i,1]=length(a) p[i,2]=t.test(a,b,var.equal=T)$p.value a=c(a,rnorm(n)); b=c(b,rnorm(n)) plot(p[,1],p[,2],ylim=c(0,1), ylab="pvalue", xlab="sample size") }