9 ms·
How sufficient is a sample size of 31 healthy white men?
by diryawish 9y ago
How sufficient is a sample size of 31 healthy white men?
- jtmarmon 9y agoThis is what P-values are for
- shock 9y agoPerhaps explain how P-values are to be used in this case, for the lay of us.
- baobrain 9y agoDISCLAIMER: EXTREME SIMPLIFICATION Given a statement and a sample with a mean and standard deviation, what is the probability that you manage to choose a sample with that mean and SD, assuming the statement is true? For example, my statement is the average weight of chicken an american eats every month is 20 pounds and my sample gives me a mean of 16 and SD of 4. The p value is the probability of choosing a group of people who eat on average 16 pounds of chicken every month if the REAL mean is actually 20. Edit to answer question: The p value doesn't "care" about the sample size - it is adjusted and remains accurate. And so if the p value if high for the sample in the study, that means it is a likely occurrence and the result is not very noteworthy.
- Shank 9y agoA p-value in a test statistic is the probability that the result of the test statistic could have been achieved through random data variations. A p-value is described in tandem with an alpha level, the minimum acceptable value to reject the null hypothesis in the test statistic. In short: typical alpha is 0.05, so a p-value < 0.05 is considered "statistically significant" and the results are not attributable to random chance in a t-test. An unpaired t-test compares two populations, and a p-value below the alpha level indicates that differences are not due to random chance. Take everything with a grain of salt, though. Low sample sizes dramatically change how test statistics produce results. P-value hacking is the practice of eliminating data from an experiment to produce statistically significant results or otherwise altering the data by adding buffers to ensure that a result confirms a belief a researcher has.
- neaden 9y agoNot really. I mean in an ideal world yes, but right now we have no idea how many other people did a similar study, found no difference, and just didn't publish anything. Let alone accidental or purposeful p-hacking.
- deleted 9y ago[deleted]
- 3pt14159 9y agoIt isn't. People talking about P values forget about p-hacking and selection bias. Any scientific study with less than 200 participants should be discarded and the failure to routinely do so has more to do with scientific laziness and the pro-profit incentive most journals have than anything else.
- tomcooks 9y agoIs that 200 a random value or is it the result of a study (with hopefully at least 200 partecipants)
- QAPereo 9y agoThis is also the cause of rotating advice for decades, eg “eggs are/good/bad/good/bad/good” which is s feedback loop between terrible science, predatory media, and a credulous public.
- canjobear 9y agoThe number of participants required to show an effect should be shown by power analysis. It depends a lot on the effect being investigated. An arbitrary number like 200 is not helpful.
- sov 9y agoThat's not entirely true and doesn't really convey the right message about p-values or study size. P-hacking and selection bias are definite concerns, but they're also concerns for studies with 200, or more, participants. Rather, we should put more of an emphasis on the statistical power of the study. We're not yet in the beautiful golden world of pre-registered studies, but even high-n studies can have huge disparities due to bad statistical power (eg: The Control Group is Out of Control).
- 3pt14159 9y agoI'm essentially right. Studies with less than 200 people are essentially useless. It is more complicated than I originally let on, but not so much more complicated to muck up the takeaway. Most of our junk science has to do with these low population studies or with studies simultaneously studying high numbers of attributes.
- anigbrowl 9y agoIt isn't, but that's how you pitch for more research funding. I don't know anything about the economics of clinical studies or grant issuance, but I imagine it's hard to get the cash for a large sample up front so researchers go through this scaling process.
- bhc3 9y agoA metric that may be useful: fragility index. How many positive results would need to be reversed to have the p-value exceed 0.05? Seems applicable for low-powered studies such as this one. Details: https://emcrit.org/pulmcrit/fragility-index-ninds/ https://emcrit.org/pulmcrit/fragility-index-ninds/
- icegreentea2 9y agoIf you're asking if this is generalization to other healthy white men? It seems okay for a first shot. There is no fancy stats in this paper (they just use unpaired t-test to look for significance). They measured at two time points (+ the baseline condition) and both "real" (I don't really feel like including the Free T/LH ratio...) measurements that they found significantly different at minus 14 days remained significant (and same direction) at 44 days. The dose response curve seems particularly convincing to me. It does seem that the particular study was centered around a single electrically stimulated exercise session. It would be difficult to try to guess what happens if combined with different (more sustained? more repeated?) exercise. If you're asking if this is generalizable to any other group? No idea. Probably okay to generalize to all healthy men. But that's about it. As for why this is so small? Looks like they just piggy backed on another study. Pretty reasonable use of resources.
- meotai 9y agoEnough to be significant.