7 ms·
It seems that the main issue here is that with pre-registration, study authors have to pick a single measure of primary benefit at the outset, whereas before, t
by dthal 11y ago
It seems that the main issue here is that with pre-registration, study authors have to pick a single measure of primary benefit at the outset, whereas before, they might have made that choice after getting results back. The original study is at PLoSONe, and it is not a difficult read[1]. From that source:
>>Prior to 2000, investigators had a greater opportunity to measure a range of variables and to select the most successful outcomes when reporting their results... Among the 25 preregistered trials published in 2000 or later, 12 reported significant, positive effects for cardiovascular-related variables other than the primary outcome.
That is, in most cases, there are large effects for some outcome, and if they get to choose the primary outcome after looking at some results, they could have been cherry-picking the outcome variables.
[1] http://journals.plos.org/plosone/article?id=10.1371/journal.pone.0132382 http://journals.plos.org/plosone/article?id=10.1371/journal....
- narrator 11y agoTo be fair, Viagra was originally being developed as a high blood pressure medication. They decided to switch to erectile dysfunction when they found out why study participants were hoarding the pills. A lot of science is about serendipity.
- alephnil 11y agoPre-registration would not hinder them doing that discovery, but they would have had to do another registration and study to get it to the market with the new use. They would likely still be able to use the outcome of the original study to assess the safety of the drug.
- zmmmmm 11y agoWhich actually makes an important but subtle point: just because the rate of positive effects identified went down doesn't mean the effects that would otherwise have been identified were all false. It just means we don't know. There may be extremely strong evidence of effects that would overcome any amount of multiple testing correction, but they still aren't allowed to use it. It might take years or even decades for them to do a follow up study to validate the result, which means a significant number of people are harmed by not having access to the drug in the interim. Just playing devil's advocate to make the point that it's not a given that we're getting an better overall outcome by being this stringent.
- eru 11y ago> [...], which means a significant number of people are harmed by not having access to the drug in the interim. Yes, that might happen. The more likely outcome though is that we are saved from a lot of drugs that don't work any better than chance (or even worse).
- dthal 11y agoFair enough... but the issue here is about findings of an effect versus no effect. It's about statistical significance. In order for single-comparison p-values and the like to be valid, there has to be a (one) single comparison. There is a way to do 'any-of-k' testing, but the required effect sizes get larger.
- bunderbunder 11y agoif they get to choose the primary outcome after looking at some results, they could have been cherry-picking the outcome variables Not just could, would. Choosing your hypothesis after you run the experiment is (or at least should be) a cardinal sin in science for good reason. At the standard p-value cutoff of .05, even when there's absolutely no effect going on the probability of getting a spurious positive result when you do n comparisons is equal to 1 - (.95^n). So that 5% chance of a type I error if you only look at one test statistic jumps to 40% if you look at ten, and to 72% if you look at 25. Here's a nice piece of gonzo journalism that deals with this issue: http://io9.com/i-fooled-millions-into-thinking-chocolate-helps-weight-1707251800 http://io9.com/i-fooled-millions-into-thinking-chocolate-hel...
- jsprogrammer 11y agoCardinals and sin are the realm of religion. Not sure why it's being brought up here. Choosing a hypothesis after the experiment is run is perfectly valid as long as your experiment is valid for that hypothesis. Besides, you would always run a new experiment again anyway.
- obastani 11y agoIt's fine if you run a separate experiment to justify the hypothesis. If you choose the hypothesis after conducting the experiment, then the p-values you obtain for that hypothesis are invalid.
- thedufer 11y agoIf you choose your hypothesis after running the experiment, you can almost guarantee a positive outcome, regardless of whether there's any effect. Are you claiming that this statement is untrue, or irrelevant? > Besides, you would always run a new experiment again anyway. Then surely the initial experiment doesn't count. Otherwise, why are you re-running it?
- jsprogrammer 11y agoYes, you can always choose any hypothesis you want. It's largely irrelevant. Every experiment will support analysis through a set of hypotheses. Just because you didn't select all of those hypotheses before the experiment ran doesn't mean you can't select it after the experiment. Imagine that an experiment has been run, but you do not know the results (or even what was done). Now you select a hypothesis, if the experiment required to validate the hypothesis is the same as what was run previously, you can now look at and use the results. A hypothesis is like running a query against a database. Many queries are valid, even though the data may not have changed. >Otherwise, why are you re-running it? Science requires it. Doctrine from one-off experimentation is religion (hard to dump).
- jrochkind1 11y ago> It seems that the main issue here is that with pre-registration, study authors have to pick a single measure of primary benefit at the outset, whereas before, they might have made that choice after getting results back Right. And then you'd have to use proper statistical reasoning for that state of affairs. Which nobody ever does, cause they're not statisticians and it's complicated and it would reduce the chance of 'statistical significance'. So they just use a standard calculation of statistical significance -- which is based on the assumption that you have picked a single hypothesis in advance and then done your test. So it's completely invalid to use it how everyone typically does. Imagine you flip a coin 50 times. Then you see, okay, did I ever get 10 heads in a row? Nope? Okay, how about 5 heads followed by 5 tails? Nope. Okay.... try a couple dozen other things, oh, look, I got exactly 3 tails followed by exactly 3 heads followed by exactly 3 tails again! Let's run my test of statistical significance to see if that was just chance, or is likely significant -- oh hey, it's significant, this is likely a magic coin not random at all! Nope. If you test everything you can think of, _something_ will come up as 'statistically significant', but it's not really, those tests of statistical significance -- which calculate how likely it is the results you got happened by random chance happenstance vs an actual correlation likely to be repeatable -- are no longer valid if you go hunting for significance like that.