3 ms·
Unfortunately, reasonable has nothing to do with it. Though I must admit my first comment is rather sarcastic and not very informative. Let me elaborate on why
by miggol 3y ago
Unfortunately, reasonable has nothing to do with it.
Though I must admit my first comment is rather sarcastic and not very informative. Let me elaborate on why the high drop out rate is so damning.
Let's use a marble analogy because they are common in statistics. You gather 132 marbles of that vary in size in a way that is representative of the marble population as a whole. You then randomly assign them to either an intervention (n=68) or control (n=64) group. This random assignment is already an intervention in its own right, but with these numbers you could still say the groups are pretty much comparable.
Now, you run your experiment. Your hypothesis is that the marbles in your intervention group will grow in size, and not the control. But you don't know that, and we should assume the null hypothesis until disproven.
It's time to measure your marbles. It would make sense to measure them all, but you don't do that. You take a sample (n=12) for the intervention group and another sample from the control (n=11) and measure those samples. The rest are dropouts.
The reasons that caused these marbles to drop out are irrelevant. I'm not suggesting the researchers are paid shills and consciously selected marbles to suit their narrative. For all we know all the other marbles were stolen, or are invalid for other causes outside the researchers' control.
The problem is that this act of subsampling has wrecked the experimental design either way. The null hypothesis suggests that both groups were not only initially equal, but also remained equal regardless of intervention. So you must ask yourself the following: if I took such a small sample from each group before I had even performed the intervention, what are the odds that one sample would be significantly bigger than the other by pure chance?
I'm not doing the maths, and they would heavily rely on the imaginary distribution of marble sizes in this analogy anyway. But I dare say that one sample would be significantly bigger than the other one often enough. And half of those times it would be the intervention group. So before even considering the effects of your own intervention, you're already fighting a losing battle. Because of subsampling the null hypothesis has grown stronger and can now explain even a very large difference between groups.
So what do you do as a conscientious researcher when faced with the hardships of campus shutdowns and participants that didn't do cognitive tests in person but via zoom? My recommendation would be to run your statistical tests on everything and publish it anyway. Acknowledge that garbage in = garbage out, there were many confounders, so you can't draw any conclusions. Sometimes pandemics get in the way of science.
You could even publish the results for all participants side by side with the restrictive subsample. If they both point to the same result, perhaps you even have grounds for some kind of valid conclusion! Makes you wonder why they didn't just do that, huh.
What you definitely shouldn't do is only publish your statistical tests on the subgroup and draw conclusions based on those alone. And then not acknowledge how the massive group of participants that you _chose_ to leave out could have affected the results, or how the choice to leave them out affects the null hypothesis.
Hope this helps!
(I'm not a researcher but I work in academia and have on occasion assisted in experimental design.)
- godelski 3y agoOh yes, I do agree that the sample size is problematic, but that's honestly not uncommon in anything with medical research. Samples sizes are always crazy small and people are making wildly too large of conclusions from them given that. But it's also not like you can expect to form a hypothesis and then get 50k people to participate in a study. You gotta show that there's something to the idea first. That's what these kinds of papers are about. And for the dropoff rate, I'm less concerned with that than 1) the (different) bias introduced by the people who chose to drop out vs not and 2) the overall sample size being smaller. Were people that dropped out uniformly distributed and there was still a large sample size, it wouldn't be of any concern. The statistical error you're making here is actually the same as the Monte Hall one. Basically, ignore the previous group and pretend they don't exist. Though there probably is a bias introduced from the dropout (which isn't biased in the Monte Hall problem), but there's so many other effects going on here that it's hard to say. (And medical students and human phys students aren't known for having the strongest stats skills. To be fair, it's a fucking tough field and they got a lot of extra factors that don't make it any easier) But again, let's never think of these works as proofs of effects but rather like a proof of concept. Even if they didn't have any dropout I'd say the same thing here. Evidence is good, more people makes the evidence stronger, but you gotta keep pushing for stronger and stronger evidence but university research papers in human studies are just never strong evidence. I am a researcher. I do think it is good that you are being critical, but it's often too easy to be overly critical and is a mistake a lot of junior people make. There are ALWAYS mistakes and are ALWAYS things to poke holes in. But that's not the point of publishing. The point of publishing is to communicate to your peers. So in that respect being overly critical is a hindrance to scientific progress. It's because the context matters. This actually ties into my other comment in the thread fwiw. Just make sure to see things for what they are (it's why I basically ignore articles that talk about papers now, because they do the same thing but usually vastly over exaggerating the conclusions.)