3 ms·
Tragically, I bet this is all too common. Wansink's mistake was blogging about it. This emblematic of a larger problem with how science is practiced: the obse
by tryitnow 9y ago
Tragically, I bet this is all too common. Wansink's mistake was blogging about it.
This emblematic of a larger problem with how science is practiced: the obsessive focus on p-value thresholds leads to irrational practices like trawling data for interesting "findings."
But on a certain level Wansink was right: a data set is not completely worthless if it showed a null result. So we need to start thinking about how to communicate the value of data even when the null is not rejected.
One way to do this is to encourage widespread sharing of data sets regardless of the outcome of the experiment. Maybe for a given study the data did not show a definitive result - but does the data point to potential future paths of study? Maybe another researcher could get ideas for new experiments.
- StavrosK 9y ago> a data set is not completely worthless if it showed a null result. Why is a null result worthless? I don't understand. It goes against my common sense that paying more for food doesn't make you eat more. The fact that this is not the case is valuable knowledge to me. I understand that it's not as glamorous as "paying more makes you eat less!", but it's still valuable knowledge that should be published.
- draw_down 9y agoI think you're saying the same thing as GP
- ad-hominem 9y ago> a data set is not completely worthless if it showed a null result. A null result is not worthless. A data set is not worthless if it showed a null result. Nobody is claiming a null result is worthless. Even the original article gives a good reason for publishing a null result.
- KingMob 9y agoNull results are "worthless" only in the sense that they're harder, if not impossible, to publish, and do less for your career. As a former neuroscientist, I had a unicorn data set of intracranial EEG data, and we'd spent so much time collecting it, that we were determined to find something. I left grad school before we found anything of interest. I believe my former prof eventually published something on it, but I analyzed the shit out of it, even knowing I was fishing, because so much time would have been lost to not use it. Fishing like this was one of the reasons I left. The pursuit of knowledge and the pursuit of your career don't align often enough. To me, what's unusual is not that Wansink pushed to keep analyzing the data, but that he was called out for it. Everyone seemed to be doing it when I was in academia.
- noobhacker 9y agoI feel the same way about current practice research, but find that industry standards are even more abysmal in terms of fishing. What field are you in now that you find an acceptable level of rigor?
- KingMob 9y agoHah, well, I returned to software dev, so I'm not too worried about research rigor these days.
- bmm6o 9y agoEven the post-hoc analysis isn't invalid as a way of discovering potential avenues of future research. But you really have to run new experiments designed to test your new hypotheses.
- appleflaxen 9y agoYou can make multiple comparisons (or post-hoc "exploratory" comparisons); you just need to make your p-value threshold sufficiently strict [1]. 1. https://en.wikipedia.org/wiki/Bonferroni_correction https://en.wikipedia.org/wiki/Bonferroni_correction
- andrewla 9y ago> a data set is not completely worthless if it showed a null result As the other commenters have pointed out, the fact that it showed a null result is not really relevant to future utility. A better way of phrasing it might be "A data set is not completely worthless after it has been used to the test the hypotheses for which it was gathered." As to how it is useful, I think a sound summary would be "a data set is _only_ useful for validating (or invalidating) those hypotheses for which it is gathered, but it may be useful for deriving new hypotheses for which new data can be gathered to validate." Or more negatively as "a data set is not useful for validating or invalidating hypotheses invented after the data set is seen, but it may be useful for deriving those hypotheses and leading to further research." To some very limited extent, if a mechanism of action is established (which is to say there is an external reason to infer a correlation) then cross-validation with an existing data set may be able to derive the parameters of that correlation.
- Uhhrrr 9y agoIt seems like a nothing to me. He didn't say, "publish any thing with p>0.05," but rather, "see if there's something interesting."