8 ms·
The article hints at this, but not publishing null results (at least in a database - somewhere!) goes hand-in-hand with the replication crisis. An experimental
by glial 2y ago
The article hints at this, but not publishing null results (at least in a database - somewhere!) goes hand-in-hand with the replication crisis. An experimental outcome is always a single sample from a distribution of outcomes that you would obtain if you repeated the experiment many times.
Choosing to only publish the most extreme positive values means that when the experiment is replicated, "regression to the mean" makes it very likely that the measured effect will be weaker, and possibly not statistically significant. This is not an evidence of scientific fraud -- rather, it is a predictable outcome of a publishing incentive scheme that rewards hype and novelty over robust science.
I've said it before but it bears repeating - replicating published results, and adding the findings to a database, should be a standard part of PhD training programs.
- setopt 2y agoRelevant XKCD: https://xkcd.com/882/ https://xkcd.com/882/ Do 20 experiments with a p<5% criterion, and it’s likely that one will be a false positive. Only publish positive results, and someone will eventually publish a false positive result without fraud.
- etrautmann 2y agoVery few studies report a single statistical test as the sole conclusion. Most papers should assess some outcome in multiple ways using complementary data, multiple analyses etc. not always of course, but there are lots of ways of making sure your conclusions a robust without relying on a single analysis result.
- kerkeslager 2y agoThat doesn't fix the problem at all. No matter how many statistical tests you run on a sample, you can't get around the fact that the sample may not be representative of the population or the underlying phenomenon. You need different samples. There isn't a statistical trick that gets around this. For example: let's say there's a cancer with 20% survival rate. You test a treatment with 25 experimental and 25 control patients, 40% in the experimental group survive[1]. You can analyze this with a bunch of statistical methods. You can ask different questions about the patients, focusing on well-being rather than simple carcinogenic remission. But ultimately, the thing that happened in this study is that 50% got better and no fiddling with numbers or changing the questions you ask is going to change that underlying phenomenon. You can check for blood markers of cancer: you get 50% have no blood markers. You ask them questions about how they feel: you get 50% feel better. You body scan the area where tumors were: you get 50% no longer have tumors. You have only tested one phenomenon in one sample, and that essentially amounts to 5 people getting better. [1] I know this is not how cancer treatment studies work exactly, this is a simplified hypothetical.
- etrautmann 2y agoThat’s not what I’m saying. Obviously you can’t run multiple tests on one set of data. If you have a hypothesis, test it on multiple data sets in multiple ways and find supporting evidence rather than finding a single p=.05 and writing a story around it.
- kerkeslager 2y agoOkay, if that's what you're saying, then I don't understand why it sounds like you're disagreeing with the poster you're responding to.
- nathell 2y ago> Do 20 experiments with a p<5% criterion, and it’s likely that one will be a false positive. That would be true if p were the probability of null hypothesis being true given the data observed, but that’s not what it is.
- lucianbr 2y agoI really have no clue what p is, but I also really believe Randall Munroe does. Not that he's above making mistakes, but come on.
- mtts 2y agop is the chance that you would have gotten this result if the null hypothesis were true.
- stonemetal12 2y ago>For example, if 1,000,000 tests are carried out, then 5% of them (that is, 50,000 tests) are expected to lead to p < 0.05 by chance when the null hypothesis is actually true for all these tests. https://eurradiolexp.springeropen.com/articles/10.1186/s41747-020-0145-y#:~:text=Keeping%20the%20significance%20threshold%20at,true%20for%20all%20these%20tests https://eurradiolexp.springeropen.com/articles/10.1186/s4174....
- kerkeslager 2y agoEven if you are right, how would this response be helpful? You're not giving the right answer, you're just saying the answer given is wrong. Nobody is coming out of this interaction with corrected knowledge in their head. This is the sort of thing I used to do when I was younger, and looking back, the reason I did it was because I was basing my sense of self-worth in being smarter than other people. Ironically, this made me dumber, because I was less open to the possibility of being wrong and therefore was slower to learn. And, I found doing this made people dislike me.
- deleted 2y ago[deleted]
- humansareok1 2y ago>someone will eventually publish a false positive result without fraud. I think most people correctly intuit that this is actually a type of very pernicious fraud.
- whack 2y agoIf 20 different people all conduct the same experiment, and the 19 negative results are never published, there is no fraud involved when the 20th person publishes his positive result without realizing it is a 5% statistical anomaly. That person probably has no idea that 19 other people tried and failed at near-identical experiments. This seems like such an obvious problem in the way science is currently done. Are people so focused on their own individual fields that they aren't thinking about and fixing such glaring meta problems?
- prerok 2y agoAs others already pointed out, there is no incentive to do so. Consider: "Hey, look, I went on top of tower of Pisa and threw down two identically shaped balls, one iron and one wooden. They dropped at the same time!" The above is the expected result and would only be interesting if the result is different from expectation. Now, if 1000 scientists did this and each published the confirmation of what we knew would happen then who would read that? But, if one scientist said: "I tried and they drop at different times!" that would be different. The 999 scientists would then try to replicate again and then the papers of the 999 would be interesting again.
- humansareok1 2y agoThis is completely different from one lab running the same experiment 20 times and publishing the one positive result.
- some_random 2y agoIt's a great xkcd but it's really wrong on one count, it's not some outside media/popular force compelling scientists to investigate something. Most of the time, researchers are looking to prove something they "know" to be true. They truly believe that jelly beans cause acne, they just need to prove it. When they get a negative result, they simply don't believe it. Something must have gone wrong, obviously, because jelly beans obviously cause acne, so maybe it's a color thing? Ah hah it's the green ones, now that we have our results we can construct other metrics to support this correct data! Eventually if we (the public) are lucky someone else in the field will disagree and run the trial again, which is how you get the alt text.
- aeternum 2y agoYes, I find that when reading a paper I think to myself "Do the authors really want this to be true?" and if the answer is yes as it often is, I boost my own acceptance criteria to p<.03 Particle physics still uses five sigma as the significance threshold.
- mistermann 2y ago> This is not an evidence of scientific fraud... In its scriptures/philosophy, science describes extremely thorough and sound principles and guidelines...but in on the ground practice (by scientists, which are a part of "science"), they are often not achieved[1]. However, this distinction is not only not advertised broadly and without aversion, it is usually (in my experience) not mentioned at all, if not outright denied using persuasive rhetorical language (like, for example, when an object level instance of not achieving it is pointed to in the wild, such as in forum conversations). This may not be fraud (that requires intent I think?), but it achieves the same end: misinforming people. I absolutely agree with your database idea, and if science would like me to take them seriously (something near how seriously they take themselves) they'd also have to go much further. [1] Not unlike in religion, a competing metaphysical framework (model of reality) to science.
- 3np 2y ago> Not unlike in religion, a competing metaphysical framework (model of reality) to science. No. Correlation fallacy.
- mistermann 2y agoFallacy fallacy. Naive Realism fallacy.
- kerkeslager 2y agoThis sort of comment is why I think a lot of philosophy is just communicating poorly to make yourself sound smart. In your footnote, for example, you translated your philosophy-speak into English (metaphysical framework -> model of reality). Why not just say that? Your entire comment goes into "philosophy mode" and communicates a few very simple ideas in overcomplicated language. Science and religion are pretty poorly understood as competing models of reality. Religion originates when people make up answers to other people's questions to gain social standing, and religion continues due to (among other things) anchoring bias--the bias people have toward continuing to believe what they already believe. While religion does result in those people having a model of reality, there is no attempt being made at any point to relate the model to reality. When religious people and scientists disagree, it's not because the religious person is trying to model reality differently--the religious person isn't even trying to model reality--it's because the religious person is biased in favor of their existing belief. You said: > In its scriptures/philosophy, science describes extremely thorough and sound principles and guidelines...but in on the ground practice (by scientists, which are a part of "science"), they are often not achieved[1]. This is presented as some sort of gotcha, but it's not: few scientists will claim that science is being practiced perfectly or even well. Outside of a few areas such as particle physics, we're quite aware that our ability to practice scientific ideals is hampered by funding, publication incentives, availability of test subjects in human studies, data privacy, etc. And we're aware that this means that our conclusions need to be understood as probabilities rather than 100%-confidence facts. There are certainly some people who treat scientific conclusions with religious absolute confidence, but doing that is fundamentally against scientific principles. The accusation you are leveling against science would be better targeted toward people: generally science journalists and the science-illiterate public rather than scientists themselves. The entire reproducibility crisis is scientists using science to show that our practice of science is too imperfect to result in high-confidence conclusions. Religious people jumping on the replication crisis because they think it disproves science is rich. The replication crisis isn't a disproof of science, it's an application of science. The reason we know that there's a replication crisis is because scientists asked "How confident can we be in the conclusions of existing studies?" and applied science to answer that question. If you really think science is invalid, then you can't use science to prove that. And the fact remains that any confidence in conclusions at all is more than religion has to offer, because again, religion isn't trying to model reality--the fact that religion produces a model of reality is merely an unfortunate side-effect.
- hcks 2y ago“We thought academia was not soul crushing enough so from now on you will additionally spend 10 hours a week replicating dumb papers from 1993”
- bumby 2y agoSnark aside, replication is a cornerstone of science. If someone doesn't want to be involved in science because they think it's soul-crushing, perhaps academia isn't the right place for them.
- matthewdgreen 2y agoThere is a major issue of limited resources to replicate results, in terms of both time and funding. For example: I would assume that most important results are replicated. As a concrete example, if someone identifies a medication that (in one small trial) shows a statistically significant effect in curing some serious medical condition, then this will drive further replication attempts. On the other hand, if someone publishes a study showing that holding a pen in your mouth makes you 1% likelier to do well on the PSATs, this study will probably languish without replication for a decade — because honestly who cares? It’s basically a curiosity. I can’t help but notice that many of the headline results that characterize the “replication crisis” were small-effect-size social science experiments that fundamentally weren’t that important outside of popular science news. I’m not saying that our current allocation of resources is optimal. I am pointing out that our resources are finite and “replicate everything” is not even a remotely practical allocation of those resources.
- nostrademons 2y ago> I would assume that most important results are replicated. GP is pointing out that the incentive structure makes this an invalid assumption. If publications reward hype and novelty when deciding what to publish, then there is no point spending your limited resources replicating other peoples' results, they won't get published anyway. And experiments that give a null result won't be published anyway. What's left are one-off results that showed something surprising simply by chance and don't replicate...but then, we generally will never know that they don't replicate, because the replication experiment is not novel, has a low chance of being published, and hence isn't worth spending limited resources on. Basically the publication process introduces selection bias into the types of research that are even attempted, which then filters down into the conclusions we take from it. A cornerstone of the scientific method is random sampling, but as long as the results that get disseminated are chosen by a non-random process, it introduces bias.
- freestyle24147 2y ago> For example: I would assume that most important results are replicated. The example you provide is solely your assumption? Seems pretty odd to provide a baseless assumption as an "example".
- YeGoblynQueenne 2y ago>> I've said it before but it bears repeating - replicating published results, and adding the findings to a database, should be a standard part of PhD training programs. Wait, why should PhD students do that work? That just sounds like pushing more grunt work to the lower rung of the academic hierarchy. Nope. If you want people to do that kind of work that is important to everyone but is not directly conducive to promoting one's research career then the solution is simple: pay them.
- bumby 2y ago>why should PhD students do that work? I think there is some reasonable argument that replicating research is the first step to learning how to do good research on your own. In an ideal world, PhD students should probably be trying to replicate similar work anyway and applying existing approaches own pet problem. In practice, many gloss over this because they are narrowly focused on doing something "new" so it can get published.
- glial 2y agoPhD students are in training, and replicating a published result is a great training exercise. PhD students ARE paid. But this work won't be prioritized by their PI unless it's also a requirement of the program.
- YeGoblynQueenne 2y ago>> PhD students are in training, and replicating a published result is a great training exercise. I just got my PhD last July and that sounds like boring drudgery rather than "good exercise". Good exercise is to have a student write their own paper, dig up their own references, formulate their own experimental hypotheses, run their own experiments, write up their own results etc. Re-doing someone else's possibly badly-done work is grunt work that should not be forced upon anyone. >> PhD students ARE paid. Haha. Good one XD
- anticensor 2y agoNot all doctoral students are paid as part of a research position.
- deleted 2y ago[deleted]
- aeternum 2y agoIt's kind of amazing that we discovered the scientific method, used it to invent the transistor and bring the information revolution. Yet we still pool scientific results using only the printing press. It's like we unlocked the tech tree but then got so caught up in chasing citations and peer review that we forgot to use the new tech we invented.
- glial 2y agoYes, so-called "social technology" sometimes doesn't feel very advanced.
- Gormo 2y agoRichard Feynman was complaining about exactly this phenomenon fifty years ago: https://sites.cs.ucsb.edu/~ravenben/cargocult.html https://sites.cs.ucsb.edu/~ravenben/cargocult.html
- naasking 2y ago> This is not an evidence of scientific fraud -- rather, it is a predictable outcome of a publishing incentive scheme that rewards hype and novelty over robust science. And knowing all this, this behavy should be considered borderline scientific fraud at this point.