14 ms·
No publication without confirmation
- deleted 10y ago[deleted]
- dorianm 10y agoI applaud for the P < 0.01. There are too many non reproducible results with real life harm: http://infoproc.blogspot.com/2017/02/perverse-incentives-and-replication-in.html http://infoproc.blogspot.com/2017/02/perverse-incentives-and...
- feral 10y agoJust to note, there's a tradeoff here - not publishing work until you are massively certain of it would also cause real life harm. Reducing the p value doesn't automatically reduce harm. Physicists require extremely low values before confirming a discovery has been made, but that's different from requiring it before publishing. The problem is with people interpreting published work as if once its published, its completely certain. Maybe each publication should come with a headline 'confidence' stat beside the title. I guess this is a step in that direction.
- mattkrause 10y agoYou really think so? I'd argue the fetishization of a specific p-value threshold, be it p<0.05, p<0.01, or even p<0.0001, is a much bigger problem. There is an excellent quote from Rosnow and Rosenthal: [D]ichotomous significance testing has no ontological basis. That is, we want to underscore that, surely, God loves the .06 nearly as much as the .05. Can there be any doubt that God views the strength of evidence for or against the null as a fairly continuous function of the magnitude of p?” Wouldn't you prefer a few correct but tenative studies that hint at an effect while paving the way for larger, more expensive replications to a scenario where the data is sliced, diced, and tortured to hit some arbitrary p-value threshold?
- untilHellbanned 10y agoNo thanks. Papers involving animals are already backbreakingly slow compared with cell-based or in vitro work. I know because I've been lapped by my colleagues using more simple systems as I slog through our paper we got rejected from Nature because the reviewers suggested another 3 years worth of experiments. Yep, year 5 into this single project, which we knew the outcome for 4 years ago. Not excited about this proposal at all. Look I'm all for rigor but how about the people trying to make money off the deal pay for all the work and keep people like me out of it. Or don't allow the people trying to make money interpret the results of such preliminary studies so liberally. It's like the education system. Scientists like teachers, both of which don't make much money and do all the labor, don't want more hoops jump through.
- jessriedel 10y agoSorry, which parts of the proposal are you responding to? The article is more specific than "more rigor". Are you objecting to the higher p-value threshold? The independent confirmation? The author argues that a single higher quality confirmatory experiment will be able to replace gathering lots of statistics for exploratory experiments: > Unlike clinical studies, most preclinical research papers describe a long chain of experiments, all incrementally building support for the same hypothesis. Such papers often include more than a dozen separate in vitro and animal experiments, with each one required to reach statistical significance. We argue that, as long as there is a final, impeccable study that confirms the hypothesis, the earlier experiments in this chain do not need to be held to the same rigid statistical standard. Do you disagree?
- untilHellbanned 10y ago> one that incorporates an independent, statistically rigorous confirmation of a researcher's central hypothesis. We call this large confirmatory study a preclinical trial. These would be more formal and rigorous than the typical preclinical testing conducted in academic labs, and would adopt many practices of a clinical trial. As you can see above and from your quotation (and like many other folks who come in to save the day), this article is heavy on plans and short on who is going to do the work. Of course I support papers where every single experiment doesn't have to play p < 0.05 games, but other parts of the article wander in other directions. That's all I'm reacting to.
- StClaire 10y agoI have an idea: if a research study doesn't go the way you thought it would, put it out there. We need a central repository like Arxiv where we dump the experiments that didn't work out so that we can quickly compare a "successful" one to ones done before. That gives us a better idea of if the data is just a fluke. The papers wouldn't have to be super involved. What did you do? What were statistical conclusions. Give an upper-level undergraduate or an early masters student some experience writing up a procedure. Shouldn't take more than a couple hours but it could save a lot of time dealing with publication bias
- jonlucc 10y agoBioRxiv [1] attempts to be this, but it isn't widely used as it is in the physics/math/compSci world. [1] http://biorxiv.org/ http://biorxiv.org/
- untilHellbanned 10y agoNot going to happen. You don't understand the forces at play when biomedical researchers (I'm one too) do the work. A major reason more specific to this type of researcher that he/she wants to hold back the data is in the event it can be re-purposed for later goals. It's capital. The time and money spent acquiring the types of negative data the article talks about (mouse experiments) take big dollars. I'm not going to throw it away in an archive. It's like donating to a Goodwill bin some really expensive clothes you saved your money to buy only to realize after you bought it that it's the wrong size. You're totally going to try to make it work often for years before you give up on it. If you take issue with my comment and have ever worked for a startup that didn't pan out (should be many of us here on HN), think of the scenario if someone told you, "hey, why don't you just be a good person and open source your app and give all your customers to XYZ?" (Setting aside the armchair quarterback guilt imposed on you) There are a few who do that, but if you've spent any length of time on your business, you're going to think of ways to re-purpose the investment you made it in other ways before you just go dumping it in an archive like Github or whatever.
- alexryan 10y agoThere seems to be a lot of sharing amongst artificial intelligence researchers nowadays and it seems to be accelerating innovation. Sharing of medical research could potential accelerate innovation that could save a lot of lives (including those of the researchers themselves and their loved ones). It's curious that these industries are so different in this respect. What would have to happen in order to increase the amount of sharing of medical research?
- deleted 10y ago[deleted]
- jeffdavis 10y agoDumb outsider question: why not just mark studies that have been reproduced versus ones that still have not been reproduced? The way I view it is two steps of publication: the cutting-edge (not reproduced yet) versus independently reproduced.
- untilHellbanned 10y agoFiguring out which is which is easy enough. The real issue is people not over-interpreting the results.
- jeffdavis 10y agoMaybe it could score labs based on how many of the studies have been reproduced successfully? If something is not reproduced, and submitted by a lab with a low score, people would take it less seriously.
- tdaltonc 10y agoThe obvious questions is "who's going to do the confirmation work?" I think that masters/bachelor students should be able to handle that work. A new grant mechanism for masters/bachelor training grants that fund replication would get the job done with a lot of nice side effects.
- untilHellbanned 10y agoAs a middle american, I'm actually serious when I say this could be a way middle america gets on it's feet again. Being a biomedical researcher at a university hospital, the hospital has proliferated with all types of trainees. What's holding back the basic research side? It's an assembly line in the same way the midwest is familiar, the automobile industry. However, universities aren't good businesses, and what's missing is big pharma companies paying employees fair wages to do this work. Though companies saving the day won't work either because putting people to work we'll mean they won't get sick which is bad for business. So oh well on to the next approach.
- tdaltonc 10y agoThat is an interesting idea. I lot of biomedical research is very repudiative. There are people working on automating it. https://emeraldcloudlab.com/ https://emeraldcloudlab.com/ It would be hard to know if the man on the line is doing it right, but you might get the benefit of accidental insights.
- pcrh 10y agoThe article refers chiefly to repeating mouse or rat experiments. The obstacle there isn't what level of training researcher has (as long as it's sufficient), but who is going to pay for it. At the scale proposed (a 6-fold greater number of mice per experiment than is usual) the cost of testing only the core hypothesis is easily over $100K. In addition there is the time involved, which can be from months to years, depending on the experiment.
- mattkrause 10y agoScaling up non-human primate experiment like this could make them span a decade, assuming the infrastructure and staff aren't also increased sixfold.
- bloaf 10y agoI don't think this is a good idea because it would increase the politicking in scientific publication. Specifically, no one is going to want to do the reproduction work, so reproduction work will be seen as a favor from one scientist to another. Moreover, in specialized fields, scientists just as frequently see each other as competitors as collaborators. I strongly suspect there would be a lot of gamesmanship where scientists refuse to (or drag their feet) do reproduction work on new studies that threaten to disrupt the status quo that has made them successful.
- Fomite 10y agoPretty much this. I cannot help but see a replication requirement like this turning things even more political. Young New Investigator's Lab, who has very little to trade, is going to struggle, while someone who can conjure a postdoc out of thin air for someone's student will probably be able to find someone.
- disgruntledphd2 10y agoI would absolutely kill to get a job doing replication. What I always hated about science was the inability for things not to work out. Even if you find something directly opposed to your hypothesis, you are somehow supposed to pretend that it worked out "just as planned". It's toxic, boring and leads to bad science. And so, for me, I would absolutely adore to be in a place where I got to run well-powered studies and aim to just figure out the right answer rather than build my career on a bunch of unrepeatable statistical flukes. That being said, my PhD is in Psychology, so they probably won't be hiring me to run animal-model studies. I really like this idea, as long as Nature put their space where their mouth is (which they won't, as they have at least one of these articles per year and it doesn't appear to have made any impact).
- rdlecler1 10y agoThe problem here is that publishing a single paper is often the product of months if not years of work. Saying: now add more work without more grant money, is going to be difficult to swallow. Even worse, it departments need to hire even more PhDs who are unemployable after they graduate.
- caseysoftware 10y agoI think we've conflated terms. The lay public thinks "peer reviewed" means that others have tried it and validated the results. What it really tends to mean is that a peer looked at the procedures and results and that it passes the "sniff test" and generally doesn't have any glaring errors. The more subtle problem is that in some circles, it isn't even that. Since fewer and fewer people want to be the person who damaged someone else's work and/or career, it's a blanket pass. We're drifting away from scientific study and critical thinking to "reasonable" approaches and not upsetting doctrine and/or your superiors. That looks less and less like science and more like religion.
- hackuser 10y agoHere's a leading scientist's description of peer review: Peer review works superbly to separate valid science from nonsense, or, in [Thomas] Kuhnian terms, to ensure that the current paradigm has been respected. It works less well as a means of choosing between competing valid ideas, in part because the peer doing the reviewing is often a competitor for the same resources ... sought by the authors. It works very poorly in catching cheating or fraud, because all scientists are socialized to believe that even their toughest competitor is rigorously honest in the reporting of scientific results ... It certainly does not ensure that the work has been fully vetted in terms of the data analysis and the proper application of research methods. From: Reference Manual on Scientific Evidence [for U.S. federal judges], Third Edition; How Science Works section by David Goodstein, CalTech Physics Professor and former Provost; published by National Academies Press (2011) https://www.nap.edu/catalog/13163/reference-manual-on-scientific-evidence-third-edition https://www.nap.edu/catalog/13163/reference-manual-on-scient...
- T-A 10y agoAlso by Goodstein [1]: Peer review is usually quite a good way to identify valid science. Of course, a referee will occasionally fail to appreciate a truly visionary or revolutionary idea, but by and large, peer review works pretty well so long as scientific validity is the only issue at stake. However, it is not at all suited to arbitrate an intense competition for research funds or for editorial space in prestigious journals. There are many reasons for this, not the least being the fact that the referees have an obvious conflict of interest, since they are themselves competitors for the same resources. This point seems to be another one of those relativistic anomalies, obvious to any outside observer, but invisible to those of us who are falling into the black hole. It would take impossibly high ethical standards for referees to avoid taking advantage of their privileged anonymity to advance their own interests, but as time goes on, more and more referees have their ethical standards eroded as a consequence of having themselves been victimized by unfair reviews when they were authors. Peer review is thus one among many examples of practices that were well suited to the time of exponential expansion, but will become increasingly dysfunctional in the difficult future we face. [1] http://www.its.caltech.edu/~dg/crunch_art.html http://www.its.caltech.edu/~dg/crunch_art.html
- ramblenode 10y agoPre-registration is nice, larger samples/greater power is necessary, and increasing the p-threshold may indirectly filter out some false positives but kind of misses the underlying issue of p-hacking, some of which would be solved by pre-registration. The authors suggestions are preventative in nature but what I would like to see above all else is requiring researchers to publish the raw data and to make their statistical analyses minimally reproducible--something which could be satisfied by publishing scripts or Excel macros along with instructions for any non-automated data stitching. Experiments frequently implode at the analysis phase which then gets intentionally or unintentionally masked in ambiguous, poorly written methods sections. Giving others access to the data allows errors to be spotted earlier after publication and alternative hypotheses and analyses to be tested against the published results. It's also sometimes the only way of spotting abnormalities resulting from the data collection process itself. Again, not a means of preventing errors, but a low friction way of discovering them. Maybe having everything in the open would light a fire under some researchers to be more thorough, though.
- Ranlot 10y agoA simple discussion to find out more about the "statistical power" mentioned in the article: https://p-value-convergence.herokuapp.com/ https://p-value-convergence.herokuapp.com/
- lngnmn 10y agoAgain, statistics applied to a partially observable and partially understood phenomena yield nonsense. If not all the variables are controlled or not all possible causes has been taken into account the result will be a mere aggregation of observations. What's true for coins and dices does not applicable for partially observable environments with multiple causation and yet unknown control mechanisms. Statistics is not applicable to imaginary models based on unproven assumptions or premises.
- lutusp 10y agoQuote: "Our proposal is a new type of paper for animal studies of disease therapies or preventions: one that incorporates an independent, statistically rigorous confirmation of a researcher's central hypothesis." This probably won't happen right away, but it's a terrific and necessary idea that we need to to move forward. It will revolutionize biology and medicine, and it will end the field of psychology as we know it. http://arachnoid.com/psychology_and_alchemy http://arachnoid.com/psychology_and_alchemy
- mattkrause 10y agoWhy would this end psychology? It might make it better, but we're nowhere near being able to describe (e.g.) group behavior from first principles or ion channel kinetics. Despite your link, there is a lot of solid psych research. There are (obviously) discredited theories and cranks too, but psychologists characterized rods and cones way before the biologists found them, for one example.
- lutusp 10y ago> Despite your link, there is a lot of solid psych research. Yes, solid, but lacking the dimension of falsifiable theories about the mind. Tax accounting is also solid research. > ... but psychologists characterized rods and cones way before the biologists found them, for one example. Those weren't psychological studies. Psychology is study of the mind and behavior. Rods and cones are neither. When a psychologist studies something biological, it's not psychology any more.
- mattkrause 10y agoReally? Psychophysics and perception research is widely considered to be part of psychology, and has strong, falsifiable theories about how sensory stimuli are encoded and processed. Using purely behavioural methods, psychophysicists figured out that there were three color-sensitive "sensors" and pinned down their properties. I'm not sure it suddenly becomes biology because someone later found the cellular substrate, nor did it become chemistry when someone figured out the structure of opsin. Likewise, I'd argue that a lot of the learning stuff (e.g., reinforcement learning) also describes the mind's operation in testable and falsifiable ways.
- beloch 10y ago"Confirmatory labs would be less dependent on positive results than the original researchers, a situation that should promote the publication of null and negative results. They would be rewarded by authorship on published papers, service fees, or both. They would also be more motivated to build a reputation for quality and competence than to achieve a particular finding." Sounds great, but how would this actually work. Nobody is going to get juicy grants from existing funding agencies for being a "confirmatory" lab. Nature sure as hell isn't going to pay for this. Most researchers probably can't afford to pay an outside lab to duplicate their research. Is Nature going to suddenly start refusing papers whose results haven't been reproduced elsewhere? That's basically suicide for their journal because researchers are frequently in a race with other researchers to publish first, so why publish with a journal that requires you to double your budget to pay a confirmatory lab and wait months or years for them to do the job? The pressure will be intense to publish elsewhere first. I have a simpler solution. Don't just slap the names of confirmatory lab authors onto other papers. Publish original papers and publish confirmatory papers with equal prominence to the original papers. Hell, devote a portion of Nature to doing just that. Currently, if you want to publish a paper about confirming someone else's original findings, not even a third rate journal will touch it unless you put at least some kind of novel-sounding spin on it. Nature should use all that scummy impact factor gaming they do to make confirmatory papers respectable. Only when the work of reproducing results gains labs respect will funding agencies start supporting "confirmation labs". At present, such "unoriginal", "hack" work is not respected at all, and Nature is a big part of the reason why.
- setrofim_ 10y ago> Most researchers probably can't afford to pay an outside lab to duplicate their research. Even if they could, we probably don't want the researchers paying for their results to be duplicated. This would create perverse incentives, similar to what happened with investment banks and credit rating agencies. If the original researchers must get their results confirmed in order to get published, and it is them who are paying for the confirmation, they will naturally tend to choose confirmatory labs that are more likely to confirm their findings. Since the labs would then rely on the researchers for funding, that would create pressure on the confirmatory labs to adapt their methodologies in ways that make it more likely that results get confirmed (even when the original study may not warrant it). We want confirmatory labs to have no special interest in either confirming or disproving a particular study, but in improving the overall quality of research. Since a journal's reputation depends (at least in part) on the quality of research it publishes, the journals would seem to be the natural candidates for the source of funding of confirmatory labs. Whether they'd actually be willing to do it another matter...
- misnome 10y agoHow about always having a professional, non-field statistician on the review panel? Non-reproduceability should probably be interpreted as criticism of the reported certainly of the results.
- clamprecht 10y agoIt's interesting the parallel between the blockchain (requiring confirmations by peers) and this discussion.
- fiatjaf 10y agoWe need less published stuff, much less. Much much less.
- deleted 10y ago[deleted]