25 ms·
Replace peer review with “peer replication” (2021)
- moelf 3y agoI wish we can replicate the LHC
- gxt 3y agoSo why haven't "science modules" been developed yet? I see a library sized piece of equipment to physically perform the lab work that can be configured akin to CNC machining. Papers would then be submitted with the module program and be easily replicated by other labs.
- katrinaroberta 3y ago[dead]
- camilamacedo 3y ago[dead]
- eesmith 3y ago> the real test of a paper should be the ability to reproduce its findings in the real world. ... > What if all the experiments in the paper are too complicated to replicate? Then you can submit to [the Journal of Irreproducible Results]. Observational science is still a branch of science even if it's difficult or impossible to replicate. Consider the first photographs of a live giant squid in its natural habitat, published in 2005 at https://royalsocietypublishing.org/doi/10.1098/rspb.2005.3158 https://royalsocietypublishing.org/doi/10.1098/rspb.2005.315... . Who seriously thinks this shouldn't have been published until someone else had been able to replicate the result? Who thinks the results of a drug trial can't be published until they are replicated? How does one replicate "A stellar occultation by (486958) 2014 MU69: results from the 2017 July 17 portable telescope campaign" at https://ui.adsabs.harvard.edu/abs/2017DPS....4950403Z/abstract https://ui.adsabs.harvard.edu/abs/2017DPS....4950403Z/abstra... which required the precise alignment of a star, the trans-Neptunian object 486958 Arrokoth, and a region in Argentina? Or replicate the results of the flyby of Pluto, or flying a helicopter on Mars? Here's a paper I learned about from "In The Pipeline"; "Insights from a laboratory fire" at https://www.nature.com/articles/s41557-023-01254-6 https://www.nature.com/articles/s41557-023-01254-6 . """Fires are relatively common yet underreported occurrences in chemical laboratories, but their consequences can be devastating. Here we describe our first-hand experience of a savage laboratory fire, highlighting the detrimental effects that it had on the research group and the lessons learned.""" How would peer replication be relevant?
- msla 3y agoWith some of the things, but admittedly not most of the things you mentioned, there's a dataset (somewhere) and some code run on that dataset (somewhere) and replication would mean someone else being able to run that code on that dataset and get the same results. Would this require labs to improve their software environments and learn some new tools? Would this require labs to give up whatever used to be secret sauce? That's. The. Point.
- counters 3y agoIn practice this is happening in many disciplines, for most research, on a daily basis. What _isn't_ happening is that the results of these replications are being independently peer reviewed, because that isn't incentivized. However, when replication fails for whatever reason, it usually leads to insights that themselves lead to stronger scientific work and better publications later on.
- eesmith 3y ago> someone else being able to run that code on that dataset and get the same results. I think when people talk about "replicate" they mean something more than that. The dataset could contain coding errors, and the analysis could contain incorrect formulas and bad modeling. Reproducing a bad analysis, successfully, provide no corrective feedback. I know for one paper I could replicate the paper's results using the paper's own analysis, but I couldn't replicate the paper's results using my analysis. > Would this require labs to give up whatever used to be secret sauce? That's. The. Point. That seems to be a very different Point. Newton famously published results made from using his secret sauce - calculus - by recasting them using more traditional methods. In the extreme cas, I could publish the factors for RSA-1024 without publishing my factorization method. "I prayed to God for the answer and He gave them to me." You can verify that result without the secret sauce. I mean, people use all sorts of methods to predict a protein structure, including manual tweaking guided by intuition and insight gained during a reverie or day-dream (à la Kekulé) which is clearly not reproducible. Yet that final model may be publishable, because it may provide new insight and testable predictions.
- nomilk 3y agohttps://web.archive.org/web/20230130143126/https://blog.everydayscientist.com/replace-peer-review-with-peer-replication/ https://web.archive.org/web/20230130143126/https://blog.ever...
- the_arun 3y agoThank you. Currently the original article is throttled. Seems like article is not about software code.
- miga 3y agoPeer review does not serve to assure replication, but assure readability and comprehensibility of the paper. Given that some experiments cost billions to conduct, it is impossible to implement "Peer Replication" for all papers. What could be done is to add metadata about papers that were replicated.
- NalNezumi 3y agoIsn't readability and comprehensibility the job of the editor/journal to check. (after all they're actually paid) maybe not for conference, but peer review is more for checking if the methodology, scope, claim, direction, conclusion and relevances is sound&trustable. At least that's my understanding
- kergonath 3y agoThe editor is often not the right person to decide based on technical details. Most often, articles they receive anre outside their field of expertise and they don’t really have a way of deciding if a section is comprehensible or not. It’s very difficult for an outsider to know what bit of jargon is redundant and what bit is actually important to make sense of the results. So this bit of readability check falls to the referees. In theory editors (or rather copyeditors, the editors themselves have to handle too many papers to do this sort of thing) should help with things like style, grammar, and spelling. In practice, quality varies but it is often subpar.
- kkylin 3y agoHighly dependent on journal / field. In mine (mathematics) most associate editors work for free, same as reviwers. The reviewer do all the things you say, and in addition try to ensure readability & novelty. Most journals do have professional copy editing, but that's separate from the content review. I don't know how refereed conference proceedings work (we don't really use these). The only journals I know of that have professional editors (i.e., editors who are not active researchers themselves) are Nature and affiiliated journals, but someone more knowledgeble should correct me here.
- hedora 3y ago
- leedrake5 3y agoPeer Review is the right solution to the wrong problem: https://open.substack.com/pub/experimentalhistory/p/science-is-a-strong-link-problem https://open.substack.com/pub/experimentalhistory/p/science-... On replication, it is a worthwhile goal but the career incentives need to be there. I think replicating studies should be a part of the curriculum in most programs - a step toward getting a PhD in lieu of one of the papers.
- vinnyvichy 3y agoFear of the frontier.. that's why instead of people getting excited to look for new rtsp superconductor candidates, we get a lot of talk downplaying the only known one. Strong link vs weak link reminds me of how some cultures frown on stimulants while other cultures frown on relaxants.
- NalNezumi 3y agoImo, A more realistic thing to do is "replicability review" and/or requirement to submit "methodology map" to each paper. The former would be a back and forth between a reviewer that inquire and ask questions (based on the paper) with the goal to reproduce the result, but don't have to actually reproduce it. This is usually good to find out missing details in the paper that the writer just took for granted everyone in the field knows (I've met Bio PHD that have wasted Months of their life tracking up experimental details not mentioned in a paper) The latter would be the result of the former. Instead of having pages long "appendix" section in the main paper, you produce another document with meticulous details of the experiment/methodology with every stone turned together with an peer reviewer. Stamp it with the peer reviewes name so they can't get away with hand wavy review. I've read too many papers where important information to reproduce the result is omitted. (for ML/RL) If the code is included I've countless of times found implementation details that is not mentioned in the paper. In matter of fact, there's even results suggesting that those details are the make or break of certain algorithms. [1] I've also seen breaking details only mentioned in code comments... Another atrocious thing I've witnessed is a paper claiming they evaluated their method on a benchmark and if you check the benchmark, the task they evaluated on doesn't exit! They forked the benchmark and made their own task without being clear about it! [2] Shit like this make me lose faith in certain science directions. And I've seen a couple of junior researcher giving it all up because they concluded it's all just house of cards. [1] https://arxiv.org/abs/2005.12729 https://arxiv.org/abs/2005.12729 [2] https://arxiv.org/abs/2202.02465 https://arxiv.org/abs/2202.02465 Edit: also if you think that's too tedious/costly, reminder that publishers rake in record profits so the resources are already there https://youtu.be/ukAkG6c_N4M https://youtu.be/ukAkG6c_N4M
- kergonath 3y ago> I've met Bio PHD that have wasted Months of their life tracking up experimental details not mentioned in a paper Same. Now, when I review manuscripts, I pay much more attention to whether there is enough information to replicate the experiment or simulation. We can put out a paper with wrong interpretations and that’s fine because other people will realise that when doing their own work. We cannot let papers get published if their results cannot be replicated. > The latter would be the result of the former. Instead of having pages long "appendix" section in the main paper, you produce another document with meticulous details of the experiment/methodology with every stone turned together with an peer reviewer. Stamp it with the peer reviewes name so they can't get away with hand wavy review Things that take too much space to go in the experimental section should go to a electronic supplementary information document. But then it would be nice if the ESI were appended to the article when we download a PDF because tracking them is a pain in the backside. Some fields are better than others about this, for example in materials characterisation studies it’s very common to have ESI with a whole bunch of data and details. Large dataset should go to a repository or a dataset journal, that way the method is still peer reviewed and the dataset has a doi and is much easier to re-use. It’s also a nice way of doubling a student’s papers count by the end of their PhD. > Another atrocious thing I've witnessed is a paper claiming they evaluated their method on a benchmark and if you check the benchmark, the task they evaluated on doesn't exit! They forked the benchmark and made their own task without being clear about it! [2] That’s just evil!
- jhart99 3y agoReplication in many fields comes with substantial costs. We are unlikely to see this strategy employed on many/most papers. I agree with other commenters that materials and methodology should be provided in sufficient detail so that others could replicate if desired.
- fodkodrasz 3y agoHow would you peer-replicate observation of a rare, or unique event, for example in astronomy?
- lordnacho 3y agoEither get your own telescope and gather your own data, or if only one telescope captured a fleeting event, take that data and see if the analysis turns out the same.
- fabian2k 3y agoI don't see how this could ever work, and non-scientists seem to often dramatically underestimate the amount of work it would be to replicate every published paper. This of course depends a lot on the specific field, but it can easily be months of effort to replicate a paper. You save some time compared to the original as you don't have to repeat the dead ends and you might receive some samples and can skip parts of the preparation that way. But properly replicating a paper will still be a lot of effort, especially when there are any issues and it doesn't work on the first try. Then you have to troubleshoot your experiments and make sure that no mistakes were made. That can add a lot of time to the process. This is also all work that doesn't benefit the scientists replicating the paper. It only costs them money and time. If someone cares enough about the work to build on it, they will replicate it anyway. And in that case they have a good incentive to spend the effort. If that works this will indirectly support the original paper even if the following papers don't specifically replicate the original results. Though this part is much more problematic if the following experiments fail, then this will likely remain entirely unpublished. But the solution here unfortunately isn't as simple as just publishing negative results, it take far more work to create a solid negative result than just trying the experiments and abandoning them if they're not promising.
- backtoyoujim 3y agoYes it would indeed mean slowing down and having more scientists. It would mean disruption is no longer a useful tool for human development.
- brnaftr361 3y agoIt may not be. I would be willing to argue that there was a tipping point and we've long exceeded its boundary - progress and disruption now is just making finding an equilibrium in the future increasingly difficult. So entering into a paradigm where we test the known space - especially presently - would 1) help reduce cruft; 2) abate undersirable forward progress; 3) train the next generation(s) of scientists to be more diligent and better custodians of the domain.
- ebiester 3y ago
- SubiculumCode 3y agoScientist publishes paper based on ABCD data. Replicator: Do you know how much data I'll need to collect? 11,000 particpants followed across multiple timepoints of MRI scanning. Show me the money.
- petesergeant 3y agoDefinitely something that needs large charitable investment, but charities like that do exist, eg Wellcome Trust
- SubiculumCode 3y agoLike 290+ million, just to get started.
- tines 3y ago"Replace peer code review with 'peer code testing.'" Probably not gonna catch on.
- infogulch 3y agoI like the idea of splitting "peer review" into two, and then having a citation threshold standard where a field agrees that a paper should be replicated after a certain number of citations. And journals should have a dedicated section for attempted replications. 1. Rebrand peer review as a "readability review" which is what reviewers tend to focus on today. 2. A "replicability statement", a separately published document where reviewers push authors to go into detail about the methodology and strategy used to perform the experiments, including specifics that someone outside of their specialty may not know. Credit NalNezumi ITT
- analog31 3y agoEvery experimental paper I've ever read has contained an "Experimental" section, where they provide the details on how they did it. Those sections tend to be general enough, albeit concise. In some fields, aside from specialized knowledge, good experimental work requires what we call "hands." For instance, handling air sensitive compounds, or anything in a condensed or crystalline state. In my thesis experiment, some of the equipment was hand made, by me. Sometimes specialized facilities are needed. My doctoral thesis project used roughly 1/2 million dollars of gear, and some of the equipment that I used was obsolete and unavailable by the time I finished.
- ahmadmijot 3y ago> My doctoral thesis project used roughly 1/2 million dollars of gear, Wow I envy you. My doctoral thesis project spent like... USD2.5k directly for gears (half of it just to buy lego bricks to build our own instrument exactly because we can't afford to buy commercial one lol)
- xioxox 3y agoI used a 3 billion dollar space telescope. I don't think NASA are going to launch another to replicate some of my results.
- janalsncm 3y ago“Concise” isn’t good enough. If other scientists are trying to read through the tea leaves at what you’re trying to say you did, that defeats the entire point of a paper. The purpose of science is to create knowledge that other people can use and if people can’t replicate your work that’s not science.
- abnry 3y agoIf scientists are going to complain that's its too hard or too expensive to replicate their studies, then that just shows their work is BS.
- alsodumb 3y agoNah, it doesn't. It just shows that it's time consuming and expensive to replicate their studies.
- abnry 3y agoIf that's the case, then don't claim confidence in the work or make policy decisions based off of it. If there is no epistemological humility, then yes, it is still BS.
- fodkodrasz 3y agoI guess if software developers will complain that it's too hard or too expensive to thoroughly test their code to ensure exactly zero bugs at release[1], then that just shows their work is BS. [1]: if you have delivered telco code to Softbank you may have heard this sentence
- abnry 3y agoReplication is not the same thing as zero bugs in software.
- gordian-not 3y agoThe incentive should be to clear the way for tenure track The junior faculty will clear the rotten apples at the top by finding flaws in their research and then will win the tenure that was lost in return This will create a nice political atmosphere and improve science
- elashri 3y agoGreat, but who is going to fund the peer replication?. The economics of research now doesn't even provide a compensation for peer review process time.
- nine_k 3y agoMaybe the numerous complaints about the crisis of science are somehow related to the fact that scientific work is severely underpaid. The pay difference between research and industry in many areas is not even funny.
- matthewdgreen 3y agoThe purpose of science publications is to share new results with other scientists, so others can build on or verify the correctness of the work. There has always been an element of “receiving credit” to this, but the communication aspect is what actually matters from the perspective of maximizing scientific progress. In the distant past, publication was an informal process that mostly involved mailing around letters, or for a major result, self-publishing a book. Eventually publishers began to devise formal journals for this purpose, and some of those journals began to receive more submissions than it was feasible to publish or verify just by reputation. Some of the more popular journals hit upon the idea of applying basic editorial standards to reject badly-written papers and obvious spam. Since the journal editors weren’t experts in all fields of science, they asked for volunteers to help with this process. That’s what peer review is. Eventually bureaucrats (inside and largely outside of the scientific community) demanded a technique for measuring the productivity of a scientist, so they could allocate budgets or promotions. They hit on the idea of using publications in a few prestigious journals as a metric, which turned a useful process (sharing results with other scientists) into [from an outsider perspective] a process of receiving “academic points”, where the publication of a result appears to be the end-goal and not just an intermediate point in the validation of a result. Still other outsiders, who misunderstand the entire process, are upset that intermediate results are sometimes incorrect. This confuses them, and they’re angry that the process sometimes assigns “points” to people who they perceive as undeserving. So instead of simply accepting that sharing results widely to maximize the chance of verification is the whole point of the publication process, or coming up with a better set of promotion metrics, they want to gum up the essential sharing process to make it much less efficient and reduce the fan-out degree and rate of publication. This whole mess seems like it could be handled a lot more intelligently.
- dmbche 3y agoYour analysis seems to portray all scientists as pure hearted. May I remind you of the latest Stanford scandal where the president of Stanford was found to have manipulated data? Today, publications do not serve the same purpose as they did before the internet. It is trivial today to write a convincing paper without research and getting that published(www.theatlantic.com/ideas/archive/2018/10/new-sokal-hoax/572212/&sa=U&ved=2ahUKEwjnp5mRtsiAAxVwF1kFHesBDC8QFnoECAkQAg&usg=AOvVaw0t_Bo31BrT5D9zHBdmNAqi).
- hgsgm 3y agoThe problem is equating publication with truth. Publication is a starting point, not a conclusion Publication is submitting your code. It still needs to be tested, rolled out, evaluated, and time-tested.
- j45 3y agoCan every thing be replicated in every field
- User23 3y agoThat’s the defining characteristic of engineering. If you can’t reliably replicate everything in an engineering discipline then it’s not an engineering discipline.
- seventytwo 3y agoThere would need to be an incentive structure where the first replications get (nearly) the same credit as the original publisher.
- waynecochran 3y agoI spent a lot of my graduate years in CS implementing the details of papers only to learn that, time and time again, the paper failed to mention all the short comings and fail cases of the techniques. There are great exceptions to this. Due to the pressure of "publish or die" there is very little honesty in research. Fortunately there are some who are transparent with their work. But for the most part, science is drowning in a sea of research that lacks transparency and replication short falls.
- cptskippy 3y agoYou'll quickly discover when you enter the workforce that the reasons we have CI/CD, Docker, and virtualization are because of a similar problem. The dread "it works on my machine" response. CI/CD forces people to codify exactly how to build and deploy something in order for it to get into a production environment. Docker and VMs are ways around this by giving people a "my machine" that can be copied and shared easily.
- janalsncm 3y agoI had a very similar experience in my masters. Really made me think, what exactly are the peers “reviewing” if they don’t even know whether the technique works in the first place.
- waynecochran 3y agoI have reviewed many papers and there is never the time to recreate the work and test. That is why I love the "papers w code" site. I think every published CS paper should require a git repo with all their code and experimental data.
- jimmar 3y agoHow do you replicate a literature review? Theoretical physics? A neuro case? Research that relies upon natural experiments? There are many types of research. Not all of them lend themselves to replication, but they can still contribute to our body of knowledge. Peer review is helpful in each of these instances. Science is a process. Peer review isn't perfect. Replication is important. But it doesn't seem like the author understands what it would take to simply replace peer review with replication.
- janalsncm 3y agoI don’t think the existence of papers that are difficult to replicate undermines the value of replicating those that are easier.
- ahmadmijot 3y agoQuite related: nowadays there is this movement within scientific researches ie Open Science where the (raw) data from ones research is open source. And even methods for in-house fabrication and development together with its source code is open source (open hardware and open software)
- titzer 3y agoIn the PL field, conferences have started to allow authors to submit packaged artifacts (typically, source code, input data, training data, etc) that are evaluated separately, typically post-review. The artifacts are evaluated by a separate committee, usually graduate students. As usual, everything is volunteer. Even with explicit instructions, it is hard enough to even get the same code to run in a different environment and give the same results. Would "replication" of a software technique require another team to reimplement something from scratch? That seems unworkable. I can't even imagine how hard it would be to write instructions for another lab to successfully replicate an experiment at the forefront of physics or chemistry, or biology. Not just the specialized equipment, but we're talking about the frontiers of Science with people doing cutting-edge research. I get the impression that suggestions like these are written by non-scientists who do not have experience with the peer review process of any discipline. Things just don't work like that.
- Maxion 3y ago> I get the impression that suggestions like these are written by non-scientists who do not have experience with the peer review process of any discipline. Things just don't work like that. Not to mention that the cutting edge in many sciences are perhaps two-three research groups of 5-30 individuals each in varying research institutions around the world.
- mike_hearn 3y agoIs PL theory actually science? Although we call it computer science, I don't personally think CS is actually a science in the sense of studying nature to understand it. Computers are artificial constructs. CS is a lot closer to engineering than science. Indeed it's kind of nonsensical to talk about replicating an experiment in programming language theory. For the "hard" sciences, replication often isn't so difficult it seems. LK-99 being an interesting study in this, where people are apparently successfully replicating an experiment described in a rushed paper that is widely agreed to lack sufficient details. It's cutting edge science but replication still isn't a problem. Most science isn't the LHC. The real problems with replication are found in the softer fields. There it's not just an issue of randomness or difficulty of doing the experiments. If that's all there was to it, no problem. In these fields it's common to find papers or entire fields where none of the work is replicable even in principle. As in, the people doing it don't think other people being able to replicate their work is even important at all, and they may go out of their way to stop people being able to replicate their work (most frequently by gathering data in non-replicable ways and then withholding it deliberately, but sometimes it's just due to the design of the study). The most obvious inference when you see this is that maybe they don't want replication attempts because they know their claims probably aren't true. So even if peer reviewers or journals were just checking really basic things like, is this claim even replicable in principle, that would be a good start. You would still be left with a lot of papers that replicate fine but their conclusions are still wrong because their methodology is illogical, or papers that replicate because their findings are obvious. But there's so much low hanging fruit.
- 6510 3y agoSeems like a great way for "inferior" journals to gain reputation. Counting citations seems a pretty silly formula/hack. How often you say something doesn't affect how true it is.
- user6723 3y agoI remember showing someone raw video of a Safire plasma chamber keeping the ball of plasma lit for several minutes. They said they would need to see a peer reviewed paper. The presumption brought about by the enlightenment era that everyone should get a vote was a mistake.
- janalsncm 3y agoFor a while Reddit had the mantra “pics or it didn’t happen”. At least in CS/ML there needs to be a “code or it didn’t happen”. Why? Papers are ambiguous. Even if they have mathematical formulas, not all components are defined. Peer replication in these fields is an easy low hanging fruit that could set an example for other fields of science.
- simlan 3y agoThat is too simplistic. You underestimate the depth of academia. Sure the latest break through Alzheimers study or related research would benefit from a replication. Which is done out of commercial interest anyway. But your run of the mill niche topic will not have the dollars behind it to replicate everyones research.just because CS/AI research is very convenient to replicate does not mean this can be extended to all research being done. That is exactly why peer review exists to weed out the implausible and low effort/relevance work. It is not fraud proof because it was not designed to be.
- bradley13 3y agoCS and ML are my field, although I'm no longer active in research. I always made a code archive available. Want to replicate? Download and run. This should be standard now, in the age of GitHub, GitLab, et al. If a paper discusses an implementation, but doesn't provide code, it is probably BS.
- hospadar 3y agoI assume that the goal here is to reduce the number of not-actually-valid results that get published. Not-actually-valid results happen for lots of reasons (whoops did experiment wrong, mystery impurity, cherry picked data, not enough subjects, straight-up lie, full verification expensive and time consuming but this looks promising) but often there's a common set of incentives: you must publish to get tenure/keep your job, you often need to publish in journals with high impact factor [1]. High impact journals [6] tend to prefer exciting, novel, and positive results (we tried new thing and it worked so well!) vs negative results (we mixed up a bunch of crystals and absolutely none of them are room-temp superconductors! we're sure of it!). The result is that cherry picking data pays, leaning into confirmation bias pays, publishing replication studies and rigorous but negative results is not a good use of your academic inertia. I think that creating a new category of rigor (i.e. journals that only publish independently replicated results) is not a bad idea, but: who's gonna pay for that? If the incentive is you get your name on the paper, doesn't that incentivize coming up with a positive result? How do you incentivize negative replications? What if there is only one gigantic machine anywhere that can find those results (LHC, icecube, etc, a very expensive spaceship)? There might be easier and cheaper pathways to reducing bad papers - incentivizing the publishing of negative results and replication studies separately, paying reviewers for their time, coming up with new metrics for researchers that prioritize different kinds of activity (currently "how much you're cited" and "number of papers*journal impact" things are common, maybe a "how many results got replicated" score would be cool to roll into "do you get tenure"? See [3] for more details). PLoS publish. I really like OP's other article about a hypothetical "Journal of One Try" (JOOT) [2] to enable publishing of not-very-rigorous-but-maybe-useful-to-somebody results. If you go back and read OLD OLD editions of Philosophical Transactions (which goes back to the 1600's!! great time, highly recommend [4], in many ways the archetype for all academic journals), there are a ton of wacky submissions that are just little observations, small experiments, and I think something like that (JOOT let's say) tuned up for the modern era would, if nothing else, make science more fun. Here's a great one about reports of "Shining Beef" (literally beef that is glowing I guess?) enjoy [5] [1] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6668985/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6668985/ [2] https://web.archive.org/web/20220924222624/https://blog.everydayscientist.com/?p=2455 https://web.archive.org/web/20220924222624/https://blog.ever... [3] https://www.altmetric.com/ https://www.altmetric.com/ [4] https://www.jstor.org/journal/philtran1665167 https://www.jstor.org/journal/philtran1665167 [5] https://www.jstor.org/stable/101710 https://www.jstor.org/stable/101710 [6] https://en.wikipedia.org/wiki/Impact_factor https://en.wikipedia.org/wiki/Impact_factor, see also https://clarivate.com/ https://clarivate.com/
- paulpauper 3y agothis would not apply to math or something subjective such as literature. only experimental results need to be replicated.
- geysersam 3y agoBoth review and replication has their place. The mistake is treating researchers and the scientific community as a machine: "pull here, fill these forms, comment this research, have a gold star" Let people review what they want, where they want, how they want. Let people replicate when they find interesting and motivating to work on.
- User23 3y agoOne thing that everyone needs to remember about “peer review” is that it isn’t part of the scientific method, but rather that it was imposed on the scientific enterprise by government funding authorities. It’s basically JIRA for scientists.
- hinkley 3y agoIs there space in the world for a few publications that only publish replicated work? Seems like that would be a reasonable compromise. Yes you were published, but were you published in Really Real Magazine? Get back to us when you have and we’ll discuss.
- hedora 3y agoThe website dies if I try to figure out who the author (“sam”) is, but it sounds like they are used to some awful backwater of academia. They have this idea that a single editor screens papers to decide if they are uninteresting or fundamentally flawed, then they want a bunch of professors to do grunt work litigating the correctness of the experiments. In modern (post industrial revolution) branches of science, the work of determining what is worthy of publication is distributed amongst a program committee, which is comprised of reviewers. The editor / conference organizers pick the program committee. There are typically dozens of program committee members, and authors and reviewers both disclose conflicts. Also, papers are anonymized, so the people that see the author list are not involved in accept/reject decisions. This mostly eliminates the problem where work is suppressed for political reasons, etc. It is increasingly common for paper PDFs to be annotated with badges showing the level of reproducibility of the work, and papers can win awards for being highly reproducible. The people that check reproducibility simply execute directions from a separate reproducibility submission that is produced after the paper is accepted. I argue the above approach is about 100 years ahead of what the blog post is suggesting. Ideally, we would tie federal funding to double blind review and venues with program committees, and papers selected by editors would not count toward tenure at universities that receive public funding.
- jltsiren 3y agoThe computer science practice you describe is the exception, not the norm. It causes a lot of trouble when evaluating the merits of researchers, because most people in the academia are not familiar with it. In many places, conference papers don't even count as real publications, putting CS researchers at a disadvantage. From my point of view, the biggest issue is accepting/rejecting papers based on first impressions. Because there is often only one round of reviews, you can't ask the authors for clarifications, and they can't try to fix the issues you have identified. Conferences tend to follow fashionable topics, and they are often narrower in scope than what they claim to be, because it's easier to evaluate papers on topics the program committee is familiar with. The work done by the program committee was not even supposed to be proper peer review but only the first filter. Old conference papers often call themselves extended abstracts, and they don't contain all the details you would expect in the full paper. For example, a theoretical paper may omit key proofs. Once the program committee has determined that the results look interesting and plausible and the authors have presented them in a conference, the authors are supposed to write the full paper and submit it to a journal for peer review. Of course, this doesn't always happen, for a number of reasons.
- 37326 3y ago[flagged]
- TrackerFF 3y agoSeems to have been hugged to death. But - a quick counterexample - as far as replication goes: What if the experiments were run on custom made or exceedingly expensive equipment? How are the replicators supposed to access that equipment? Even in fields which are "easy" to replicate - like machine learning - we are seeing barriers of entry due to expensive computing power. Or data collection. Or both. But then you move over to physics, and suddenly you're also dealing with these one-off custom setups, doing experiments which could be close to impossible to replicate (say you want to conduct experiments on some physical event that only occurs every xxxx years or whatever)
- dongping 3y agohttps://web.archive.org/web/20230130143126/https://blog.everydayscientist.com/replace-peer-review-with-peer-replication/ https://web.archive.org/web/20230130143126/https://blog.ever...
- fastneutron 3y agoAs much as I agree with the sentiment, we have to admit it isn't always practical. There's only one LIGO, LHC or JWST, for example. Similarly, not every lab has the resources or know-how to host multi-TB datasets for the general public to pick through, even if they wanted to. I sure didn't when I was a grad student. That said, it infuriates me to no end when I read a Phys. Rev. paper that consists of a computational study of a particular physical system, and the only replicability information provided is the governing equation and a vague description of the numerical technique. No discretized example, no algorithm, and sure as hell no code repository. I'm sure other fields have this too. The only motivation I see for this behavior is the desire for a monopoly on the research topic on the part of authors, or embarrassment by poor code quality (real or perceived).
- oneshtein 3y agoGravitational waves are confirmed by watching of distant quasars: https://scitechdaily.com/gravitational-waves-detected-using-cosmic-clocks-and-unseen-spatial-distortions/ https://scitechdaily.com/gravitational-waves-detected-using-... Hydrodynamic quantum analogs can be uses to study quantum particles at macro-scale: https://en.wikipedia.org/wiki/Hydrodynamic_quantum_analogs https://en.wikipedia.org/wiki/Hydrodynamic_quantum_analogs ESA Euclid near-infrared telescope launched few weeks ago: https://www.esa.int/Science_Exploration/Space_Science/Euclid_overview https://www.esa.int/Science_Exploration/Space_Science/Euclid...
- Nevermark 3y agoReproducibility would become a much higher priority if electronic versions of papers are required (by their distributors, archives, institutions, ...) to have reproduction sections, which the authors are encouraged to update over time. UPDATABLE COVER PAGE: Title Authors Abstract Blah, blah, ... State of reproduction: Not reproduced. Successful reproductions: ...citations... Reproduction attempts: ...citations... Countering reproductions: ...citations... UPDATABLE REPRODUCTION SECTION ATTACHED AT END Reproduction resources: Data, algorithms, processes, materials, ... Reproduction challenges: Cost, time, one-off events, ... Making this stuff more visible would help reproducers validated the value of reproduction to their home and funding institutions. Having a standard section for this, with an initial state of "Not reproduced" provides more incentive for original workers to provide better reproduction info. For algorithm and math work the reproduction could be served best with downloadable executable bundle.
- andromaton 3y agoI think Tim Berners-Lee would approve.
- freeopinion 3y agoMy mind automatically swapped out the words "peer" for "code". It took my brain to interesting places. When I came back to the actual topic, I had accidentally built a great way to contrast some of the discussion offered in this thread.
- dongping 3y agoIn the sense of replicating the results, we do have CI servers and even fuzzers running for our "code replication".
- freeopinion 3y agoI don't want to derail the science discussion too much, but what if you actually had to reproduce the code by hand? Would that process produce anything of value? Would your habit of writing i+=1 instead of i++ matter? Or iteration instead of recursion? Would code replication result in fewer use after free, or off by one than code review? Or would it mostly be a waste of resources including time?
- dongping 3y agoI'm not sure if it is meaningful to divert the topic to an analogy that is never precise. But I think we always rerun the same code (equivalent to the same procedure in the papers) in the CI or your workstation to debug. Replicating the result of the program doesn't mean that one would have to rewrite the code, as in replicating the result of a computer vision paper may only include code review and running the code from the paper.
- cycomanic 3y agoWhile I agree with the general sentiment of the paper and creating incentives for more replication is definitely a good idea, I do think the approach is flawed in several ways. The main point is that the paper seriously underestimates the difficulty and time it requires to replicate experiments in many experimental fields. Who will decide which work needs to be replicated? Should capable labs somehow become bogged down with just doing replication work? Even if they don't find the results not interesting? In reality if labs find results interesting enough to replicate they will try to do so. The current LK-99 hurrah is a perfect example of that, but it happens on a much smaller scale all the time. Researchers do replicate and build on other work all the time, they just use that replication to create new results (and acknowledge the previous work) instead of publishing a "we replicated paper". Where things usually fail is in publication of "failed replication" studies, and those are tricky. It is not always clear if the original research was flawed or the people trying to reproduce made an error (again just have a look at what's happening with LK-99 at the moment). Moreover, it can be politically difficult to try to publish a "fail to reproduce" result if you are small unknown lab, if the original result came from a big known group. Most people will believe that you are the one who made the error (and unfortunately big egos might get in the way, and the small lab will have a hard time). More generally, in my opinion the lack of replication of results is just one symptom of a bigger problem in science today. We (as in society) have essentially turned the scientific environment increasingly competitive, under the guise of "value for tax payer money". Academic scientists now have to constantly compete for grant funding, publish to keep the funding going. It's incredibly competitive to even get in ... At the same time they are supposed to constantly provide big headlines for university press releases, communicate their results to the general public and investigate (and patent) the potential for commercial exploitation. No wonder we see less cooperation.
- throwawaymaths 3y agoHow about we create a Nobel prize for replication. One impressive replication or refutation from last decade (that holds up) gets the prize split up to three ways among the most important authors.
- GuB-42 3y agoPeer review is not the end. When replication is particularly complex or expensive, peer review may just a way to see if the study is worth replicating.
- Hiromy 3y agoHola te amo
- staunton 3y agoLet's get people to publish their data and code first, shall we? That's sooo much easier than demanding whole studies to be replicated... and people still don't do it!
- pajushi 3y agoWhy shouldn't we hold science more accountable? "Science needs accounting" is a search I had saved for months which really resonates with the idea of "peer replication." In accounting, you always have checks and balances, you never are counting money alone. In many cases, accountants duplicate their work to make sure that it is accurate. Auditors are the corollary to the peer review process. They're not there to redo your work, but to verify that your methods and processes are sound.
- SonOfLilit 3y agoMy first thought was "this would never work, there is so much science being published and not enough resources to replicate it all". Then I remembered that my main issue with modern academia is that everyone is incentivized to publish a huge amount of research that nobody cares about, and how I wish we would put much more work into each of much fewer research directions.
- ayakang31415 3y agoOne of the Nobel prizes in Physics was the discovery of Higgs Boson at LHC. It cost billions of dollars just to build the facility, and required hundreds of physicists working on it to just conduct the experiment. You can't replicate this. Although I fully agree that replication must come first when it is reasonably doable.
- jonnycomputer 3y agoIf they are, in fact, implying that another lab should produce a matching data-set to try to replicate results, well, I'm sorry, but that won't work, at least in a whole lot of fields. Data collection can be very expensive, and take a lot of time. It certainly is in my field. If on, the other hand, they just want the raw data, and let others go to town on it in their own way, that's fine, probably. Results that don't depend on very particular details of the processing pipeline are probably more robust anyway.
- peteradio 3y agoWhat field is it too expensive or difficult to reproduce the data?
- 542458 3y agoAs reviewers are paid nothing and get no substantial credit for their work, I’m going to say “Every field”. Why would you do a significant chunk of a paper’s work for no reward? Replication studies are typically a bad deal even when you get to pick a notable study to reproduce and you get a paper to your name out of it - replicating what will probably be an obscure paper for no credit is not something most academics (let alone commercial research labs) would entertain.
- jonnycomputer 3y agoLots? Human subjects research, for one e.g. very often they involve clinical populations that are very hard to recruit. You can spend tens of thousands in advertising, and multiples more in labor, to get a hundred participants in, over the course of an entire year of effort, and that's not even counting the money spent on a clinician doing a diagnosis. And then, when you do, you may, say, pay $1000 for the MRI per subject, plus the $100 bucks you pay directly to the participant themselves.
- deleted 3y ago[deleted]
- TrackerFF 3y agoOn the top of my head: Say you want to research biological data from the Mariana trench. Or even worse: Some passing by comet, or planet/moon/whatever in the solar system. And just to make things EVEN worse, you need to analyze the data in some destructive way. Certainly very plausible scenarios, but also some which could prohibitively expensive to do multiple times.
- ugh123 3y agoWhy not just develop a standard "replication instructions" format that papers would need to adhere to? All methods, source code, ingredients, processes, etc are documented in a standard way. This could help tease out a lot of bullshit just by reading this section.
- jxramos 3y agoYou know what I would love to see is metadata attributes surrounding a paper such as [retracted], [reproduced], [rejected], etc. We already have the preprint thing down. Some of these would be implied by being published, ie not a preprint. Maybe even a quick symbol for what method of proof was relied upon—-video evidence, randomized control trial, observational study, Sample count of n>1000 (predefined inequality brackets), etc. I think having this quick digest of information would help an individual wade through a lot of studies quickly.
- tonmoy 3y agoWe just need a second LHC with double the number of particle physicists in the world to replicate observation of the Higgs Boson, no big deal
- andsoitis 3y agoWhy would you bother replicating someone else’s work (thereby validating it), when you could use that time and resources to do something novel?
- 10g1k 3y agoPeer review is also not part of the scientific method. It's nice, but it's not strictly part of the method. It may be more accurate to suggest that repeatability is part of the scientific method. But even that is not strictly true. Consider, the single longest running scientific work was not repeatable, and was not shared with anyone outside the cadre of people doing it. Around 3000 years ago, a secretive caste of astrologers/scribes watched the heavens, and recorded their observations for several centuries. They did not publish their findings, thus making them anecdotal (yes, that's what anecdotal means, just that it wasn't published). The exact circumstances and variables were never repeatable, due to the movements of the celestial bodies, precession, etc. Similarly, the UQ pitch drop experiment, having not yet completed, has not been repeated. But it's still an entirely valid scientific experiment.
- okaleniuk 3y agoBack when I was active in academia, our publishers were reluctant to print source code or even repository links (that was largely before GitHub) but they could still share a paper source on demand. If you reference someone else's paper and want to quote some formula, it is easier and less error prone to copy rather than retype. At that point I thought about making a TeX interpreter so one could easily "run a paper" on their own data to see if the papers claims hold. As it turned out, people often write the same formula in multiple ways and to make a TeX interpreter you'd have to specify a "runnable" subset and convince anyone to use that subset instead of what they got used to. So the idea stalled. In a few years, publishing a GitHub link along the paper became the norm, and the problem disappeared. At least in applied geometry, people do replicate each other results all the time.
- kjkjadksj 3y ago“Running a paper” is honestly challenging these days because of the resource requirements of a lot of scientific code, or the size of certain datasets. One group might have access to a beefy cluster and there’s no pressure for very performant code when they can parallelize the work across a few dozen xeons or have access to tbs of memory. Another group might be running their code on a laptop. Maybe if your data is much larger than the authors data, their code doesn’t even work since it was designed for much smaller datasets. Tools like nextflow or snakemake help with respect to having a one liner to generate all data in a paper potentially, handle dependencies, list resource expectations, use your own profile to handle your environment specific job scheduling commands and parameters. However, this still doesn’t do anything for whether you have access to the resources needed.
- wcerfgba 3y agoWhat do we recommend for qualitative research, where replicability is not a quality criterion?
- deleted 3y ago[deleted]
- bobby_the_whale 3y ago[dead]
- bobby_the_whale 3y ago[dead]
- JR1427 3y agoI think this wouldn't work, because many experiments need such specific equipment and expertise, that it would be hard to find labs that already have said equipment.
- JR1427 3y agoOne thing I think people are missing, is that labs replicate other experiments all the time as part of doing their own research. It's just that the results are not always published, or not published in a like-for-like way. But the information gets around. In my former field, everyone knew which were the dodgy papers, with results no-one could replicate.
- m-watson 3y agoThat is something that I have struggled to convey. Working in a scientific field looks a lot different than just reading things from that scientific community. Sub-fields and up real small and you know what is going on, what the problems are, and who are what kind of players in your field.
- amai 3y agoIt would already be a step in the right direction, if papers would also publish a VM with all their code, data and dependencies. It is nice to have the code (https://blog.arxiv.org/2020/10/08/new-arxivlabs-feature-provides-instant-access-to-code/ https://blog.arxiv.org/2020/10/08/new-arxivlabs-feature-prov...), but without necessary dependencies, the correct OS, compiler Version, etc. replication is even with code often impossible. Having running demos is another step in the right direction (see https://blog.arxiv.org/2022/11/17/discover-state-of-the-art-machine-learning-demos-on-arxiv/ https://blog.arxiv.org/2022/11/17/discover-state-of-the-art-...). But outside of computer science replication is even more difficult. Maybe if people would use standardized laboratories and robots, one could replicate findings by rerunning the robots code on another standard robot lab ( Basically the idea here is to virtualize laboratory work). But even then for the biggest most complex experiments this will not work: Replicate CERN anyone?
- whatever1 3y agoWe can have tiers. Tier 1 peer reviewed. Tier 2 peer replicated. We can have it as a stamp on the papers. All PhD programs have requirement for a minimum number of novel publications. We could add to the requirements a minimum number of replications. But truth to be told, a PhD in science/ engineering will probably spend their first two years trying to replicate the SOTA anyway. It’s just that today you cannot publish this effort, nobody cares, except yourself and your advisor.
- bartwr 3y agoI review 10-20 papers a year. It's a ton of unpaid, volunteer work, if I want to be a high quality reviewer then it's at least a day (at least 3 thorough reads, taking notes, writing the review, reviewer discussions, post rebuttal, back-and-forth for journals). I am lucky and privileged that my employer counts this towards work time. Only 20% papers get accepted in my domain. Now if I had to spend a week on replicating a paper - and this is CS/graphics, where it's easy and "free" - I'd never volunteer to being a reviewer. You'd need professional "replicators", but who will pay for them? And who will be them - you need experts, and if you are an expert, you don't want to merely replicate others people work full time, instead of working on your own innovation.
- hooby 3y agoCurrently BOTH is being used - peer review is the first pass, reproduction the second. Peer review might (or might not) weed out a few papers before they ever get to being reproduced - and that a paper "passed" peer review often means very little. (In some journals more, in some less). You can't replace peer review with peer replication. Reviewers often do volunteer work - supporting their field and the journal by checking submissions just for any grave errors/mistakes. They often spend just 10 to 15 minutes per submission - for hundreds of submissions. It's not realistic to ask those reviewers to do a full replication attempt for hundreds of submissions. So any attempt to "replace" review with replication, would end up basically removing review altogether, without increasing the amount of replication attempts being made.
- hooby 3y agoI did some work for online journals where papers got published even IF the peer review was bad and rejections were exceptionally rare. The review score of the abstract was only used to decide on the best topics to invite for a presentation or talk - and the review score of the paper was used to hand out awards, decide "highlighted" papers, and it also influenced how high up in the search results a given paper might appear.
- husamia 3y agoI review articles all the time. I look for things that tells me about their real work. there are nuances to some experiments that can't be known without replication.