15 ms·
AI intensifies fight against ‘paper mills’ that churn out fake research
- deleted 3y ago[deleted]
- dmingod666 3y agoNot even a mention of the galactica model. I suspect the author probably isn't aware of its existence..
- dallasg3 3y agoI didn’t know what the Galactica Model was. https://www.technologyreview.com/2022/11/18/1063487/meta-large-language-model-ai-only-survived-three-days-gpt-3-science/ https://www.technologyreview.com/2022/11/18/1063487/meta-lar...
- SiempreViernes 3y agoTo be fair, it is stressful to keep track of all the faceplants the new saviours perform
- dmingod666 3y agoFair enough, but this was a major 'event' in both tech and academic circles.. Funny that Nature publishes an article about fake papers and fails to mention the one model specifically trained exclusively on white paper to write (fake?) white papers..
- api 3y agoYou get what you incentivize. We incentivize papers so we get papers. Also anything used as a metric tends to cease to be a good metric.
- caddemon 3y agoI think the main problem is not in itself incentivizing papers or using papers as a metric. The problem is that they're in many ways the only thing incentivized and the only metric given serious consideration. And the general expectations for/form of a paper are so homogenized across a given field. This makes it very feasible to game the system in a particular way that is bad for science. It's also a serious misallocation of incentives purely from a system functioning perspective. At the very least you need to incentivize peer review alongside incentivizing papers. Currently there is hardly any incentive at all to help carefully review papers, so we have no checks and balances in place to deal with the natural result of incentivizing the writing of papers. If there were greater heterogeneity in scientific roles it wouldn't be so bad for some to have heavy paper incentives -- because there'd be others with heavy incentives to do a good job highlighting the best papers and finding flaws in the bad ones.
- riku_iki 3y agoThere is a good metric in AI research: reproducible results on well known benchmarks (Big Bench for example).
- sebzim4500 3y agoThere's no good benchmark that tests for most of the issues with LLMs, for example. Maybe if we had a benchmark that told you how often a model hallucinated we would have solved the problem by now.
- riku_iki 3y agoThere are plenty of benchmarks with factual questions.
- tikkun 3y agoThe upside from this downside is that it increases the performance of papers that can be replicated easily. In fact, once AI can replicate papers, that'll be helpful for verifying a lot of research.
- deleted 3y ago[deleted]
- nradov 3y agoHow will AI be able to replicate papers?
- captainbland 3y agoThis seems like a bleak dystopian fiction concept actually. AIs being in control of human experimentation on the basis of well meaning but poorly interpreted ethical guidelines (with, of course, fully hallucinated loopholes). Maybe add in that this has become one of the few sources of income for people as more and more jobs have been lost to the machines so many are compelled to participate.
- samstave 3y ago"Take all the research done in [field] and compare the papers written by N authors, and look into their citations and compare the outcomes of what was published and find where the summation of the reads of their papers as shared citations and summarize the research where they agree, disagree and compare with international organizations and their papers coming from specifically the leads in this field from countries X,Y,Z" Create a paper based on this information and coalesce all this data into a new outcome. Do not Lie or make up data. list where you think you are lacking in data - or based on searching, which datasets or companies should be included in the list. Create a table for all the citations sourced, with links and a comment why it is or is not included in your findings" EDIT: Someone down voted this, and I am learning 'Prompt Engineering' - can someone ELI5 why this is a bad prompt? Seriously, can someone explain what sucks about the question? Is it a stupid premise or is the crafting of the prompt lame?
- x3874 3y ago[flagged]
- naillo 3y agoThe problem is the people doing this. The AI doesn't have agency, it's just an autocomplete tool.
- burkaman 3y agoThat doesn't really seem relevant to the discussion. AI obviously makes it 1000x easier to generate realistic-looking fake papers, so it "intensifies" the issue as this article explains. It is newsworthy and discussion-worthy, and as you said AI doesn't have agency so we don't need to worry about hurting its feelings by referencing it in a headline.
- juve1996 3y agoThe only reason it's a problem is that papers became the measure - more papers = more prestige. The measure will simply change because now it won't have a strong signal.
- burkaman 3y agoThat is not the only reason and it's not the most important reason. Papers are useful in the real world and people use them to learn stuff, and also reference them as evidence. If you're trying to learn about a topic and you have to wade through 90% bullshit AI papers when search for the topic, when you aren't experienced enough to immediately tell which papers are real, that's a problem. If malicious actors are able to spread disinformation that cites very real-looking research, or pitch journalists with very real-looking papers, that's a problem.
- juve1996 3y agoNo, papers are not always useful. That isn't an inherent property. They can be useful. It depends on the content of the paper. If you can't tell whether a paper is real or bullshit it doesn't matter if an AI or a human created it. Journalists will, heaven forbid, have to do real journalism instead of blindly trusting some piece of paper they found somewhere.
- huijzer 3y agoInstead of fighting against "paper mills", let’s fight against journals. There are strong arguments against the need for peer-review for example [1]. Science does not get worse when there are more bad papers, science gets better when there are more good papers. Most papers in AI aren’t even reviewed by peers or editor and guess what: we have lots of progress happening. I’m not saying that is because the lack of review, I’m saying that reviews are not necessary for progress. Furthermore, was the iPhone good because it was reviewed by a board of "independent" reviewers? No it wasn’t. Let’s just ignore Nature. Good papers are good because they have good arguments and if they are not good, then time will tell. Papers are not good simply because Nature (TM) put a stamp on it. [1]: https://www.experimental-history.com/p/science-is-a-strong-link-problem https://www.experimental-history.com/p/science-is-a-strong-l...
- ethanbond 3y agoThe cases you mentioned have clear and immediate commercial incentives that push them toward quality. Not quite so for basic research that might (optimistically) be decades away from commercialization and will probably be commercialized from someone very very far away in the value chain from the original researcher. FWIW I'm also really critical of peer review/journals, just don't think this is a great line of analogy.
- SQueeeeeL 3y agoI've found most people who are deeply critical of peer review as a process (and not just how shady modern journals are) have never engaged in serious scientific work. They have no idea how lost in the weeds researchers can get, and how useful having a somewhat uninformed 3rd party see if they can parse your prose is for having a work be understandable. Commercialization really only applies to a very small subset of scientific discoveries
- webnrrd2k 3y agoRe: "Science does not get worse when there are more bad papers"... I don't think that this is true at all. Weeding through bad papers is, at a minimum, an opportunity cost, as is a good paper built on top of a bad one. Also, there is a societal cost in that bad research can get picked up and believed by people, like the anti-vax crowd. Or, bad research can be used to push an agenda, like anti-climate change.
- ushtaritk421 3y agoFeels like AI makes the fight against paper mills easier. A group could submit AI-generated papers to different places and create a public index based on how successful they are at getting their nonsense papers accepted.
- bagels 3y agohttps://en.wikipedia.org/wiki/Sokal_affair https://en.wikipedia.org/wiki/Sokal_affair
- SpaceBuddha 3y ago"Later, after Sokal revealed the hoax in Lingua Franca, Social Text's editors wrote that they had requested editorial changes that Sokal refused to make, and had had concerns about the quality of the writing: 'We requested him (a) to excise a good deal of the philosophical speculation and (b) to excise most of his footnotes.' Still, despite calling Sokal a 'difficult, uncooperative author", and noting that such writers were 'well known to journal editors', based on Sokal's credentials Social Text published the article in the May 1996 Spring/Summer 'Science Wars' issue"
- noobermin 3y agoThat even makes it worse, I somehow didn't know this part of it.
- BeetleB 3y agoSokal did not publish in a peer reviewed journal.
- EGreg 3y agoYour comment seems self-contradictory. AI makes the fight against spam harder since it is harder to detect fakes. The group would only get more successful over time. The “AI detection” tools get worse over time and overhyped already anyway, as we learned from the guys running that sci fi submission site.
- pierat 3y agoThen, lets use the 'AI' as its own proving ground. "AI, here's the text of this paper. We need to determine if it's accurate and true. It may be a false paper, and our goal is to determine that. AI, please analyze for logical fallacies, and list and cite them here. Next, for the potentially falsifiable claims, please give me a list of potential experiments that I can do or verify to check the claims being made." Edit: Wow, at -2, and attracting all the haters. Thought this place was about inquisitiveness and curiosity about tech. Oh well.
- worrycue 3y agoI don’t understand why you would think current AI can do any of that. Current AI can’t even stop hallucinating.
- pierat 3y agoBecause I was sub'ed to openAI until yesterday. And my experiments with it were pretty promising. I took the whole corpus of Magic the Gathering rules ( https://magic.wizards.com/en/rules https://magic.wizards.com/en/rules ) as a textfile, and fed it into 4.0 . parsed it in a few seconds. I was then able to send it cards from Gatherer (MtG card database), and then ask it pointed questions about multiple card interactions. I also compared it to what DCI judges have made ruling on as well, and matched 100%. It quite impressed me. I was thinking next is to give it the rules text, and a JSON of every card. And then ask for all combos. But I'd run out of response before it could.
- jfengel 3y agoOut of curiosity, did you try the same experiment without specifically training it on the MtG rules? Could it have all of that, including card interaction decisions, from its training data (sourced from the whole Internet)?
- pierat 3y agoI only tried it after giving the URL of the rules. There's multiple rules documents, and I wanted to make sure to use the current rules. It might have given good results without explicit rules provided. Or it could have spouted garbage. I did however ask very pointed questions about timing and layers. The ones that had DCI judge writings matched 100% (could be overfit with matching these documents). And the ones not written about also appeared to be completely accurate as well, since it also cited the rules that it came to its decision. However the larger problem is that GPT4 has been degrading quite a bit recently. It also precipitated my decision to unsubscribe. And I'm not the only one to notice this https://news.ycombinator.com/item?id=36134249 https://news.ycombinator.com/item?id=36134249
- dsr_ 3y agoIt might be time for a Research Web of Trust.
- LadyCailin 3y agoIt feels like AI is exposing Kessler Syndrome for other areas - that is, some small scale amount of junk isn’t a problem necessarily, but if you scale up that problem, it fundamentally changes the thing for everyone permanently. I guess the jury is still out whether it’s net good or bad, but it feels like it’s going to force us to confront some fundamental issues we have in society, which have perhaps always been there, but are now unavoidable, and will demand quick societal change. That almost never goes over well.
- lacker 3y agoThis seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well. It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results. It's common to use things like manual Jupyter notebooks, unlabeled CSVs, and a bunch of disorganized data files, in a one-off process that a scientist manually summarizes to produce a paper. To me, in an ideal world each paper would sit in a GitHub repository, with an integration test that verifies the code actually produces the results used in the paper. That isn't really what academics prioritize, but perhaps things will move in this direction as more people realize that we have a replication crisis, and also as scientists tend to have more software engineering skills over time.
- CobrastanJorji 3y agoImagine a world where all the raw data sits in a public repository next to a script to process the data which was reviewed as part of the publication process, which would accompany the paper and could be quickly and easily replicated with new data. Anyone could produce the graphs shown in the paper simply by downloading both and running them together. What a wonderful thing that would be. And you could imagine the government funding the storing and serving of all the data. What a wonderful daydream of an idea.
- jxramos 3y agoseems attainable, doesn't sound far fetched technically. Maybe get a few universities to start the trend.
- venv 3y agoEven this would require somehow verifying the raw data. It's plausible a bad actor could "reverse engineer" their data from a pre-determined conclusion. But yes, overall more openness is good. Still, the cost losing trust in society is very high (as you need to verify everything).
- photochemsyn 3y agoThe Jan Hendrik Schon scandal of two decades ago and the fallout from it point to how to limit the spread of fraudulent research into the accepted literature: https://en.wikipedia.org/wiki/Sch%C3%B6n_scandal https://en.wikipedia.org/wiki/Sch%C3%B6n_scandal The basic issue is that all raw data must be preserved and made available for scrutiny by other researchers after publication, as must experimental materials. Why? > "The committee requested copies of the raw data, but found that Schön had kept no laboratory notebooks. His raw data files had been erased from his computer. According to Schön, the files were erased because his computer had limited hard drive space. In addition, all of his experimental samples had been discarded or damaged beyond repair." The opposition to this standard is strongest in the corporatized patent-centric research sectors, which is most of applied science in the USA and China, etc. Ambitious academics in non-commercial sectors don't really like it either as it means competitors can jump-start their research by having access to their raw data and experimental protocols. Regardless, implementing standard practices with respect to laboratory notebooks, raw data, and experimental materials in any institution receiving federal research money makes a lot of sense - along with regular audits, with failure leading to a cutoff in funding. This problem is much broader than just the paper-mill outfits the article focuses on; some highly public and contentious related issues are public access to the raw data, research records and database sequences from the Wuhan Institute of Virology from c.2016-2019, the raw clinical trial data from Pfizer/Moderna/J&J Covid vaccine trials in 2020, and so on.
- luckydata 3y agoAndrew Huberman will remember this
- nologic01 3y agoThis is No.2 in the list of existential AI risks [1] > A deluge of AI-generated misinformation and persuasive content could make society less-equipped to handle important challenges of our time. One way to think about it as significant chunks of information exchange turning into a Market for Lemons. Namely the information asymmetry between the producer of AI junk and the receiver of said junk means that the receiver cannot distinguish between a high-quality message (a "peach") and a zero (or negative) value "lemon". Then receivers are only willing to pay a fixed price for a message that averages the value of a "peach" and "lemon". Given the zero marginal cost of producing junk, this will mean that in the limit receivers will be willing to pay exactly zero. Information exchange is completely discredited. But is this really an "existential risk" or an opportunity to think deeply about human relations, trust and the meaning of exchange? Maybe the transactional, "a fool is born every second", buyer beware, caveat emptor society we have built was never fit-for-purpose in the first place? [1] https://www.safe.ai/ai-risk#Misinformation https://www.safe.ai/ai-risk#Misinformation [2] https://en.wikipedia.org/wiki/The_Market_for_Lemons https://en.wikipedia.org/wiki/The_Market_for_Lemons
- omginternets 3y agoPrediction: as a result of fake research and it’s ilk (fake credentials, fake trade-journal pubs, fake reviews, etc.), we will observe a renewed interest in referrals and meatspace networking. It will become correspondingly difficult for outsiders to enter professional circles.
- godelski 3y agoOn the other hand, maybe we'll abandon peer review and return to a pre 1950's like style? Everyone in ML just uses arxiv because the field is moving so fast. All work is realistically peer reviewed once it is out there. We treat conferences (more important than journals) as very noisy signals. Unfortunately, we still use those as metrics for completing degrees or hiring people. The way I would try to make this better is just to post to OpenReview rather than arxiv or integrate the two. That way discussions can be had on the works, in the open, and public.
- stubybubs 3y agoIt's easy to do this with anything in computer science as others can implement or recreate whatever is being discussed easily. Not so with a 6 month experiment with several groups of humans, or even a 2 week one with rats.
- godelski 3y agoThis is not always true. ML has reproducibility issues despite the status quo being open models. Even if we don't include proprietary datasets there are still issues. Compute is one issue, but we can even ignore that[0]. The Lottery Ticket[1,2] plays a big role, in that you can just get lucky. This really should have resulted in treating benchmarks as weaker indicators but see other comments about reviewing incentives. Another issue is that it is status quo to optimize hyperparameters on test data and this results in information leakage. While this won't affect results of running a checkpoint there are issues in reproducing the work. It also adds noise. Generative papers also have a large issue in not showing random uncurated samples, which introduces huge biases and we can argue this is a reproducibility issue as you can't reproduce the results show in the works. Anyone that's played with things like Stable Diffusion (or any generative model) will be familiar with this. There are more nuanced issues and things that require intimate domain knowledge to fully understand, but I wanted to push back at this comment because I see a lot of people just brush off reproducibility concerns by pointing to GitHub. While it helps, it definitely doesn't make reproducibility "easy" or concerns nonexistent. [0] Sometimes we can't though as scaling has more effects than obvious. e.g. GANs can't scale batch on a per GPU level but scaling by multiple GPUs/nodes does give an advantage in quality (not just training times). This is non-obvious and many things can play a role. [1] https://arxiv.org/abs/1803.03635 https://arxiv.org/abs/1803.03635 [2] https://arxiv.org/abs/2109.08203 https://arxiv.org/abs/2109.08203
- godelski 3y agoThere's really only one solution to this. Luckily it is the same solution to a lot of other things. Unluckily it is the same solution that no one has sought to implement for decades. Reviewing takes a lot of nuance, care, and time. You're not going to get this when all the incentives for the reviewers are currently to reject works. You can only get a group to work on ethics alone when that group is small and accountable. Incentivize the reviewers to be high quality and worthwhile work. Incentivize the chairs to review the reviewers and force high quality reviews. An expert paying close attention to a work makes the papermill's job exponentially more difficult. It also just makes the review process actually useful in the first place, as it ensures authors get actual feedback.
- noobermin 3y ago1) You must therefore pay the reviewers. Probably a good starting point. 2) It is in fact, hard to publish things already in top journals especially if you're not established. Given the fact you also want the pool of reviewers to shrink, you're also making difficulty of publishing works even harder. Yes, the issue is paper mills, got it, but you're making it even harder for sincere scientists who aren't part of a mill in doing this. 3) The flood of work that will fall to the smaller, more accountable group will require culling, which of course, the top journals will cull the horde based on the biases they already have: pedigree, which will further calcify the issues in science and research already, or even intensify it.
- godelski 3y ago1) I mostly agree. There are at least other rewards that can also be offered, such as conference discounts. There's also the token system (pay tokens to submit works, receive tokens for reviewing). But some incentive structure needs to happen, agreed. > especially if you're not established This is commonly stated but I don't think many internalize what it means. I don't think a meritocracy can exist -- this indicator supports that -- and shows that there's strong factors influencing acceptance beyond the quality of work. 2) I do think we need to be very careful about how we sell the prestige of publications/venues. As fields get more popular and our metrics rely more on them (publish or perish) then this not only encourages paper mills, but cheating in general. As two simple examples, look at how ML publishes works where papers use proprietary datasets/models (reveals authors' lab and thus violates ethics), or how it is status quo to tune hyperparameters on test data results (information leakage). You're are a huge disadvantage if you don't cheat. We have a similar problem in schools and this is why students cheat. If it is hard to catch and bad actors aren't punished (high risk because high false positive) then you actively reward cheaters. This needs to be a serious conversation and I don't think we are having this (this is several domains, even outside academia). 3) There's a coupling effect here though, that is pretty destructive and can have bigger social ramifications. As the noise in venue publication increases, the trust in the venue decreases. But I think it increases first, as people are metric chasing. But I suspect it'll be like a rubber band, and snap back hard. The larger ramification is social trust around science, where only the large venues matter and not the small unknown journals. We already have a growing distrust as the "just asking questions" anti-science strategy has been growing, and we need to be pretty careful and think beyond a local (spatially and temporally) window.
- mtkhaos 3y agoCheers to an Open Simulation Testing Platform. No reason not to have each aspect of our knowledge base to be testable in an reasonable amount of time.
- dontupvoteme 3y agoThe easiest solutions won't be appreciated -- chain of trust from older professors/academics who co-author(or otherwise sign off that it's legit), metrics based on historical outputs from university+lab, and i'm sure some will start looking at names.
- warkdarrior 3y agoThat will help entrench rich/old/"well-established" labs and academics even more.
- joshspankit 3y agoI had hoped this was from the angle of “AI does better review than the journals” and “AI is being used to weed out a lot of papers that are just plain wrong”. Maybe next time.
- ISL 3y agoIn the short term, the solution may be reputation, distributed by keychain. When author C submits a paper and they're unknown to the journal, the journal consults the keychain to see whom has vouched for C as a credible researcher. If authors A and B have vouched for them, and A and B are in similarly good standing, C is regarded as being in good stead and the editorial process can move forward as usual. If C is later found to have maliciously faked data, it isn't just C who gets dinged on the keychain, so do A and B by extension, providing an enforcement mechanism. This is, in effect, how it has worked for centuries. Editor D calls/writes to A and B to ask, "hey, there's this new author C with a provocative paper that's in your subject area, but I've never heard of them. Are they the real deal?" If A and B vouch for C, but C turns out to be a fraud, Editor D will take any future consultation with A and B with a large grain of salt. It is possible that AI will soon be able to generate entirely-credible looking papers. What AI will struggle to do, until it starts funding research of its own, is to generate novel research that reflects whatever is actually true of nature. A reputation system is the last bulwark against entropy, if no automated tools can sort wheat from chaff. Lest you think versions of these systems aren't already in place, try posting something to the arXiv as a new author..... https://info.arxiv.org/help/endorsement.html https://info.arxiv.org/help/endorsement.html
- hospitalhusband 3y agoThis is web of trust.
- MengerSponge 3y agoSpiderman pointing at spiderman.png
- jxramos 3y agoThere was an interesting claim I heard from Eric Weinstein regarding peer review that he characterized as a cancerous infiltration into all of science from some medical area? Aha, found it in my notes... """ 1:32:28 Eric: but let me jump in--peer review is a cancer from outer space. It came from the biomedical community, it invaded science. The old system, because I have to say this because many people who are now professional scientists have an idea that peer review has always been in our literature and it absolutely [mff __ ] has not. Bret: right Eric: Okay. It used to be that the editor of a journal took responsibility for the quality of the journal which is why we had things like nature crop up in the first place because they had courageous, knowledgeable, forward-thinking editors. And so I just want to be very clear because there is a mind virus out there that says peer review is the sine qua non of scientific excellence, yada, yada, yada, bs, bs, bs. And if you don't believe me go back and learn that this is a recent invasive problem in the sciences. Bret: recent invasive problem that has no justification for existing in light of the fact Eric: Well not only does it have no justification for existing... When Watson and Crick did the double helix and this is the cleanest example we have. The paper was agreed should not be sent out for review because anyone who was competent would understand immediately what its implications were. There are reasons that great work cannot be peer-reviewed. Furthermore you have entire fields that are existing now with electronic archives that are not peer reviewed. Peer review is not peer review, it sounds like peer review it is peer injunction. It is the ability for your peers to keep the world from learning about your work. Bret: keep the world from learning about your work. Eric: because peer review is what happens, real peer review is what happens after you've passed the bs thing called "peer review". """ https://youtu.be/JLb5hZLw44s?t=5608 https://youtu.be/JLb5hZLw44s?t=5608 This is a fascinating bit of history to know about if true. I wonder if anyone has ever documented the evolution of its adoption across fields. This would be interesting too to see the HN community react to this characterization. It was the first time I ever heard of it from his wild interview with his brother.
- yellowcake0 3y agoWell it's certainly in keeping with the view he has of himself as an under appreciated scientific genius, however I don't think it makes for a very compelling critique of peer-review. Frankly, all his boohoo-hooing about being shut out of the in-group at Harvard probably has more to do with him being an insufferable narcissist, rather than any attempt by the establishment to prevent heterodox views in physics from reaching the wider scientific community.
- ftxbro 3y agoso the ai is making bad research papers honestly i don't know if i like that more or less than if they were somehow doing amazing research
- mohave529 3y agoFrom my perspective as an outsider, I have always been amazed by the use of papers in academic research as a means of communicating findings to the wider world. I find it problematic that these papers are often formatted in a way that makes them highly unreadable, with two columns and compressed text. In my opinion, adopting more modern methods of publishing research could greatly enhance the overall quality of research by making papers more accessible and increasing the likelihood of them being read. Imagine a scenario where there is a standardized format for academic papers, where the conclusion is explicitly derived from specific data and accompanied by confidence intervals. This standardized schema would enable easier referencing of other papers and easy incorporation of additional data through features like autocomplete. Implementing such a system could potentially reverse the trend of academic papers that use excessive and unnecessary language to appear more intellectually rigorous, even when the actual information being conveyed is limited. By embracing these changes, we could create a more transparent and efficient research environment that promotes clearer communication and enhances the impact of academic findings.
- jltsiren 3y agoWhatever you think academic research is, it's not always like that, and it's usually easy to find counterexamples. That's why the attempts to standardize the processes usually fail. There is not necessarily any data behind the paper. Even if there is data, the conclusions may not be about the data. Even if the conclusions are about the data, the paper may not use quantitative methods. Even if it uses quantitative methods, confidence intervals may not make sense. Even if confidence intervals do make sense, adding new data might not. And so on.
- valec 3y agoyou don't know how to read papers then. first read the abstract. then the conclusion/discussion. then the methodology to see if the study was conducted reasonably. papers are made to be read by scientists -- dumbing them down to be understandable by layman would waste time for little gain. that is the job of science communicators and journalists -- and even these often do a pretty poor job of it.
- Balgair 3y agoFor 'lay' people here on HN, I just want to make a quick point: Peer-review does not mean that the reviewer is re-doing experiments or re-running analyses. They are only reviewing to see if the paper merits inclusion in the journal. Often this means telling the authors to do more experiments or check other things. But, to be clear, the peer-reviewer does not re-do things and check if they are 'right'
- JoshuaJB 3y agoAlthough there is a type of peer review that includes redoing experiments and analysis: artifact evaluation. All of the top conferences in my field (real-time embedded systems) include this as an opt-in option, and papers get a special badge if they also pass artifact evaluation. I strongly believe that other fields in computer science would benefit by including and normalizing this process. Besides the reproducibility benefits, artifact evaluation forces documentation of the experiments and process; I've found this enormously useful when on-boarding new students to an existing project.
- Balgair 3y agoI'm nearly certain that your's is the only field that does anything like re-experimentation then. I'm in biotechy fields and it's a totally different beast out here man.
- __MatrixMan__ 3y agoThe problem is letting this be a step that directs our attention in the first place: > platform X published this so it must be good It's a root-of-trust scheme, and those create high value targets which fail to corruption. Better would be: > human Y cited this, it's about Z, and you've configured Y to be trusted in domain Z Webs of trust require maintenance, which isn't convenient, but if roots of trust continue to degrade in trustworthiness, then that maintenance will eventually be a price worth paying.
- gtop3 3y agoThe height of trust rot is Science publishing Woo Suk Hwang's 2005 stem cell research. This was one of the top scientific journals publishing research that appeared to be a nobel-prize track line of research that would have been a breakthrough in medical treatment. Instead it tainted a line of research. The research results were fraudulent, claiming a much higher success rate at generating a stem cell line than what was achieved. They lied about the number of stem cell lines generated, the number of oocytes used to generate the stem cell lines, and the number of donors the oocytes came from. As if bad data wasn't enough, they lied both to their donors, lied about their donors, and miscredited authors.
- __MatrixMan__ 3y agoI wasn't aware of the this kind of thing in 2005. How do you think the trustworthiness of papers has been trending since then, generally speaking. It sounds like that was a bit of an outlier.
- gtop3 3y agoI don't think a single worse example has occurred since, even though there have been other high profile instances of falsifying data and other unethical conduct. I think after the replication crisis the scientific community is more aware of the limitations and faults of the current system. This has impacted different fields in different ways, to varying degrees of improvements. Now, I think a lab like Hwang's would be meet with more skepticism from both readers and the top tier journals. The biggest cause for skepticism now would be Hwang's claimed success rate with the techniques. If other labs couldn't reach similar levels of success then the most generous assumption would be that the techniques weren't described well enough in the articles. Compare this to CRISPR gene editing, which is a slightly more modern advancement in genetics that is valued because of how easily other labs can incorporate it.