6 ms·
This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure
by lacker 3y ago
This seems like an extension of the replication crisis. In many fields, most published research is already bogus. The idea that peer review is enough to ensure that research is valid is perhaps not scaling well.
It would be great to have things like open data sharing. At least in astronomy, which I'm somewhat familiar with, it doesn't seem like we're that close. Most scientists cannot even reproduce their own results. It's common to use things like manual Jupyter notebooks, unlabeled CSVs, and a bunch of disorganized data files, in a one-off process that a scientist manually summarizes to produce a paper.
To me, in an ideal world each paper would sit in a GitHub repository, with an integration test that verifies the code actually produces the results used in the paper. That isn't really what academics prioritize, but perhaps things will move in this direction as more people realize that we have a replication crisis, and also as scientists tend to have more software engineering skills over time.
- CobrastanJorji 3y agoImagine a world where all the raw data sits in a public repository next to a script to process the data which was reviewed as part of the publication process, which would accompany the paper and could be quickly and easily replicated with new data. Anyone could produce the graphs shown in the paper simply by downloading both and running them together. What a wonderful thing that would be. And you could imagine the government funding the storing and serving of all the data. What a wonderful daydream of an idea.
- jxramos 3y agoseems attainable, doesn't sound far fetched technically. Maybe get a few universities to start the trend.
- venv 3y agoEven this would require somehow verifying the raw data. It's plausible a bad actor could "reverse engineer" their data from a pre-determined conclusion. But yes, overall more openness is good. Still, the cost losing trust in society is very high (as you need to verify everything).
- ChainOfFools 3y ago> It's plausible a bad actor could "reverse engineer" their data from a pre-determined conclusion. I've already heard of someone planning a product ( initially targeted at lazy^H^H^H^Hbusy high schoolers and undergrads) that will use AI to reverse discover citations that fit a predetermined narrative in a research paper. Write whatever the hell you want, and the AI will do its best to backsolve a pile of citations that support your unsourced claims and arguments. The founder, and I use that term very generously, expects the academic community to net support this because it will boost citation counts for the vast majority of low citation, low visibility works.
- reprociboi 3y agoSome journals in fact implement this idea (e.g. having the raw data underlying each figure one click away). That is however not the crux of the reproducibility crisis; it would be great if it was just "I can't make an exact copy of Supplementary Figure 5D", but rather "I can't confirm that protein X causes dementia using orthogonal techniques". There is no easy code fix for that problem.
- simonw 3y agoFunding the storing and serving of all of that data doesn't sound like a difficult problem to me. That has gotten SO cheap over the past couple of decades. There are plenty of well funded institutions that can support that kind of resource.
- jrumbut 3y agoYou'd be surprised. One headache there is what did you tell study participants you would do with the data? Did you say you'd keep it forever? Did you say 5 years? Who's in charge of making sure that this centralized repository isn't inappropriately holding and distributing data? Funding organizations also have different requirements. Then a script is only useful if paired with a set of libraries of a particular version, a specific compiler/interpreter, an OS, also there may be specialized hardware involved. Some of the languages used in science like SAS, Stata, SPSS, and Matlab, etc aren't free and open source so you can't always just bundle it. And even data storage isn't trivial. For a recent small conference abstract I processed ~150GB of data. Hundreds (thousands? Tens of thousands?) of other papers have looked at that same data. You would really want some way to deduplicate that storage, but that introduces some additional complexity. I do like this vision but I think it would be a major undertaking that would require a lot of well funded institutions coming together rather than any one in particular doing it on their own.
- Fomite 3y agoThis. "Your data is open and available in perpetuity for whatever use" is in deep conflict with how we think about human subjects data, often for very good reason.
- jrumbut 3y agoOne thing that was really jarring for me moving from the startup/consulting/advertising space to the research space was that human subjects data really gets deleted. It's not just the deleted_at column being set, there's no backup, it's really gone. Every copy, forever. I appreciate the ethics of it, and part of my reason for working in this area is because of these ethics, but even 5 years in there is so much reluctance to press that delete button.
- noobermin 3y agoI hate to be sincere, but the reality is the data is our product. If we open source our data too much, we won't have anything left to publish as others who would have to expend nowhere near the resources we must do to produce it can scrape it and publish it. (that literally is how "AI" of the current hype bubble works today, lol, why would I want that online?) The incentives are definitely bad, and that's where the actual fix should be.
- diognesofsinope 3y agoIf any level of government funds your institution you should have to releases the data. The code I write is not my code, it's the banks.
- caddemon 3y agoIf that happened it would indirectly change the incentives anyway, because everyone would start being required to release data, so bad practices that are currently incentivized would become impossible. However I think it would be better to directly incentivize data release rather than require it, at least in biomedical sciences. Because of patient privacy issues there is no way raw data release can be required across the board. And I certainly do not trust the NIH to come up with a coherent set of rules for when it is versus isn't allowed, which would mean loopholes and more corruption.
- ChainOfFools 3y agoMost projects are funded by a patchwork of different sources, not all of which would agree to the same terms of release of the information, but all of which are required sufficiently fund its creation. Not to mention maintaining ongoing storage and accessibility. single source government grants for scientific research are the exception, not the rule.
- deleted 3y ago[deleted]
- Fomite 3y agoWhat if I use your medical records, which contain pretty trivial identifying information in them?
- habitzreuter 3y ago(PhD student in STEM here) I think most people have the wrong idea about what peer review is. My advisor teaches us to treat it as a first check, but it is not a guarantee of correct results. Most of the time I spend on research is actually trying to understand the literature and reproducing their results. If I can't do it, it probably means I don't understand enough about the work I'm reading, but there is also the small chance that the published analysis is wrong, which already happened to me. EDIT: typo.
- gowld 3y agoMost people have the wrond idea about what publication is, too. For almost all papers, for almost all readers, it doesn't matter if results are correct. It's just part of the game of academia. If you need to build something based on a paper, and you need it to work correctly (so, outside of public policy or macroeconomics), you need to do the work yourself. But if you just want to cite the work and write your own paper on top of it, go ahead. Published science is like ChatGPT ;-)
- slt2021 3y agoscientific publication is like a JIRA ticket. to get visibility of your work you need to close JIRA tickets, the more the better, since people look at aggregate metrics. The more your JIRA tickets are quoted in company documentation the better. but JIRA ticket is not code, in the same way publication is not science
- caddemon 3y agoIt's never going to be a guarantee, but it's also implemented very shittily at present. There's a difference between what peer review is and what it should be. It makes sense to look at it realistically in day to day practice, but in discussion of systemic problems it should be viewed in a different light.
- BeetleB 3y agoPeer review cannot catch if they performed the procedure correctly (e.g. in a lab), but it is there to check things like: - Validity of the experimental setup - Validity of the statistical analysis methodology - Validity of the conclusions Take a look at the items in this comment:[1] > Afflicted by studies with small sample sizes, tiny effects, invalid exploratory analyses All these can be caught with peer review. > together with an obsession for pursuing fashionable trends of dubious importance In contrast, this problem is intensified with peer review. > In their quest for telling a compelling story, scientists too often sculpt data to fit their preferred theory of the world. Or they retrofit hypotheses to fit their data. Not the role of peer review to catch these, but making your data and analyses scripts open will allow these to be caught by anyone who wishes to. The problem we have with current publishing is that I could use this technique to get faulty results, publish in a prestigious journal, get a 100 citations on my paper, and one of those citations will be from the person who looked at my data and saw obvious biases in my selection (which cannot be caught by a mere peer review). That one citation pointing out how wrong I am is lost in the noise. We need a way to highlight that citation - merely making a simple graph of connections via references doesn't get us there. [1] https://news.ycombinator.com/item?id=36140540 https://news.ycombinator.com/item?id=36140540
- JacobThreeThree 3y agoThe current scientific system has long been known to have serious problems of incorrect results and conflicts of interest. This article seems like an attempt to pin the crisis on an AI scapegoat. From 2015, the Editor of The Lancet: The case against science is straightforward: much of the scientific literature, perhaps half, may simply be untrue. Afflicted by studies with small sample sizes, tiny effects, invalid exploratory analyses, and flagrant conflicts of interest, together with an obsession for pursuing fashionable trends of dubious importance, science has taken a turn towards darkness. As one participant put it, “poor methods get results”. The Academy of Medical Sciences, Medical Research Council, and Biotechnology and Biological Sciences Research Council have now put their reputational weight behind an investigation into these questionable research practices. The apparent endemicity of bad research behaviour is alarming. In their quest for telling a compelling story, scientists too often sculpt data to fit their preferred theory of the world. Or they retrofit hypotheses to fit their data. Journal editors deserve their fair share of criticism too. We aid and abet the worst behaviours. Our acquiescence to the impact factor fuels an unhealthy competition to win a place in a select few journals. Our love of “significance” pollutes the literature with many a statistical fairy-tale. We reject important confirmations. Journals are not the only miscreants. Universities are in a perpetual struggle for money and talent, endpoints that foster reductive metrics, such as high-impact publication. National assessment procedures, such as the Research Excellence Framework, incentivise bad practices. And individual scientists, including their most senior leaders, do little to alter a research culture that occasionally veers close to misconduct. https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(15)60696-1/fulltext https://www.thelancet.com/journals/lancet/article/PIIS0140-6...
- tga_d 3y agoHow do you see this as scapegoating? The headline specifically says "intensifies", the article very clearly positions AI-fabricated data as an extension of existing problems, and I don't see anything in the article downplaying those existing problems (the entire closing section is about how the summit was on issues broader than AI).
- mike_hearn 3y ago
- jvanderbot 3y agoThe unfortunate truth is that academic lineage matters. In a world with decreasing SNR from paper content, the SNR from finding the paper on a reliable researcher's homepage massively increases (because sadly, people are added to papers without their permission in an effort to boost their own signal).
- noobermin 3y agoNo offense, but astronomy is ripe for being rife with this sort of thing because the validity of astronomical research doesn't matter for the "real world" and can't really easily be tested. I feel like it could be better, but since you guys do not have any external pressures beyond yourselves (as academics), you can still publish however you feel like. A lot of the more applied fields cannot survive scrutiny. Either you have to produce something that leads to a product or something a PM will scrutinize (especially if you work for a national lab or the like, or for pharma, and so on). There, a lot of published research might be bogus, but it is nowhere near the majority.
- gtop3 3y agoAstronomy is fantastic for preventing this type of fraud. The data is incredibly open. Anyone can download datasets from telescopes, or build their own. Papers built on invalid data can be refuted by simply pointing telescopes at the location the paper is lying about. There are very few possible discoveries or theories in an astronomy journal that would immediately impact the wider culture. Compare astronomy to sociology, in which questionnaire results can be fabricated and are essentially unconformable without using other unconformable questionnaires. Compare astronomy to economics, which the articles can be used bolster political positions.
- lumost 3y agoThe way academic research is setup is not currently conducive to effective research. Research must intrinsically be allowed to fail. If you incentivize success, then what happens when someone is 2-4 years in and their only skill is apparently worthless? Without Tenure or the ability for a scientist to accumulate their own savings, there is a strong incentive for most scientists to charitably interpret results/papers to maintain relevance.
- mycologos 3y agoIt might help to expand on "bogus". Bogus has a few levels, going from "not good, but possible from a well-intentioned author trying to do the right thing" (low-level bogus) to "outright deception" (high-level bogus). Small sample sizes, statistical errors, and flawed (but honest) experiment design are all, I suggest, low-level bogus. Faking data and plagiarism are high-level bogus. I think peer review is capable of, eventually, mitigating low-level bogus. The quantitative standards in fields where low-level bogus is a problem (e.g., but definitely not only, medicine) are rising. Peer review is not a scalable solution to high-level bogus. Figuring out high-level bogus seems to be almost a full-time job [1]. You cannot expect this level of effort from researchers, especially if they are reviewing for free; I would even argue that it's easier to fake data than to figure out it's fake. It also requires more expertise to assess quality research than to write a low-level bogus paper and submit it. There's a mismatch here. There are not enough expert reviewers to handle all the low-level bogus papers. The solution therefore seems to require some kind of reputational component. There needs to be a cost to engaging in high-level bogus. But this is a hard problem. Do you ban any lead author of a paper with demonstrated high-level bogus? Publicize their names? Ban any author of a paper with demonstrated high-level bogus? Throttle the submissions any one person can make to a conference/journal at a time? I don't know. But the current model will have to change. [1] https://www.ft.com/content/32440f74-7804-4637-a662-6cdc8f3fba86 https://www.ft.com/content/32440f74-7804-4637-a662-6cdc8f3fb...
- caddemon 3y agoIt's also possible to be functionally bogus while doing everything 100% by the book. If you control things super well you can inadvertently make your result so narrow that it's basically meaningless, at least with the way the scientific system functions today. If people worked together in a more concrete way to build on prior results this effect might not be so bad. One of my favorite examples is a mouse study that did not replicate between two genetically identical mouse populations raised in the same conditions run by the same lab. The difference was the supplier the two sets of mice originally came from, and the researchers were able to pin down the cause as differences in the gut microbiome between the two (and in fact one particular bacterial strain). That is an example of great research, but the vast majority of studies will never catch something like this before they publish because they will only use one mouse supplier as they keep things controlled while minimizing costs. Because designed replication studies are fairly rare and people often do not officially publish when they find things that don't replicate in biology, we are approaching interpretation/downstream use of these highly controlled studies in an extremely inefficient way. But that isn't really the fault of individual researchers. Technically they're applying the scientific method correctly to their niche problem. It's the definition of the problem coupled with how we combine results across groups that causes the inefficiencies. As problems of interest increase in complexity we can't define them on the scale of individual labs anymore, and for some reason we've addressed it by breaking them up into subproblems in this homogeneous way. Then we just assume that piecing together the little results will work great... Anyways, I agree there is also straight up fraud and blatantly bad practices on the level of individual papers, it's definitely a continuum. Sometimes such bad results slip into the mainstream of science or even have a huge impact on subsequent funding directions like with the fabricated amyloid beta paper. But I do suspect that for the most part the blatantly bad work stays on the fringes, and the largest negative impact on scientific productivity actually comes from a level of abstraction up.
- mschuster91 3y ago> Most scientists cannot even reproduce their own results. It's common to use things like manual Jupyter notebooks, unlabeled CSVs, and a bunch of disorganized data files, in a one-off process that a scientist manually summarizes to produce a paper. The only way to get that under control is if universities had a career track of data scientists and programmers, basically a shared resource pool of specialists that all researchers could use. But most US universities seem to prefer to "invest" their endowments into professional athlete teams.
- Salgat 3y agoIt's ironic and deeply shameful when whitepapers side-step the most fundamental aspect of the scientific method: reproducibility.
- dndn1 3y agoMy instinct about numbers (i.e. incl. statistics) is that they should be reproducible: on-demand, with results and workings easily subjected to interrogation. In reality in many fields - my experience is in finance, but also in science - calculations are scattered across different languages and systems. This adds friction to any process about reproducing, understanding, analysing numbers. This is some motivation for calculang, a language for calculations I develop. https://github.com/calculang/calculang https://github.com/calculang/calculang For the HN crew it's an under-development(!) functional language with properties to permit flexible designs that can scale. It's for numbers and if you share the model alongside your numbers people can check them according to the model and see the workings. It goes well with visual number [dev/person]tools that I will release one of soon, and in the future watch for a browser extension for the workings behind numbers you are reading. Important, it won't address the raw data part of the problem, but where numbers following from that are concerned, it might get closer to that instinct.
- matthewdgreen 3y agoI love the idea. But do be aware that the high-quality software engineering like this will require has a cost: be prepared for basic science costs to scale appropriately (or for productivity to be reduced.) More to the point, not all science can be usefully verified by unit tests. Even machine-verified mathematical proofs are entirely dependent on the definitions being correct, and that requires expert human analysis.
- Fomite 3y agoGiven the non-modular budget for an NIH R01 has not been updated since Clinton was president, there is absolutely no preparation for basic science costs to scale appropriately with inflation, let alone adding something that is a lot of work for relatively minimal per-paper return.
- RosanaAnaDana 3y agoI think if the data is valid, open access, and reproducible, it shouldn't matter if the paper was written by AI.
- tga_d 3y agoThe article says this: > Kahn says that, although there will undoubtedly be positive uses of AI to support researchers writing papers, it will still be necessary to distinguish between legitimate papers written with AI and those that have been completely fabricated.
- mike_hearn 3y agoUnfortunately this assumes that the only thing which can go wrong is lack of reproducibility. Not so. I read a lot of public health papers during COVID and a staggering quantity (IMHO nearly all of them) should not have been published; many of them would have been reproducible despite that. Other things that can and do regularly appear in reproducible, peer reviewed papers: • Nonsensical methodologies • Logical fallacies • Mis-representation of their data • Incorrectly implemented software • Source datasets that are cherry-picked One might think that things like incorrectly implemented software would fall under the umbrella of irreproducible research, but that won't work. Some fields don't really recognize a distinction between model implementation and specification. The model is the implementation and if a description was once published it's quite possibly either too vague to implement, out of date, wrong, or all three. IIRC Prof Neil Ferguson's team actually rejected an attempted replication of one of their epidemic models on the basis that only they were qualified to use it! Pseudo-science like this goes unremarked in universities, only outsiders seem to care. Sometimes you get an honest academic who knows what they're supposed to do and actually does it, but there's no observable benefit to them from doing so because the ones who don't bother don't seem to suffer any consequences. tl;dr The biggest problem with the replication crisis is the framing of it as being about replication. What the world actually faces is a misleading research crisis. You can drive the replicability rate to 100% and you'll still find whole fields consisting of logical fallacies and misinformation.
- RosanaAnaDana 3y ago
- ftxbro 3y ago> "The idea that peer review is enough to ensure that research is valid is perhaps not scaling well." As others have said, this is not what peer review has ever been for, at all. It only checks for gross omissions and violations of form and syntax that are obvious to other scientists who are in adjacent fields (not even necessarily the same one). It's a relict from the times when publications were not target metrics. Peer reviewers never replicate the work. If you are a grad student tapped for peer review and you spend the time to replicate the work, you are ruining your career and your advisor will also be mad.
- Fomite 3y agoI have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation. I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same.
- ftxbro 3y ago> I have had reviewers and editors dig throughly into work, including code, and it has never been met with anything other than appreciation. Yes obviously that kind of free and unexpectedly diligent labor would be appreciated. > I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same. Yes these are the kinds of things peer reviewers do. Notably not replicating the work.
- Fomite 3y ago"Yes obviously that kind of free and unexpectedly diligent labor would be appreciated." Then perhaps characterizing doing work - including often doing partial replication, independently deriving results, checking code to make sure it's both complete and consistent, etc. as "ruining your career and your advisor will also be mad" is unfair. "I've sent reviews in detailing what further test or figure might be needed to make a point more persuasive, or noting inconsistencies, and I encourage my students to do the same." Some of this is replication. "I cannot get from X to Y in your paper without the addition of Z" is, at its essence, a statement that a result cannot be replicated given what has been provided.