15 ms·
Entities enabling scientific fraud at scale (2025)
- pixl97 7mo agoThis is Goodhart's law at scale. Number of released papers/number of citations is a target. Correctness of those papers/citations is much more difficult so is not being used as a measure. With that said, due to the apparent sizes of the fraud networks I'm not sure this will be easy to address. Having some kind of kill flag for individuals found to have committed fraud will be needed, but with nation state backing and the size of the groups this may quickly turn into a tit for tat where fraud accusations may not end up being an accurate signal. May you live in interesting times.
- armchairhacker 7mo agoThere’s an accurate way to confirm fraud: look for inconsistencies and replicate experiments. If the fraudsters “fail to replicate” legitimate experiments, ask them for details/proof, and replicate the experiment yourself while providing more details/proof. Either they’re running a different experiment, their details have inconsistencies, or they have unreasonable omissions.
- wswope 7mo agoYeah, but this happens all the time. >>95% of the time, the fraudsters get off scot-free. Look at Dan Ariely: Caught red-handed faking data in Excel using the stupidest approach imaginable, and outed as a sex pest in the Epstein files. Duke is still giving him their full backing. It’s easy to find fraud, but what’s the point if our institutions have rotten all the way through and don’t care, even when there’s a smoking gun?
- pixl97 7mo agoOf course this is slightly messy too. Fraudsters are probably always incorrect, of course they could have stolen the data. But being incorrect doesn't mean your intentionally committing fraud.
- john_strinlai 7mo agothat approach is accurate, but not scalable. the effort to publish a fraudulent study is less (sometimes much less) than the effort to replicate a study.
- awesome_dude 7mo agoIs it that easy? Machine Learning papers, for example, used to have a terrible reputation for being inconsistent and impossible to replicate. That didn't make them (all) fraudulent, because that requires intent to deceive.
- itintheory 7mo agoWhat do you think it is about machine learning that makes it hard to replicate? I'm an outsider to academic research, but it seems like computer based science would be uniquely easy - publish the code, publish the data, and let other people run it. Unless it's a matter of scale, or access to specific hardware.
- renewiltord 7mo agoA lot of things are easy if you ignore the incentive structure. E.g. a lot of papers will no longer be published if the data must be published. You’d lose all published research from ML labs. Many people like you would say “that’s perfectly okay; we don’t need them” but others prefer to be able to see papers like Language Models Are Few-Shot Learners https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165 So the answer is that we still want to see a lot of the papers we currently see because knowing the technique helps a lot. So it’s fine to lose replicability here for us. I’d rather have that paper than replicability through dataset openness.
- armchairhacker 7mo agoBut the lab must publish at least the general category of data, and if that doesn't replicate, then the model only works on a more specific category than they claim (e.g. only their dataset).
- awesome_dude 7mo agoEven with the exact same dataset and architecture, ML results aren't perfectly replicable due to random weight initialisation, training data order, and non-deterministic GPU operations. I've trained identical networks on identical data and gotten different final weights and performance metrics. This doesn't mean the model only works on that specific dataset - it means ML training is inherently stochastic. The question isn't 'can you get identical results' but 'can you get comparable performance on similar data distributions.
- ertgbnm 7mo agoThat would be great if journals bothered publishing replication studies. But since they don't, researchers can't get adequate funding to perform them, and since they can't perform them, they don't exist. We can't look for failed replication experiments if none exist.
- mike_hearn 7mo agoThat only confirms a very small subset of fraud. There are many ways to do scientific fraud that will yield internally consistent papers that pass replication as practiced today. An example is papers which claims of the form, "We proved X by doing Y" where Y is a methodology that isn't derived from and can't prove X. This sort of paper will replicate every time because if you re-derive a correct methodology the original authors say you didn't really replicate their study and your work should be ignored, but if you use their broken methodology you'll just give an intellectually fraudulent paper the stamp of replication approval. This kind of problem is actually much more widespread than work that looks scientific but in which the data is faked.
- bwfan123 7mo ago> This is Goodhart's law at scale. Also, Brandolini's law. And Adam Smith's law of supply and demand. When the ability to produce overwhelms the ability to review or refute, it cheapens the product.
- otherme123 7mo ago> Number of released papers/number of citations is a target There was this guy, well connected in the science world, that managed to publish a poor study quite high (PNAS level). It was not fraud, just bad science. There were dozens of papers and letters refuting his claims, highlighting mistakes, and so... Guess what? Attending to metrics (citations, don't matter if they are citing you to say you were wrong and should retract the paper!), the original paper was even more stellar on the eyes of grants and the journal itself. It was rage bait before Facebook even existed.
- bonoboTP 7mo ago> Number of released papers/number of citations is a target Only in stupid university leaderships is that truly what gets you hired or promoted. It's simply not true. Junior researchers in fact are believing it stronger than the facts actually support. Yes, you have to have a solid amount of publications, but doing a ridiculous amount of low-impact salami-sliced stuff or getting your name on a ton of papers where you did no real work is not going to win you a job. People who evaluate applications also live in this world and know that these metrics are being gamed. It's a cat and mouse game but the cats are also paying attention. You can only play this against really dumb government bureaucracies that mechanically give points for publications and have hard numerical criteria etc. Good institutions don't do that. Good evaluators actually read the papers themselves. Of course you can't read the papers of every single applicant if there are many. But once the applicant gets into the a somewhat filtered down list, reading the paper(s) or having an interview about it, or having them give a talk is much more informative than the number of the papers. Still not perfect, because some people can't communicate well, but communicating is part of the job, so maybe that's super bad but somewhat bad. Evaluators will use also other evidence such as recommendation letters (informally being aware of the reputation of the recommender), previous fellowships or grants obtained, etc. None of these are foolproof in themselves. But someone who has super few publications relative to their career stage will need some other piece of evidence in favor. In machine learning and AI, peer reviews are known to be quite random. If you have a good Arxiv-only paper that makes sense and you can give a good talk on it and answer questions, that will get you further than having a rubberstamp on some paper that's "meh, so what". There are some players in this game (which includes funding agencies, journals, university administration, hiring committees, conference organizers, students, etc) that are more ossified and slow-moving than others. And it's also true that double blind peer review and the rubberstamp of a top-tier conference was mostly beneficial to small, not well connected research groups, as it puts the paper on an equal footing with the big labs. The more this system erodes, to more we fall back to reputation and branding of big labs and famous researchers. Again, because there is no infinite time and infinite wisdom available to pick from applicants and there never will be. There are only tradeoffs.
- deleted 7mo ago[deleted]
- gjsman-1000 7mo agoThe future of science, the Internet, and all things: The Library of Babel by Jorge Luis Borges. Some things should not have been democratized. Silicon Valley assumes that removing restrictions on information brings freedom, but reality shows that was naïve.
- honeycrispy 7mo agoYou shouldn't just assume that the inverse would be free from fraud. The incentives for fraud still apply even when the system is not democratized.
- gjsman-1000 7mo agoExcept with AI, a fraudulent gatekept world would still be a smaller percentage of fraud than what is coming. Infinite scale fraud. The soviets may have rigged a few studies; but the democratized world now faces almost all studies being rigged.
- honeycrispy 7mo agoI think it'd be a different form of fraud that would be much harder to discredit. Think sugar industry blaming fat for health issues. More of that.
- rdevilla 7mo agoTearing down gatekeeping (i.e. "high standards") in pursuit of maximal inclusivity is just another way of saying "regression to the mean." The gate has been removed from the signal chain, and now the noise floor is at infinity.
- qsera 7mo agoThere is a saying in my native language that goes something like "If you mix poison and milk, the milk will turn poisonous, instead of poison becoming milk (aka beneficial)". I guess, to convert it into this context, we can say that if you mix the high minded and infantile (which I think is what Internet and social media did), the high minded becomes infantile, instead of the other way around.
- RobotToaster 7mo agoIt kinda skips over how large mainstream journals, with their restrictive and often arbitrary standards, have contributed to this. Most will refuse to publish replications, negative studies, or anything they deem unimportant, even if the study was conducted correctly.
- tppiotrowski 7mo agoMaybe we need a journal completely dedicated to replication studies? It would attract a lot of attention I think.
- speefers 7mo ago[dead]
- temporallobe 7mo agoMy wife completed her PhD two years ago and she put a LOT of work into it. Many sleepless nights, and it almost destroyed our marriage. It took her about 6 years of non-stop madness and she didn’t even work during that time. She said that many of her colleagues engaged in fraudulent data generation and sometimes just complete forgery of anything and everything. It was obvious some people were barely capable of putting together coherent sentences in posts, but somehow they generated a perfect dissertation in the end. It was common knowledge that candidates often hired writers and even experts like statisticians to do most of the heavy lifting. I don’t know if this is the norm now, but I simultaneously have more respect and less respect for those doctoral degrees, knowing that some poured their heart and soul into it, while others essentially cheated their way through. OTOH, I also understand that there may be a lot of grey area. My eyes have been opened!
- titzer 7mo agoI found the article and your third-hand anecdotes troubling. The good news is that it does not match any of the years of experience in my field. Fraud is just not that rampant. At PhD-granting institutions, the level of fraud you describe here is very seriously punished. It's career-ending. The violations that you are serious enough that any institution would expel said students (or harshly punish faculty--probably firing them). She did no one any favors by not reporting them. Unfortunately I don't think a dialogue around vague anecdotes is going to be particularly enlightening. What matters is culture, but also process--mechanisms and checks--plus consequences. Consequences don't happen if everyone is hush-hush about it and no one wants to be a "rat".
- qsera 7mo ago>It's career-ending.. That is where being good at politics come into play. And if you are good at it, instead of being career-ending, fraud will put you in the highest of the positions! No one wants a "plant" who cannot navigate scrutiny!
- delichon 7mo ago> The good news is that it does not match any of the years of experience in my field. I worked for exactly one academic, and he indulged in impossible-to-detect research fraud. So in my own limited experience research fraud was 100%. It was a biology lab, and this was an extremely hard working man. 18 hours per day in the lab was the norm. But the data wasn't coming out the way he wanted, and his career was at stake, so he put his thumb on the scale in various ways to get the data he needed. E.g. he didn't like one neural recording, so he repeated it until he got what he wanted and ignored the others. You would have to be right in the middle of the experiment to notice anything, and he just waved me off when I did. This same professor was the loudest voice in the department when it came to critiquing experimental designs and championing rigor. I knew what he did was wrong, because he taught me that. And he really appeared to mean it, but when push came to shove, he fiddled, and was probably even lying to himself. So I came away feeling that academic fraud is probably rampant, because the incentives all align that way. Anyone with the extraordinary integrity to resist was generally self-curated out of the job.
- fastaguy88 7mo agoIt is useful to distinguish between "effective" scientific fraud, where some set of fraudulent papers are published that drive a discipline in an unproductive direction, and "administrative" scientific fraud, where individuals use pseudo-scientific measures (H-index, rankings, etc) to make allocation decisions (grants, tenure, etc). This article suggests that administrative scientific fraud has become more accessible, but it is very unclear whether this is having a major impact on science as it is practiced. Non-scientists often seem to think that if a paper is published, it is likely to be true. Most practicing scientists are much more skeptical. When I read a that paper sounds interesting in a high impact journal, I am constantly trying to figure out whether I should believe it. If it goes against a vast amount of science (e.g. bacteria that use arsenic rather than phosphorus in their DNA), I don't believe it (and can think of lots of ways to show that it is wrong). In lower impact journals, papers make claims that are not very surprising, so if they are fraudulent in some way, I don't care. Science has to be reproducible, but more importantly, it must be possible to build on a set of results to extend them. Some results are hard to reproduce because the methods are technically challenging. But if results cannot be extended, they have little effect. Science really is self-correcting, and correction happens faster for results that matter. Not all fraud has the same impact. Most fraud is unfortunate, and should be reduced, but has a short lived impact.
- qsera 7mo ago>methods are technically challenging. And finanacially too.. >Science really is self-correcting.. When economy allows it....
- perfmode 7mo agoThe distinction between effective and administrative fraud is useful and I think underappreciated. A lot of the conversation in these threads conflates the two, which makes it hard to reason about what actually needs fixing. I want to push back a little on "science is self-correcting" though. It's true in the limit, but correction has a latency, and that latency has real costs. In fields like nutrition, psychology, or pharmacology, a fraudulent or deeply flawed result can shape clinical guidelines, public policy, and drug development pipelines for a decade or more before the correction lands. The people harmed during that window don't get made whole by the eventual retraction. The comparison I keep coming back to is fault tolerance in distributed systems. You can build a system that's "eventually consistent" and still have it be practically broken if convergence takes too long or if bad state propagates faster than corrections do. The fraud networks described in TFA are basically an adversarial workload against a system (peer review) that was designed for a much lower rate of bad input. Saying the system self-corrects is accurate, but it's not the same as saying the system is healthy or that the current correction rate is adequate. I think the practical question isn't whether science corrects itself in theory but whether the feedback loops are fast enough relative to the rate of fraud production, and right now the answer seems pretty clearly no.
- pfdietz 7mo agoOne approach is more integration of researchers with businesses. Fraud (or simple incompetence) by researchers negatively affects businesses, as they expend effort on things that aren't real. I understand this is a constant problem in the pharmaceutical industry.
- robmccoll 7mo agoIt's quite possible to be very successful marketing and selling things that aren't real. The market consists of humans, not perfectly rational machines.
- pfdietz 7mo agoEven so, businesses largely compete based on whether their products are worth buying. That means bad research is bad for business.
- stanford_labrat 7mo agothe problem is two-fold in my opinion. firstly, there are basically no legal repercussions for scientific misconduct (e.g. falsifying data, fake images, etc.). most individuals who are caught doing this get either 1) a slap on the wrist if they are too big to fail or in the employ of those who are too big to fail or 2) disbarred, banned, and lose their jobs. i don't see why you can go to jail for lying to investors about the number of users in your app but don't go to jail for lying to the public, government, and members of the scientific community about your results. secondly, due to the over production of PhD's and limited number of professorship slots competition has become so incredibly intense that in order to even be considered for these jobs you must have Nature, Cell, and Science papers (or the field equivalent). for those desperate for the job their academic career is over either way if they caught falsifying data or if they don't get the professorship. so if your project is not going the way you want it to then... sad state of things all around. i've personally witnessed enough misconduct that i have made the decision to leave the field entirely and go do something else.
- noslenwerdna 7mo agoI unironically agree, p-hacking should be a criminal offense.
- pjdesno 7mo agoPerhaps relevant to this - if you go to this global ranking of publications: https://traditional.leidenranking.com/ranking/2025/list and select "Mathematics and Computer Science", you'll find the top-ranked university is the University of Electronic Science and Technology of China. My Chinese colleagues have heard of it, but never considered it a top-ranked school, and a quick inspection of their CS faculty pages shows a distinct lack of PhDs from top-ranked Chinese or US schools. It's possible their math faculty is amazing, but I think it's more likely that something underhanded is going on...
- Atlas667 7mo agoAlmost as if capitalism makes everything into a market, and the profits make it self sustaining. How many will see the connections between this and our capitalist mode of production? Probably few since modern lit/news is allergic to systemic analysis. The blatant flaws of capitalism can't be ignored for much longer.
- pooooka 7mo agoWhat I get from this is that the professional academic community -- as a whole -- has hit critical mass, which has produced a cottage industry of paper mills and fraudulent services to support said surplus. Socialism wouldn't be the answer to this because socialism is famous for struggling with surpluses and shortages. All socialism would do is clamp down (hard) on academic's, which case you wind up with the famous shortage where not enough PHD's are available to produce research for an industry. And that's not a problem specific to just socialism, that's the fallacy of central-planning. The US government clamped down on welfare fraud and the result were freak government social workers sniffing people's bed sheets and rooting through drawers and forcing everyone to document partners. This is the situation where there needs to be a market correction because the alternative could be far worse.
- Atlas667 7mo agoIt's the tax-payer funded business model, the NGO trap. Subsidies, grants, tax-breaks, credit, deductions, exemptions, etc. A whole class of profiteers live in this sector. Even though academia funding isn't strictly categorized as an NGO, it still fits/foots the bill. Public funding of private gains is the oldest trick in the book. Ask any capitalist, they know. And I'm not saying I'm against public funding, but this is often codified into a mafia of sorts when enough money flows through. The real problem here is the fundamental lack of democratic control over our agencies. That our political organization is intensely lagging behind our productive organization. That our whole political will involves TRUSTING strangers to not be corrupt instead of directly democratizing these processes as much as possible. But besides that, you cannot remove history from historical analysis. The reason socialism countries struggled in the beginning wasn't an inherent flaw in its organization, but the fact that they were under constant war war by capitalist countries through out their existence. Also keep in mind that most socialist countries did NOT have a whole section of the world where-from to extract riches through murder (S.America, Africa, Middle east, etc), like western capitalist countries had. This is convenient for you to ignore. Maybe because you don't know, or don't care about the super-exploitative history of these places and how they tie into western capitalism. But they are inherent to western wealth and these countries' whole history is struggle against this exploitation. Not to mention that most of the countries on earth are capitalists and are very very very poor. To add: Socialism has nothing to do with "clamping down" on X or Y industry, as you hypothetically claim would happen. Socialism is almost exclusively about removing the need to generate capital from production. It unleashes production from its historical ball and chain that is profiteering. In a single sentence: Instead of production being held back by capitalists generating wealth we can produce for our own needs. It is self sustaining production. Central planning is not fallacious. Your problem is with corruption, not democratic central planning. The US Govt is a pro-capitalist entity that pro-capitalists try to distance themselves from (ironically). So using them as an example isn't saying anything at all. Central planning is not "allow a small group of people to decide things", as happens in the US Govt. Central planning is to take into account all sources of information on production to plan said production democratically. This will always beat the highly highly inefficient speculation of capitalism. Where trillions vanish on a whim and cause of a tweet, where crisis occur every 8-10 years, and where its whole trade market is built to hide that it is mostly insider trading. Again, your problem is with corruption not democratic central planning. And the way to deal with corruption is to create more democratic bodies where avg people hold real power. I don't see you asking for that either. We call that socialism.
- assaddayinh 7mo ago[dead]
- deleted 7mo ago[deleted]
- pmarreck 7mo agowhy would anyone actually interested in scientific research come to this, since it literally undermines the whole practice of science?
- cyberjerkXX 7mo agoPublish or perish. Academia requiring PhDs to publish or be fired. It's made entire fields echo chambers and prone to political influence.
- pmarreck 7mo agoso a perverse incentive, basically. shocker. got it
- zahlman 7mo agoIt's strange to me that in places full of smart people, it seems to be well understood that this happens and there are lots of anecdotes relating to it; yet the same people will be confused that their political adversaries don't trust "the science" on one issue or another. Maybe it's the scientists they don't trust?
- Hendrikto 7mo agoThat’s the beautiful thing about science: You do not have to (and should not) trust any individual. And even if you don’t trust “the consensus” of “the scientific community”, you can empirically verify yourself.
- zahlman 7mo agoCan ordinary civilians feasibly measure, for example, global trends in mean temperature without relying on the data of others?
- fc417fc802 7mo agoNo, but the literature is open for you to read. Thus you can judge the stated reasoning for yourself. You can also assess how many independent groups are making the same (or closely related) claim. If only one person claims X then it might be fraud. If large numbers of seemingly unrelated people all claim X then you're forced to decide between X and a global conspiracy to misrepresent X. To your example. Importantly, even if you deemed one of the global mean temperature datasets to be untrustworthy there are other related (but different) datasets. There are also other pieces of evidence related to the downstream claims that don't look directly at temperature.
- mike_hearn 7mo agoThis is a common misconception in discussions about scientific fraud. You don't have to be able to do a thing correctly to detect when it's being done incorrectly. You shouldn't trust any claims by scientists about global trends in mean temperatures. We can say this with confidence, without being able to compute a better timeseries, by just looking to see if the basics of the scientific method are being followed by those who do it. If we do that check we find that they don't follow the scientific method. Specifically, they edit past observations to bring them into line with theory instead of deriving theory from data believed to be robust. https://retractionwatch.com/2021/08/16/will-the-real-hottest-month-on-record-please-stand-up/ https://retractionwatch.com/2021/08/16/will-the-real-hottest...
- gadders 7mo agoThis is what happens when people argue past each other on "Trust the science". Science is good, but it's mediated via corruptible humans.
- MarkusQ 7mo agoAlso, "science" isn't some sort of dogma that you should trust, it's a process you should follow. "Trust the science" is anathema to the process. If anything, the chant should be "Doubt the science! Give it your best shot, refute it with data, with logic, provide a better explanation!"
- bonoboTP 7mo agoRealistically, the vast majority of people will not have a real chance to "refute" or even evaluate scientific claims. Maybe given a lot of time and foundational work to learn the field, some percentage of people can usefully think about them, but the vast majority can't. A lot of people are functional illiterates. They will pick based on trust and gut feelings either way. For example, when deciding whether to give your kids certain vaccines or not, you really can't expect that new parents will read the primary literature and try to refute or confirm the conclusions based on the numbers and will trace through the citations and so on... Any of those claims will also have some online account on social media refuting it with equally scientifically sounding words. In the end it will come down to heuristics and your model of how the world works, which set of people operate with what kind of intention. Like maybe you know people working in the field who you trust and hear from them that generally this sort of stuff can be trusted. Or maybe you had some bad experiences getting screwed by "the establishment" (maybe even unrelated to medicine) and now you lump all this together and distrust them.
- MarkusQ 7mo agoWhich is why we need people "doing science" to also focus on getting rid of bad ideas rather than just coming up with more. The present incentive structure is such that we reward people for coming up with shocking new ideas even if they are obviously rubbish and don't do enough to reward the ones who put in the effort to debunking existing bad ideas. Coming up with ideas is the easy part of science, but most new ideas are wrong. Getting rid of the ones that aren't actually correct is hard, yet we shower praise on people doing the easy part and ignore the ones doing the hard part.
- barbazoo 7mo agoIt always comes back to Goodhart's Law and our apparent inability to create sustainable incentive structures.
- butILoveLife 7mo agoIndustry >> Academia Profits are the deciding factor, not honor.
- jjk166 7mo agoMore broadly, an incredible amount of our society's systems are built around actors being uncoordinated. Redesigning institutions to resist networks of coordinated action between seemingly unlinked individuals will, in my opinion, be one of the great social challenges of this era.
- ilovesamaltman 7mo ago[flagged]
- ukoki 7mo agoIf you get paid by the government to do research you should make all your raw data, code, results etc, accessible to the public. If it then turns out any of it is fabricated, you should be personally liable for paying it back
- canjobear 7mo agoI ran into an interesting incident of this recently. I got a Google Scholar alert about a paper with some experiments related to a paper I had published a while ago, by one "N. Tvlg". I read the paper with interest but I started noticing that although the arguments sounded good, they didn't really make sense, and also the descriptions of the results didn't really match the figures. Eventually I came across a cluster of citations to completely unrelated papers---my field is computational linguistics and these were citations to, like, studies of battery technologies for electric cars. I looked up "N Tvlg" on Google Scholar and they had "published" several articles very recently in totally divergent fields, and upon inspection, all of them had citations back to this materials science research buried deeply somewhere. Clearly these were LLM generated papers trying to build up citation count and h-rank for someone's career.
- reactordev 7mo agoWhere there’s a ranking, there’s someone out there trying to cheat at it. Citation count is a joke.
- matthewdgreen 7mo agoThe purpose of scientific publication used to be to deliver useful scientific results to one's peers. This meant that everyone ran their own personal filter of which peers were working on interesting things, and which collections (journals) were reproducing the most interesting ones. This system still works relatively well for most conscientious researchers. The idea that we should also use publication metrics to rank researchers was never part of this system, and it obviously leads to all sorts of spam (that most scientists just work around) but that seems to really upset non-scientists.
- fph 7mo agoAre these "entities" named and shamed somewhere? I just scanned the paper but couldn't find explicit mentions.
- ventuss_ovo 7mo agoThis is the part that feels hardest to fix: once a system starts rewarding throughput over scrutiny, fraud stops looking like individual misconduct and starts looking like a supply chain problem.
- holdomanoovr 7mo ago[dead]
- alansaber 7mo agoThere has always been a lot of bad science. I would suggest that percentage has only marginally increased.