9 ms·
Data detectives spotted fake numbers in a widely cited paper
- feikname 5y agoIt always suprises me how people don't do minimum effort for hiding this stuff. And no one notices. (or don't care about the obviously suspicious aspect) Do scientists actually read what they cite?
- smitop 5y agoIf they did a good job hiding it then you wouldn't know about it. For every case like this where they made a obvious mistake, there are cases where they didn't, and nobody noticed. But you only tend to hear about the ones with obvious mistakes.
- tasogare 5y ago> Do scientists actually read what they cite? I depends on ones probity, but yes the phenomenon is widespread. There are numerous reasons to cite a paper and just read it superficially: padding the references list, pleasing a reviewer by adding a paper he recommended, citing friends, etc. In my own lab they frequently cite a certain theory which if you actually read about it has nothing to do with what they are doing.
- ackbar03 5y ago> Do scientists actually read what they cite? No. What are you kidding me. There's like typically over 50 papers referenced in a typical publication, no way I have the time to read all of them carefully
- matheusmoreira 5y ago> Do scientists actually read what they cite? Yes, but few actually scrutinize the methodology of the studies. Statistics is really hard. It's easier to assume peer reviewers would have rejected the paper if it was bad.
- samhw 5y agoBenford's law (https://en.wikipedia.org/wiki/Benford%27s_law https://en.wikipedia.org/wiki/Benford%27s_law) is a pretty well-known test for fraudulent numbers. Of course, it's not infallible and (depending on the nature of the fake) it may be possible to tailor the numbers to 'pass' it, but it's a good heuristic. I'd be curious to know whether it would have detected this - and, if so, whether it was indeed used.
- uuidgen 5y agoDid YOU even check what you cite? The AUTHORS of the original paper got a dataset from a company. They didn't assume the fraud from the start and published the paper based on it. Later, when they tried to analyze the issue more in-depth, they couldn't replicate the results. THE ORIGINAL AUTHORS PUBLISHED a paper about a failure to replicate. It was just then that someone looked at the original data and found that it was faked.
- smitty1e 5y agoDivison of Labour[1] is powerful. The need is to incentivise feedback loops to QA the data on the front end. The fear of reputational implosion is apparently insufficient. [1] https://en.m.wikipedia.org/wiki/Division_of_labour https://en.m.wikipedia.org/wiki/Division_of_labour
- bobcostas55 5y agoIt's seems highly likely that Ariely did it, not the company.
- fighterpilot 5y agoWhy do you say that?
- bobcostas55 5y agoHe created the excel file. If he wanted to clear his name he could publish the original data as it was sent from the company, but he hasn't done so. And the company obviously has no incentive to falsify the data.
- feikname 5y ago> Did YOU even check what you cite? I did completely read what was available to me without having an account. > The AUTHORS of the original paper got a dataset from a company. They didn't assume the fraud from the start and published the paper based on it. My comment is not about who the culprit is or isn't. Indeed, I don't mention anything about it. Rather, it's about how, as the title says, a WIDELY cited paper has fabricated data following rather (IMO) obvious red flag patterns and none of the people -who cited the paper- raised issues about that. Thus, I questioned whether scientists read or not the papers they cite in the parent post. The question is not a judgment, I'm just truly curious since I'm not part of the formal academia, just an undergraduate.
- vidarh 5y agoThe problem is often that reading it is insufficient. Unless you try to replicate, it may well look plausible, and often missing data or details creates barriers that makes you need to want to replicate really badly to put in the effort.
- Moorhouse3 5y agoI value and respect your opinion. https://www.mygroundbiz.us/ https://www.mygroundbiz.us/
- hrhdkdlfnrne 5y agoA lot of science today is basically parallel construction. You start with a sexy story that you know will get you a lot of press, like "promising you will be honest actually makes you behave in an honest way" and then you just make that paper happen, however you can. Under the publish or perish system, scientists don't have time to actually research the topic, and imagine if it fails to confirm - you just wasted a lot of time and didn't publish anything. Too risky, it's much easier to just fake it till you make it, especially since you know peer reviewers never ever will accuse you of fraud. Any reviewer accusing a scientist of fraud will just be excluded from the community, since it's very important to uphold the narrative that "scientists are always honest, they never cheat like politicians, which is why we must always trust scientists and never question them".
- thrwyoilarticle 5y agoThere's a lot of value being abused in the term 'science'. Science is a highly valued concept but it's the result of following the scientific method, not the output of anyone with a postgrad.
- dr_dshiv 5y agoScience is what scientists do, like politics is what politicians do? Either science can be critiqued as a social construct or it is an unimpeachable Platonic aspiration. I can see both perspectives. But, communicating that science itself is a somewhat messy social phenomena might be better as a long-term message for the public.
- guerrilla 5y agoPolitics is definitely not defined as what politicians do and science is, as the GP said, when somone follows the scientific method which is something that happens all the time, far from a Platonic aspiration.
- jbjohns 5y agoIs there such a method? https://www.discovermagazine.com/planet-earth/the-scientific-method-is-a-myth https://www.discovermagazine.com/planet-earth/the-scientific...
- snakeboy 5y agoI think the original blog post [0] or Andrew Gelman's discussion of it [1] are both better sources for technical details and some historical context. In particular, this is not the first such issue for Dan Ariely, as Gelman points out, he has a history of sketchy scientific ethics like doing media tours for studies that he knows failed to replicate. [0] http://datacolada.org/98 http://datacolada.org/98 [1] https://statmodeling.stat.columbia.edu/2021/08/19/a-scandal-in-tedhemia-noted-study-in-psychology-first-fails-to-replicate-but-is-still-promoted-by-npr-then-crumbles-with-striking-evidence-of-data-fraud/ https://statmodeling.stat.columbia.edu/2021/08/19/a-scandal-...
- ZeroGravitas 5y agoThe datacolada post makes the very reasonable request that all data should be released, and scientists should make that a standard thing to do by doing it themselves and requesting others do it. It feels like this could be applied retroactively too. In this case the 2012 authors still had the data that they released in 2020 which is how the analysis got done that showed evidence of fraud. Might be worth just asking a whole bunch of people to release data they previously hadn't and collectively putting some time and effort into that.
- huitzitziltzin 5y agoIt’s not possible to release data in all circumstances. If you work with health data (I have worked with birth certificates, EMRs, inpatient discharge abstracts, drug prescription histories and other data) you can’t post it publicly. You have to promise not to include a table in the paper with a cell size of fewer than ten individuals! For what it’s worth, the Trump administration attempted to make issuing new health and environmental regs harder by requiring public data disclosure. They did this entirely because they knew that much of the data could not be disclosed. So if you were studying, eg, the effects of some pollutant on a health outcome using private data, you wouldn’t be able to rely on that study in a regulatory context bc the data could not be published. It’s a worthy idea, but there are exceptions for good reasons.
- nomoreplease 5y ago
- blunte 5y agoSeems to me... Any author of a published paper who will not stand behind the paper should have their name removed. When there is just one name left, that person either accepts responsibility for the content, or they too disavow it and get removed. When there are no names left, the paper is retracted.
- new_guy 5y agoMore than that, they should have their degree revoked too. The spam in these journals puts Buzzfeed to shame.
- iwebdevfromhome 5y agoData detectives sounds like a great job, how do I become one ?
- oenetan 5y agohttps://archive.is/YQguM https://archive.is/YQguM
- dang 5y agoPrevious threads on this. Others? A study on dishonesty was based on fraudulent data - https://news.ycombinator.com/item?id=28271805 https://news.ycombinator.com/item?id=28271805 - Aug 2021 (42 comments) Noted study in psychology fails to replicate, crumbles with evidence of fraud - https://news.ycombinator.com/item?id=28264097 https://news.ycombinator.com/item?id=28264097 - Aug 2021 (102 comments) A Big Study About Honesty Turns Out to Be Based on Fake Data - https://news.ycombinator.com/item?id=28257860 https://news.ycombinator.com/item?id=28257860 - Aug 2021 (90 comments) Evidence of fraud in an influential field experiment about dishonesty - https://news.ycombinator.com/item?id=28210642 https://news.ycombinator.com/item?id=28210642 - Aug 2021 (51 comments)