3 ms·
You'd think so. I used to believe this too. Surely you try and build on a result which creates an implicit replication study, then your experiments don't work a
by mike_hearn 1mo ago
You'd think so. I used to believe this too. Surely you try and build on a result which creates an implicit replication study, then your experiments don't work and the problem is discovered pretty quick. Hence, self correction.
It often doesn't happen. The reasons seem to vary by field. We do see it in a few very hard fields like materials science.
In the softer end of social sciences like anthropology, education or the humanities it's obvious why not: they don't make empirically testable claims about reality anyway so there's nothing to replicate. This is the Arday problem, where he was publishing what they call auto-ethnography: effectively just blog posts about the author's own feelings. He even interviewed himself in the third person.
In the harder social sciences it's because nothing actually builds on anything else (see "A problem in theory" [1]). Psychology and related fields like sociology, criminology, etc are basically just a pile of random ideas pooped out in brainstorming sessions. There are hardly any overarching theories generating testable predictions, so studies test whatever random claims the researcher thinks is interesting enough to get in the news. Everything is just rocks lying on the ground rather than towers of understanding, so when one paper is overturned it has no impact on anything else. Citation counts obscure this fact, but do some reverse-citation checks and you'll find a lot of citations are worthless or outright invalid.
In modeling based fields like economics, epidemiology or climatology, it's because they can't put the thing they study in a lab and do high speed experimentation. "Replication" comes to mean repeating the same analysis on the same data set. Often they literally just run the same program on the same input file, get the same result and call it replicated - or sometimes they call it replicated even if the outputs don't match! See [2]. But the program is just a set of assumptions, not real discoveries, so building more papers on top of that program's output doesn't detect invalidity. It will all replicate in some very narrow technical sense without that proving anything useful.
Sometimes outsiders notice what's happening but over time academia has built up a complex web of lore that protects them psychologically. For example, they routinely reject or ignore feedback from outside their field on the basis that it doesn't come from "experts", meaning themselves, or it comes from "right wing" people, meaning anyone who isn't an academic or close political ally.
[1] https://www.nature.com/articles/s41562-018-0522-1 https://www.nature.com/articles/s41562-018-0522-1
[2] https://dailysceptic.org/2020/05/06/code-review-of-fergusons-model/ https://dailysceptic.org/2020/05/06/code-review-of-fergusons...
- tomnicholas1 1mo agoWhilst I weakly agree with this characterisation, it's not at all fair or accurate to put epidemiology and climatology in the same category as your other examples. Yes they have some weaknesses in replication practices (especially for purely computational work), yes the feedback loop for self-correction can be slow, but it would be a mistake to conclude or imply that the major results of those fields are therefore unreliable. That's because those fields do still make objective empirical predictions about the world, which might take years to be testable but are still ultimately testable. For example, it's been convincingly shown that climate models from decades ago predicted our current climate pretty well. [0] These fields are more comparable to subfields of physics which only get to do big experiments a handful of times, e.g. because they can only observe so many planets/supernovae/universes, or because they have to spend decades building a new billion-dollar machine to test new hypotheses. We should also not confuse these arguments with uselessness - soft social science work can still be societally valuable without being easily empirically testable. [0] https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2019GL085378 https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/201...
- mike_hearn 1mo agoDisagree, those two fields are extremely unreliable. It takes a lot of things happening simultaneously to be scientific. Yes, those two do make testable predictions, sort of. But that isn't enough. Modern (computational) epidemiology is rife with unscientific practices. They ignore data that shows a model was invalidated so the fact they make testable predictions isn't really useful. They also engage in a lot of circular reasoning, buggy coding and logical fallacies. During COVID I wrote a whole report on this topic for some politicians [1]. But the biggest issue is "A problem in theory" again - epidemiologists conflate fitting a curve in R with developing a hypothesis, so the field is overrun with overfit models that aren't based on any refinable theory of disease, just misuses of statistics. Even if you prove a paper's predictions were wrong it changes nothing because nothing built on it anyway. The problem in climatology is that when the models don't fit the data they just change the data and claim victory, e.g. September 2013: https://www.spiegel.de/international/world/climate-scientists-face-crisis-over-global-warming-pause-a-923937.html https://www.spiegel.de/international/world/climate-scientist... June 2015: https://www.nature.com/articles/nature.2015.17700 https://www.nature.com/articles/nature.2015.17700 [1] https://plan99.net/~mike/epidemiology.pdf https://plan99.net/~mike/epidemiology.pdf