5 ms·
This is the same behaviour I've seen time and time again in biology labs. People there are re-doing the same experiment over and over until it gives them the r
by TisButMe 11y ago
This is the same behaviour I've seen time and time again in biology labs.
People there are re-doing the same experiment over and over until it gives them the result they want, and then they publish that. It's the only field where I've heard people saying "Oh, yeah, my experiment failed, I have to do it again". What does it even mean that an experiment failed? It did exactly what it was supposed to: it gave you data. It didn't fit your expectations? Good, now you have a tool to refine your expectations. But instead, we see PhD students and post-doc working 70h hours week on experiments with seemingly random results until the randomness goes their way.
A lot of them have no clue about statistical treatment of data, making a proper model to try and test assumptions against reality. Since they deal with insanely complicated system, with hidden variables all over the place, a proper statistical analysis would be the minimum expected to be able to extract any information from the data, but no matter, once you have a good looking figure, you're done. In cellular/molecular biology, nobody cares about what a p-value is, so as long as Excel tells you it's <0.05, you're golden.
The scientific process has been forgotten in biology. Right now it's basically what alchemy was to chemistry.
I very happy to see efforts like this one. Sure, they might show that a lot of "key" papers are very wrong, but that's not the crux of it. If there is a reason for biologists to make sure that their results are real, they might try to put a little more effort into checking their work. And when they figure out how much of it is bullshit, they might even try to slow down a little on the publications and go back to the basics for a little while.
I'm sorry about this rant, but I've been driven away from a career in virology by those same issues, despite my love for the discipline, so I'm a bit bitter.
- JoshTriplett 11y agohttps://xkcd.com/882/ https://xkcd.com/882/ Hopefully the people trying to reproduce results don't follow the same strategy. Also see https://en.wikipedia.org/wiki/Data_dredging https://en.wikipedia.org/wiki/Data_dredging .
- TisButMe 11y agoYeah, I've seen this xkcd. It's on point as usual, but it doesn't show the frightening thing: it's the scientists who are going "Whoa" and believing in the results they've produced...
- XorNot 11y agoBy what? Publishing them?
- return0 11y agoOn the other hand, reproducing badly-designed, not-well thought of, and probably inadequately executed experiments won't be of much help to anything. In biology its not often clear what is known, what is not, and what are the open questions, and the scientists themselves have a lot of the blame for it (for religiously sticking to closed publishing models and a general lack of initiative to share their data).
- jboggan 11y agoSpot on with the alchemy remark, I've made similar comparisons before. Coming into bioinformatics/computational biology with a strong discrete math background I found a lot of professors excited to work with me until I started telling them how their ideas and models and experiments didn't imply what they wanted them to. Just like the startup world is awash with "it's like Uber for X" the biology world is full of "let's apply X mathematical technique to $MY_NICHE" and somehow this is supposed to always generate novel positive results worthy of publication. Then you tell them that you applied such-and-such mathematical/statistical model to their pet system and that the results contradict their last 10 years of published papers . . . and they ask you to do it again. I remember one professor studied metabolic reaction networks modeled with differential equations. The networks themselves were oversimplifications and relied on ~5N parameters (N being the number of compounds in the network). The problem was that while all the examples in publication converged on a nice steady state (yay homeostasis is a consequence of the model!) it was trivial to randomize the parameters within their bounds of experimental measurement and create chaotic systems. Did this mean the model wasn't so great? No, it just meant those couldn't be the real-life configuration of those parameters . . . sigh. And now I'm a data engineer and no one asks me to get data from an API that doesn't actually provide it and I'm much happier.
- TisButMe 11y agoI hoped that it wasn't as bad in computational biology, or ecology, or any other biology field where systems and models are actually defined. It saddens me to read that your experience was as bad as mine...
- donovanr 11y agoI'm a CompBio PhD student, and my experience is that folks in that field are much more careful with statistics than in, say, molecular biology labs, but it varies from lab to lab. My PI is exceedingly meticulous about stats -- for instance, we don't report p-values, but rather entire distributions -- but that's because our work is all in silico, so it's easy to run tons of replicate simulations. Wet lab work that's finicky should definitely be held to high statistical standards, but I don't think it's fair to presume everyone in the field guilty until proven innocent.
- mbreese 11y agoYou should know that there are a million of tiny ways for an experiment to "fail", requiring one do repeat it. Reagents could be bad, a machine could have broken mid-cycle, a positive (or negative) control could have been wrong... Basically meaning, "it didn't work". In this case, any "data" that would have gotten would be incredibly suspect and you'd need to repeat the experiment. In the vast majority of labs, this is nothing nefarious, it's the way science is done. If something fails, you try to figure out why it failed, try to fix the issue, and then repeat the experiment. Only once everything works correctly can you get valid data to then try and interpret. And if you get an interesting result, you still need to repeat the experiment 2-3 times to be sure. Repeating failed experiments isn't an issue - and has nothing to do with alchemy. It's just basic troubleshooting.
- TisButMe 11y agoI absolutely agree that sometimes, you need to redo an experiment for good reasons. In most cases I've seen, people do not know why they redo the experiment, though. They know it hasn't produced the data they expected, so they redo it. Maybe it was because a reagent was bad, or a co-worker left the incubator open overnight, or maybe it was because the model is stupid. Who knows? That's my point, actually. Biologists are playing with systems they do not understand, changing parameters somewhat randomly without any control over them, and they then try to interpret whatever comes out, but ONLY it fits what they wanted. If it doesn't, then "Oh, the PCR machine is at it again!", and they throw the results away.
- mbreese 11y agoIt seems like you had a really bad experience in lab. I'm sorry for that. But it's a mistake to paint the entire field in a negative light because of this. Not all labs are bad, and some produce really outstanding work. Sometimes the issue is the PCR machine. Sometimes it's the water (my favorite troubleshooting experience from grad school). And figuring out where the issues are (is the the protocol? the reagents? or is this real signal?) can be difficult. Playing with systems we don't understand is kinda the point.
- logicallee 11y agoIsn't this actually an attractive ethical hazard† (in a very broad incentive sense) - and as such, to counteract it couldn't we actually encode an ethical obligation to immediately publish the data from any experiment, to counteract this hazard? Just in any old place, not as a full paper. You could re-run your experiment of course if you thought there was some experimental methodological error, but as you disclose somewhere your first, your second, and your third dataset all showing the same data with more or less the same methodology, you would have to show increasing confidence in your fourth or fifth dataset (the one you would otherwise publish alone), because it has to explain all of the earlier datasets as statistical flukes only: you would no longer have the incentive to run the fourth experiment alone and publish it without reference to earlier trials. To give an example, suppose anyone is rewarded by publicity if they show that flipping U.S. coins favors heads by more than 2%. This is temptingly easy to do if you don't publish experiments that don't statistically prove that: you just keep re-running the experiment until you get the result you want at the p-value you want, so if the reward for the result is more than the number of experiments you need to do to get it times the cost of each experiment (which can be truly tiny), the experiment presents an attractive ethical hazard. But if you are ethically forced to publish (or even summarize) your earlier complete experiment and dataset really in any old place, then if there is no actual difference it becomes vanishingly unlikely that you can suddenly prove your theory and explain all of the earlier experiments as statistical outliers. You would stop exploring after your second or third experiment. (Which you would still quickly summarize.) This does however present an added burden to researchers, especially if they quickly test something before really refining the methodology to do so. So these quick disclosures could still be considered quite dirty and not very meaningful. However their disclosure would give a good indication regarding how strong a result really is. (i.e. by glancing at how much dirty data precedes the actual experiment being published.) - † I'd like to recall the words "moral hazard" because it does indicate that people are tempted toward the bad behavior. But economically I think that term is too specific (meaning risk-taking where someone else has to clean up in case of failure) -https://en.wikipedia.org/wiki/Moral_hazard https://en.wikipedia.org/wiki/Moral_hazard
- zzleeper 11y agoI see a very similar mechanism in economics. Many fields (macroeconomics, IO, etc.) write models that end up with some type of calibration which is implemented in thousands of lines of code. Those lines are written by RAs with almost no programming experience, so what happens is this: While (result!=expectation) Ask assistant to look for bugs and repeat simulation The end result is that you don't stop when there are no more bugs, you stop when you've got the coefficient signs you want, and then get published. Only if your paper has a high impact, has the code and number open, AND was coded in a very easy language, people discover the flaws: "the paper ... was, and is, surely the most influential economic analysis of recent years [...] First, they omitted some data; second, they used unusual and highly ques ionable statistical procedures; and finally, yes, they made an Excel coding error" http://www.nytimes.com/2013/04/19/opinion/krugman-the-excel-depression.html http://www.nytimes.com/2013/04/19/opinion/krugman-the-excel-...
- Houshalter 11y agoSee Motivated Stopping: http://lesswrong.com/lw/km/motivated_stopping_and_motivated_continuation/ http://lesswrong.com/lw/km/motivated_stopping_and_motivated_...
- vobios 11y ago> we see PhD students and post-doc working 70h hours week on experiments with seemingly random results until the randomness goes their way. There are known and tested protocols that can fail. Not every step can be accurately recorded. It's very common that an experiment will not work well the first time it's performed (even when supervised by someone experienced). Over time, researchers improve their skills and achieve better results following the same exact protocol. Does that mean that the science behind the experiment is bad?
- nkurz 11y agoDoes that mean that the science behind the experiment is bad? No, the science might be solid. But if attempts by peers to reproduce the results fail more often than they succeed, the paper describing the science is (by definition?) inadequate. The level of detail required in the paper varies from field to field, and experiment to experiment, but if the techniques aren't described well enough for others to follow them, then the paper needs more detail.
- vobios 11y agoEven commercial products (where they have a financial incentive to provide clear and comprehensive instructions) often take a lot of training before they work properly. How can we expect a small research lab to be better?
- nkurz 11y agoYou are right to point out that giving clear instructions for a complex task is difficult. And elsewhere in this thread, 'nycticorax' makes some great points. I fear my answer is along the lines of Rutherford's often ridiculed quote about statistics and experimental design: "If you need to use statistics, then you should design a better experiment." If the experiment that you describe in your paper is too difficult for others to reproduce, perhaps it shouldn't yet be published as a paper? Would the public interest be better served by funding researchers who do simpler but reproducible work, rather than complex work where the results essentially need to be taken on faith? Carried to the extreme, this is a terrible rule, but I feel there is a kernel of truth to it. I guess the right strategy depends on how much faith you have in the correctness of published results, evaluated solely on plausibility and the reputation of the researcher, and thus how much value there is in a conclusion based on irreproducible results. I think there is a currently a justified crisis of belief in science, and that many fields would do well to get back on to solid ground. But it's a wicked problem.