4 ms·
Is machine learning really to blame for the reproducibility crisis? I'm not in academia, but it seemed to me that the problem was entirely present without machi
by kevcampb 8y ago
Is machine learning really to blame for the reproducibility crisis? I'm not in academia, but it seemed to me that the problem was entirely present without machine learning being involed.
For example, Amgen reporting that of landmark cancer papers they reviewed,
47 of the 53 could not be replicated [1]. I would have assumed that most of them didn't involve 'machine learning'
[1] https://www.reuters.com/article/us-science-cancer/in-cancer-science-many-discoveries-dont-hold-up-idUSBRE82R12P20120328 https://www.reuters.com/article/us-science-cancer/in-cancer-...
- hannob 8y agoThe problem was there before, but there are reasons why Machine Learning is amplifying bad practices. In the past people were manually fishing for results in available datasets. Now they have algorithms to do it for them. In medicine a popular way to use ML is to improve diagnosis. Now there's already a problem in medicine that the benefits of early diagnosis are overrated and the downsides (overtreatment etc.) usually ignored. You get more of that. And TBH computer scientists aren't exactly at the forefront when it comes to scientific quality standards. (E.g. practically noone is doing preregistration in CS, which in other fields is considered a prime tool to counter bad scientific practices.)
- barry-cotter 8y agoMedicine may be better than ML but there’s not much in the difference. > COMPare: Qualitative analysis of researchers’ responses to critical correspondence on a cohort of 58 misreported trials > Background > Discrepancies between pre-specified and reported outcomes are an important and prevalent source of bias in clinical trials. COMPare (Centre for Evidence-Based Medicine Outcome Monitoring Project) monitored all trials in five leading journals for correct outcome reporting, submitted correction letters on all misreported trials in real time, and then monitored responses from editors and trialists. From the trialists’ responses, we aimed to answer two related questions. First, what can trialists’ responses to corrections on their own misreported trials tell us about trialists’ knowledge of correct outcome reporting? Second, what can a cohort of responses to a standardised correction letter tell us about how researchers respond to systematic critical post-publication peer review? > Results > Trialists frequently expressed views that contradicted the CONSORT (Consolidated Standards of Reporting Trials) guidelines or made inaccurate statements about correct outcome reporting. Common themes were: stating that pre-specification after trial commencement is acceptable; incorrect statements about registries; incorrect statements around the handling of multiple time points; and failure to recognise the need to report changes to pre-specified outcomes in the trial report. We identified additional themes in the approaches taken by researchers when responding to critical correspondence, including the following: ad hominem criticism; arguing that trialists should be trusted, rather than follow guidelines for trial reporting; appealing to the existence of a novel category of outcomes whose results need not necessarily be reported; incorrect statements by researchers about their own paper; and statements undermining transparency infrastructure, such as trial registers. https://trialsjournal.biomedcentral.com/articles/10.1186/s13063-019-3172-3 https://trialsjournal.biomedcentral.com/articles/10.1186/s13...
- hannob 8y agoI am just in the process of digging into this paper and covering it in an article, so I'm quite familiar with it. But as bad as this is: What the COMPare project is doing here is documenting the flaws of a process to counter bad scientific practice. The reality in most fields (including pretty much all of CS and ML) is that no such process exists at all, because noone even tries to fix these issues. So you have medicine where people try to fix these issues (and are - admittedly - not very good at it) versus other fields that don't even try.
- mistrial9 8y agoits quite a stretch to say that CS does not have reproducibility ! .. probably want to define that a bit
- bjourne 8y agoThis is definitely true. A common "blueprint" for articles in applied CS is first they propose a "novel" algorithm. This algorithm may be very similar to an existing algorithm and in many cases is identical to one. Then they benchmark the algorithm and shows that it performs better on some metrics than existing solutions. These benchmarks are often quite poor, and if you vary them a little, the purported performance increases vanishes.
- Amygaz 8y agoAre you planning on touching on autoML? Do you think that could help?
- neltnerb 8y agoIf we broaden Machine Learning to apply to data fitting tools, I can see how it would apply broadly and with no ill intent to produce errors. Consider the simple task of peak fitting to determine the result for some data. You're probably using a commercial tool to identify the peak position, calculate the baseline, and come up with parameters for your model. But if there's an error, and at least when I was in grad school the tools often would get stuck in weird local minima that take experience to recognize, it could easily just never be noticed. If your baseline is way off, good luck calculating your peak areas reproducibly... Data analysis is hard, and it's easy to trust algorithms to be at least more reproducible than doing it more manually. Plus side if you provide your dataset and code others can at least redo the analysis! Really excited to see more Jupyter notebooks used for publications in the future.
- evrydayhustling 8y agoCan you give a citation about ML being used this way and producing overtreatment? I'm aware of experimental results with e.g. Watson that were first overstated and then rejected, bit that seemed like a situation where an institution experimented with a poor use of ML and successfully rejected it, not where they falsely accepted an ML result.
- cle 8y agoIt seems what ML is really doing is exposing weaknesses in our scientific processes. The appropriate response here is to fix the processes, instead of blaming the latest fad and imploring people to "try harder". If the root cause isn't fixed, the next fad after ML will cause the same thing again. What feasible systemic changes can we make so that scientists can't get away with publishing sloppy results? It's not an interesting question for many scientists, who prefer focusing on technical solutions over political ones.
- Certhas 8y agoI think it is a very interesting question for a lot of scientists, but also an extremely hard one to answer. And an even harder one to implement. Even obvious wins that almost everyone can agree on, like getting publishing out of the hands of for-profit entities that add no value is taking forever, because cultural, social and political institutions are hard to move.
- skywhopper 8y agoNo but it’s making it worse because it’s giving false confidence in results and amplifying failures in experiment design. It’s also been held out as a fix for the reproducibility crisis but just as lots of statistical analysis has been done by people who don’t really understand the math but are just cargo culting other experiments, machine learning is taking that ignorance-of-your-toolset risk to the next level. Not even the people who write the software understand what patterns are being found. All they can do is point to the results that seem good at a glance. But while we can train software to beat humans within a constrained dataset with careful checking, these techniques cannot find new insights. The software does not understand what the data it’s processing represents and thus it doesn’t recognize the abstractions and limits of the data. The patterns it finds are in the low resolution data, not in reality that data is a poor copy of. But science needs to be analyzing the real world and to do that you must comprehend the errors in your data and what they mean. We are nowhere close to making software that can do that.
- Cacti 8y agoNot even the people who write the software understand what patterns are being found. All they can do is point to the results that seem good at a glance. This isn’t really true. For example, we can pass in an image to a convolutional net and see which filters are activated; this can give us a clear indication if it’s edge detectors that are activating or textures or specific shapes (eg a dog would activate edge detectors, textures that look like fur, and shapes that resemble a dogs face). We can also train models to disentangle its representations and make specific variables stand for specific things (eg for a net trained on handwriting,values in one variable can represent the slant of writing, another one the letter, another the thickness, etc.). There is also a ton of work being done in training causal models. We also have decent ways now of visualizing high dimensional loss surfaces. the field has come a long way since 2012, and the whole “it’s magic, we don’t understand why it works or what it learns” is no longer true.
- sfossa 8y agoSome of the points are true specifically the last part but you are wrong about the fact that people who write these software doesn't understand what patterns are being found. We can clearly see in ML and Deep models why the decision was made by the hypothesis using various libraries such as eli5, Tensorboard and others. Deep Learning models are in general harder to debug but still possible. Therefore we know why hypothesis produces wrong results but sometimes it not possible to mend the model due to outliers, rare events, lack of data and/or randomness that surrounds our world. Just as you point out that statistical analysis is done by people who don’t really understand the math, these false result can be due to scientists using ML without understanding its advantages and limitations.
- rjf72 8y agoMachine learning trivializes p-hacking. Take a database of random datums. Pick e.g. 3 input datums at random and map them against one manually chosen output datum. Run the machine learning system and observe the error rate. If it decreases below some value 'p' you now have a [most likely completely spurious] correlation. Spin up an explanation for it - the more sensationalized the better. Claim that the process was done in the reverse order, claim it's science, publish -- you now have a ground breaking hypothesis that was validated by experimentation. One of the big reasons that the more modeling, variables, and filtering in a study - the more you should discount it. It's too easy to prove something when there's nothing actually there. An even bigger risk here is that you can engage in the above process and spot check against other data sets to see if it can be validated elsewhere. And you can find correlations that are predictive, yet are in no way whatsoever causal. If we took a sample with enough data on all individuals in the US you'd be able to find some correlation that people who have an E as the second letter in their name, a last name of five characters in length, and went to a high school whose third letter is 'A' have a 23% higher earned income average than those outside the group. And it predicts going forward. You'd be mapping onto something that obviously has nothing to do with these variables in and of themselves. Perhaps the real issue would be it's simply a very obscure proxy to a certain group of individuals in a certain subset of educational institutions. But the problem is that this is only obviously spurious (even if predictive) because these sort of variables clearly cannot have any sort of a causal relationship. When instead you only look at a selection of variables that, in practically any combination, could be made to seem meaningful through some explanation or another - you open the door to completely 'fake' science that provides results and even predictivity, but has absolutely nothing to do with what's being claimed. So people might try to maximize towards the correlations (which are/were predictive) only to find nothing more happens than if people started actively making sure the second letter of their children's name was an E and legally changed their last name to 5 letter ones. --- As a pop culture example of this something similar to this happened with video game reviews. Video game publishers noticed that there was a rather strong correlation with positive game reviews and high sales. So they started working to raise average game scores through any means possible, eventually including 'incentivizing' game reviewers to provide higher scores. As a result game reviews began to mean next to nothing, and the strength of the correlation rapidly faded. Because obviously the correlation was never about high review scores, but about making the sort of games that organically received high review scores. Though in this case we already see "obviousness" fading, since there was some argument to be made that the high review scores were what was driving sales in and of themselves - though that was clearly not the case.
- randcraw 8y agoThe problem with ML is the same one that brought about the reproducibility crisis (RC) — believing that simply exceeding one predefined threshold for some probabilistic metric (like correlation or p value) is 'good enough'. Of course the problem is compounded if we also fail to propose a causal mechanism and you don't try to validate it — something I see data scientists doing all too often since we seldom employ anything like Design Of Experiment practices, and the data we're working with is very rarely created by us. IMHO, the RC is a reminder to scientists that to confirm a hypothesis you need to pass more than one test, and a reminder to data scientists that we must test using more than one model.
- davidgl 8y agoThe mandatory xkcd significance link https://xkcd.com/882/ https://xkcd.com/882/