5 ms·
Leakage and the reproducibility crisis in ML-based science
- jokoon 4y agoMachine learning isn't really science, since it's only statistical methods. It doesn't provide insight into what intelligence is. It's only techniques, so it's just engineering. It's brute force hacking at best, and when it sort of works, it's impossible to figure out why it does because it's black boxes all the way down. So of course there are cool things like gpt, but it's not like it's scientific progress. It doesn't really to understand how brains work, and how to understand what general intelligence really is.
- deleted 4y ago[deleted]
- randomwalker 4y agoIt's possible you may have misunderstood the title of the post. It isn't about the science of ML, or GPT-3, or brains. Rather, it's about using ML as a tool to do actual science, like medicine or political science or chemistry or whatnot. The first sentence of the post explains this.
- nestorD 4y agoMachine learning is not about getting insight into what intelligence is (it might do so as a byproduct but very few people are using it with that goal in mind). However, ML is useful to generalist science as long as you are be aware of its shortcomings and not just trying to replace something with ML without thinking about it. To give you an example I worked on (to be published): I worked with some physicists that use an incredibly slow and expensive iterative solver to get information on particules. We introduced a machine learning algorithm that predicts the end result. It does not replace the solver (you could not trust its results, contrary to a physics based numerical algorithm) but, using its guess as a starting point for the iterative solver, you can make the overall solving process orders of magnitude faster.
- YeBanKo 4y ago> It does not replace the solver (you could not trust its results, contrary to a physics based numerical algorithm) And I guess the outcome variable in the train set for the ML model was produced by the solver?
- nestorD 4y agoBy an unmodified solver, yes (we did have a test/train split in case people are wondering after having read the above article).
- notrealyme123 4y agoStatistics are the backbone of many natural sciences. It is also valid to make scientific progress just inside of a field and not in the grand scheme of things.
- deelowe 4y agoI feel that the deterministic computing theologists are going to be in for a rude awakening over time. Computing need not be perfect to work and the thing about recent advancements in ML is that they scale Extremely well.
- a-dub 4y agothis is sort of one of the weird problems that shows up at the intersection between science in the public interest and a market driven system of production. pure science that is publicly funded in the public interest would publish all raw data along with re-runnable processing pipelines that will literally reproduce the figures of interest. but, the funding is often provided by governments with the aim of producing commercializable new technology that can make life better for society. the problem is that if you do the science in the open, then it can be literally picked off by large incumbents before smaller inventors have a chance to try and spin up commercialization of their life's work. so we have this system today where science is semi-closed in order to protect the inventors, but sometimes to the detriment of the science itself.
- a-dub 4y ago...also, if a technique appears in a paper, an expert on that technique should be a reviewer and/or a standard rubric should be applied (i think nature and science have gotten much more rigorous about this in recent years in the wake of the psychology replication crisis).
- adminprof 4y agoI think you're missing two fatal problems in this "publish all raw data and code" mindset. I don't think the desire of commercialization is high on the list of fatal problems preventing people from publishing data+software. 1) How do you handle research in domains where the data is about people, so that releasing it harms their privacy? Healthcare, web activity, finances. Sure you can try to anonymize it, anonymization is imperfect, and even fully anonymized data can be joined to other data sources to de-identify people; k-anonymity only works in a closed ecosystem. If we live in a world where search engine companies don't publish their research because of this constraint, that seems worse than the current system. 2) How does one define "re-runnable processing"? Software rots, dependencies disappear, operating systems become incompatible with software, permission models change. Does every researcher now need a docker expert to publish? Who verifies that something is re-runnable, and how are they paid for it?
- nicoco 4y ago
- antipaul 4y agoNot a bad checklist (“model info sheet”) But rather than stand-alone, it should be incorporated into publications. In my experience, only a minority of applied machine learning papers provide even a minority of the info requested by the info sheet. Meaning, you really have no proper idea how cross validation was done, what preprocessing was done etc. - in actually published papers
- randomwalker 4y agoOP here. I totally agree that ideally authors should report most of this information in the paper itself. One advantage of a standalone document (we suggest putting it in an appendix) is that it's easy for reviewers to check that all of this information has been reported. Of course, authors could answer some of the questions by pointing to the sections of the paper in which they have been answered.
- AtNightWeCode 4y agoThere is even tech that claims to solve the train-test split “under the hood”. You also get surprised with the low amount of data points some of these ML people think is necessary. Far off from what you learn in basic statistics classes. To not provide accurate ways of reproducing something claimed in a paper means that the paper is invalid.
- dekhn 4y agoRecently, I saw that people were tagging their input records (test records in git repos) specifically so that later data loaders would reject those records in appropriate conditions. I forget what the tech was called but it was interesting.
- telotortium 4y agoSome people are just adding a certain well-known string to their data so that it will not be used: https://news.ycombinator.com/item?id=30927569 https://news.ycombinator.com/item?id=30927569.
- throwoutway 4y agoWhat is "data leakage"? Do the authors define it? They reference Kaufman et al, but that makes it sound just like "errors". But what are the errors?
- Flashtoo 4y agoWhen you evaluate an ML approach, you should use one part of the data to train your model and a completely separate part to evaluate it. Otherwise, your model can just memorize parts of the data (or overfit in some other way), resulting in artificially high performance. Data leakage is when there is a problem in this separation and you somehow use information about the evaluation dataset in the model training process. The table in the article lists various examples. The simplest would be to just not have a separate evaluation set. A more subtle one is if you normalize your input data based on both the training and evaluation sets; this way the normalization will be better suited to the evaluation set than it should be if you had no knowledge of it, resulting in artificially high performance.
- YeGoblynQueenne 4y agoGreat, now could you please email the authors and explain to them how to explain "lekage" in their draft paper, so the rest of us can read it also?
- YeGoblynQueenne 4y agoI'm reading the draft paper linked from the article and "data leakage" is used multiple times without any attempt at defining it. Oh well, I guess this is only meant for insiders who understand the in-group jargon. The rest of us need not be interested at all. Edit: for context, here is the paragraph titled "Leakage" from the draft paper: Leakage. Data leakage has long been recognized as a lead- ing cause of errors in ML applications (Nisbet et al., 2009). In formative work on leakage, Kaufman et al. (2012) provide an overview of different types of errors and give several rec- ommendations for mitigating these errors. Since this paper was published, the ML community has investigated leak- age in several engineering applications and modeling com- petitions (Fraser, 2016; Ghani et al., 2020; Becker, 2018; Brownlee, 2016; Collins-Thompson). However, leakage oc- curring in ML-based science has not been comprehensively investigated. As a result, mitigations for data leakage in scientific applications of ML remain understudied https://arxiv.org/pdf/2207.07048.pdf https://arxiv.org/pdf/2207.07048.pdf Yes, alright. But what the flying fuck is "leakage"? Am I supposed to go read the "formative work on leakage" cited? What if that work also leaves it undefined and points to an earlier source? What the hell is this paper about? Hhow hard it is to explain what your main subject is, so I can read your paper knowing what you're talking about? How frustrating.
- photochemsyn 4y agoI think the most significant scientific result in which machine learning is playing a major role is computational protein folding, and that at least doesn't seem to have these reproducibility problems: https://alphafold.ebi.ac.uk/ https://alphafold.ebi.ac.uk/ It's a very well-defined problem and the datasets it uses are very well-characterized (the protein crystallography database, maybe some NMR structures as well), so perhaps that helps.
- mike_hearn 4y agoA useful sounding of the alarm about the expertise crisis, but: "we advocate for a standard where bugs and other errors in data analysis that change or challenge a paper's findings constitute irreproducibility." Please no. One of the biggest problems I face when talking to people about bad science in any context is the belief that "peer reviewed and reproducible" is a synonym for correct. The phrase the authors are looking for here is not irreproducible, but rather something like: flawed, incorrect, biased, pseudo-scientific, or even intellectually fraudulent. The danger of trying to redefine the word reproducible to mean more than "do people get the same results" is threefold: 1. Researchers will reject claims their work is not reproducible by saying "actually they didn't do the same things we did so of course they didn't get the same results" and that will be a convincing rebuttal to outsiders who don't dig into the details. 2. It would further undermine trust in academic research. Way too frequently, I read a paper that makes an interesting claim, only to discover that their paper or maybe entire field has redefined common words in ways that make the claims misleading. 3. It doubles down on the unhelpful and probably counter-productive "reproducibility crisis" framing. Why unhelpful, well, there isn't really a reproducibility crisis. What we have here is actually an intellectual fraud crisis. After reading a ton of papers from outside CS in past few years it became impossible to avoid the uncomfortable conclusion that in many fields the majority of observable errors are not really errors at all, but are actually deliberate. Or at least, they are deliberately fooling themselves which amounts to the same thing. To highlight just a few examples of mistakes where you think, how can nobody have noticed this: • (from the linked paper) "a recent study included the use of anti-hypertensive drugs as a feature for predicting hypertension. Such a feature could lead to leakage because the model would not have access to this information when predicting the health outcome for a new patient. Further, if the fact that a patient uses anti-hypertensive drugs is already known at prediction time, the prediction of hypertension becomes a trivial task" • A widely used ML model from social science that claimed to predict if a Twitter account is a bot was put online and found to have an FP rate of 50%+. When this was pointed out by third parties, the response was to claim the testers were "academic trolls". Nothing was ever retracted and the model continued to be used in new papers across the field. • A COVID modelling paper blithely computed that the average Brit lives with 7 people. This was obviously wrong both in absolute values and just being nonsensical design to begin with (that should be an input taken from census data not an output), and the peer reviewer even noticed this but approved the paper anyway. • The Ferguson Report 9 model that directly led to lockdowns in the UK and other countries, was full of computational bugs like typos in their custom PRNG, out of bounds memory reads, bugs in a custom Fisher-Yates shuffle, thread safety errors and more. Nobody in the academic world appeared to care about this. The paper authors suggest making researchers fill out more paperwork to get published. I find it impossible to believe that this will work. Mandatory signed data sharing statements didn't work: there was a study posted on HN in the past few months in which someone tested this to see if the data was genuinely made available on request and something like >90% of scientists refused (in epidemiology I think). Similar results were found in other fields. In this light the mass adoption of ever more opaque and buggy statistical/computational techniques is not merely an accidental drift, correctable with a minor bit of bureaucratic oversight. These techniques seem to be popular exactly because they grant so many angles of freedom to get away with scientific murder.
- vba616 4y agoHow about instead of calling it "data leakage", refer to the "Clever Hans effect"?