8 ms·
You would think that computer science can't have replication failures, but it can. And I'm talking about my own field: machine learning. There is so much hype
by MAXPOOL 8y ago
You would think that computer science can't have replication failures, but it can. And I'm talking about my own field: machine learning. There is so much hype that I suspect that people are pushing papers and trying to actively hide the irrelevance of the methods and algorithms they develop.
Artificial intelligence faces reproducibility crisis
http://science.sciencemag.org/content/359/6377/725 http://science.sciencemag.org/content/359/6377/725
Reproducibility in Machine Learning-Based Studies:
An Example of Text Mining
https://openreview.net/pdf?id=By4l2PbQ- https://openreview.net/pdf?id=By4l2PbQ-
Missing data hinder replication of artificial intelligence studies
http://www.sciencemag.org/news/2018/02/missing-data-hinder-replication-artificial-intelligence-studies http://www.sciencemag.org/news/2018/02/missing-data-hinder-r...
>In a survey of 400 artificial intelligence papers presented at major conferences, just 6% included code for the papers' algorithms. Some 30% included test data, whereas 54% included pseudocode, a limited summary of an algorithm.
- softwaredoug 8y agoYes. As a practitioner it’s very frustrating I think the situation is pretty analogous to other area with replication crises like psychology and nutrition: local variables dominate. Just like our own body’s there’s a lot that just depends on the domain or the specifics of the problem being solved... This has been my experience in a related field (information retrieval). There are trends and best practices, but the market and Impossible to recplicate overhyped academic research reinforce each other.
- androidgirl 8y agoWhy isn't more research and research materials open source in the AI world? I don't really understand. If one doesn't have your training dataset or your code, how could they possibly replicate your results?
- __s 8y agoSo the argument is that running buggy code makes your replication tainted. So ideally experiments replicate everything, meaning part of replication is implementing the code & collecting a dataset. The exact code & dataset aren't suppose to be required for the results: you should be able to replicate the results by substituting your own code & data. When you have a rat maze experiment, it isn't expected that they include a 3d printer blueprint of the maze & genomes of the rats involved Or, more succinctly, code & data are left as an exercise for the reader
- porphyrogene 8y agoIt is common to publish the layouts and dimensions of mazes in such studies. The mazes[1] themselves often reflect the nature of the study (e.g. a simple fork to test desicion-making or a complex maze with dead ends to test navigation and/or memory.) 1. http://www.ratbehavior.org/RatsAndMazes.htm http://www.ratbehavior.org/RatsAndMazes.htm
- djsumdog 8y agoBut if more journals archive artifices (code + data) and there was a bigger push to keep these around, when replication fails, they can at least go back to the exact code and say, "huh .. they implemented this differently. I think this is the problem with their algo" or "oh, we forgot to account for this." .. ideally after you've written your own, looking only at the methodology and without looking at the code for the original.
- secabeen 8y agoHaving the journals archive artifices seems like a good idea, but the librarians hate it. The last thing they want to do is to give the for-profit journals another thing to monetize against the academy. What makes more sense is to have the universities and the library systems archive that stuff.
- lumost 8y agowhy not use a public artifact repository? e.g. github, bitbucket, or a new academic collaboration.
- androidgirl 8y agoI see. That does make sense, like making a MIT version of GPL reference code. Having more researchers avaliable to audit code seems like it would help prevent flaws from slipping through, too, to prevent false conclusions. Thank you for explaining a bit more.
- porphyrogene 8y agoResearch is often funded by firms that want the study to serve as a “scientific” basis for the efficacy of a product or service. Experiments that cannot be replicated are meaningless. That standard should not be compromised.
- taeric 8y agoAmusingly, this is often taking to the other extreme. Just being able to run the exact same code on the exact same data doesn't really tell me much about if it replicates anywhere else. That is, I agree with you that the code and the data should, ideally, be available. I lose confidence when people just rerun the same code on the same data. The slides a while back about why someone didn't like notebooks resonates well with me. Something like "Shift enter through the lesson. Is this learning?"
- jackcosgrove 8y agoAre the 30% of papers including training data actually hosting the data set somewhere for download, or are they referencing a public data set elsewhere? Both approaches are acceptable in my book. If a study uses a private data set, or one that the researcher controls and only gives out to approved partners, that study should be discounted. I understand corporate labs cannot give away their data in many cases, but corporate research carries less authority than academic research anyways. Academic research should always make the data set publicly available.
- 8note 8y agoI think if you're going to do a replication study, you should collect a different data set about the same subject, and use the same method, then see if you get the same results.
- mamon 8y agoYes, simply compiling the provided code and running it against provided data set does not really count as replication. Writing your own code and creating your own dataset is the simplest way to rule out the situations where there is something fishy in either the original code or data. And the paper itself should contain enough details to make it possible to recreate the experiment this way.
- spc476 8y agoI think compiling the provided code and running it against the provided data set does do something---you know the code reports on the data. If you get a different result with the provided code and data, then there's something different with your environment vs. the original researcher, perhaps the rounding mode, or some assumption [1]. Once that's straightened out, then you can use new data and see if you can replicate the result with the provided code and new data. Think of the provided data as a sanity test. [1] I recently fixed a bug wherein I was inadvertently relying upon Linux-specific behavior that failed when tested under Solaris.
- MAXPOOL 8y agoIf you have a good paper with important result, providing code and data is not necessary. Providing the code to replicate is good form. It shows good faith and confidence. Exact replication (exactly replicating the study) is just the starting point to check that the code works and no obvious mistakes were made. replication / reproducibility / hyperparameter sensitivity If the research yields something really important and the method is well documented, usually it can be easily checked without having the data and the code. Things like dropout, batch normalization, residual learning, .. work over multiple different datasets and hyperparameters. You can reproduce the results without faithfully replicating the experiment. If the claimed result vanishes unless you have the exact data, code or the hyperparameters, the research can't be said to be meaningfully reproducible in the scientific sense. Hyperparameter sensitivity is ML equivalent to P-Hacking.
- lumost 8y agohow many papers important results are simply bugs? Numerical code is already bug prone due to subtle and hard to test errors, research code that's not code reviewed or necessarily even tested for correctness can easily generate important results erroneously. It's also counter-productive not to publish the underlying source code for these papers, as it adds a barrier to other researchers applying the algorithm in new situations. I'd be interested in seeing if those 6% of papers which include the code get more citations than the population of papers which do not include code.
- MAXPOOL 8y ago> important results are simply bugs Probably none. If the paper is important and collects citations, the algorithm is in use. Computer science != working code. Code is required when you produce something where the scientific importance is less clear. There is need to provide more evidence. Many papers are just "Hey I made some some tweaks and it works in this particular case." Those papers should have working code.
- lumost 8y agoMost papers leverage a results table which compares the newly proposed approach with existing approaches. This section is baselines the results of the new approach with prior work and helps determine whether a new result is actually important or yet another way to achieve the same results as previous work. e.g. the tables on page 7 of this paper https://www.semanticscholar.org/paper/Automatic-Acquisition-of-Lexical-Formality-Brooke-Wang/823a397b29bf596e2734d3dff7ab5abec2f60ac9 https://www.semanticscholar.org/paper/Automatic-Acquisition-... These tables are generated using real implementations that may or may not be correct, and should be subject to review when the paper is published.
- marcosdumay 8y agoThere is some famous citation from a physicist about people that used to replicate each other experiments before, but now they share their fortran models, so they can agree on all the bugs.
- nonbel 8y agoYea, I have always found academic machine learning papers to have very little value. The helpful people create a github and write a blog post, not publish a paper.
- jackcosgrove 8y agoSelf-publishing code repositories or notebooks is clearly the superior approach, but unfortunately researchers are given less institutional credit if they are not published in an academic journal. The best approach is to get the benefits of peer review and then publish all your research artifacts alongside the paper. I am not sure if the academic journals prohibit this kind of thing though.
- PascLeRasc 8y agoI'm pretty certain PLOS allows and even encourages this: https://www.plos.org/open-access https://www.plos.org/open-access Elsevier is by far the worst publisher regarding open access though: https://twitter.com/protohedgehog/status/1028819653982736389?lang=en https://twitter.com/protohedgehog/status/1028819653982736389...
- Voloskaya 8y agoBecause you are not the target audience for these papers. ML researchers (or any researcher for that matter) don't have time to go through a long form blogpost or skim through a GitHub repo. There are tens of interesting results that come out everyday. Papers with abstract and sections is the most efficient way to understand the overall idea. Blog post are nice for reaching outside of the research community. Code should always be linked for people that want to reproduce, but they don't replace a paper.
- nonbel 8y ago>"ML researchers (or any researcher for that matter) don't have time to go through a long form blogpost or skim through a GitHub repo." I don't see why this would take any more time than properly studying a paper, in fact it should be quicker since the information is presented in the proper format (not translated from code to math/prose).
- yurolis 8y agoI do research in psychology and also a little in machine learning (classification using text samples). I've actually been working with the Many Labs data (the project that's the focus of the article). My impressions of machine learning match up with yours. There's just so many parameters that there's so many opportunities to capitalize on chance, combined with [necessarily] huge datasets that preclude replication, that it's hard to avoid. I have also been surprised at how much tweaking there is of parameters; I'm used to work in traditional statistics where more is derived from theoretical principles, as opposed to trying a bunch of values to see what "works." This all lends itself to overfitting. I strongly believe that a lot of the adversarial input work is basically capitalizing on this overfitting. To be fair, I don't think people are necessarily being nefarious, I think people across all fields just don't appreciate the dangers of overfitting. The one upside of this Many Labs work, mentioned in the paper, is that it tends to show that the common criticism of "how can you generalize from this sample of undergrads" is not so much of a problem. Not that it's not an issue at all, but if you're studying some basic cognitive process it's probably not going to matter that much if you use undergrads versus some perfectly representative sample of the population. People have shown this in different ways before but it's useful to know more about. Obviously with some things sociogeographic variation will matter more though. One thing that might not be totally obvious from this article and others is that although some effects clearly replicate, and others do not, there are some effects that seem to be in a grey area, where the effects are probably real but much smaller than originally reported. The distribution of effect size estimate distributions is more continuous, with shades of grey, than these news reports would have you think. Whether or not it matters that an effect is tiny versus zero might not practically matter, but at some level it is important to be mindful of.
- forapurpose 8y ago> One thing that might not be totally obvious from this article and others is that although some effects clearly replicate, and others do not, there are some effects that seem to be in a grey area, where the effects are probably real but much smaller than originally reported. The two prior replication studies that I know of - the first was several years ago and raised the issue in the public mind afaik, then second maybe in the last year - both replicated almost all of the effects. The problems were that some effects were weaker than in the original studies. Are you saying that this most recent attempt at replication shows a large number of studies with no significant effects at all?
- peter303 8y agoAt a recent Boulder deep learning talk a distibuished scientist asked the speaker 'how can I use your result?' There are not easy ways to publish and distribute such results / models.
- selimthegrim 8y agoIs there a video of this talk?
- RA_Fisher 8y agoIt's for this reason that I like publishing straight to my blog: http://statwonk.com/weibull.html http://statwonk.com/weibull.html Let the code / math do the work. I intentionally make it easy to pick up, modify and inspect. The benefit to me is precisely being alerted if it isn't replicable.
- the8472 8y agoSo a replication crisis in artificial neuropsychology.
- jhanschoo 8y agoYou may be joking, but I'm tired that everytime neural networks comes up this unfortunate conflation of different aspects of AI occurs; one based in accurate biological modeling, the other based in tractable statistical methods. Whereas it is true that both are inspired by neural activation and that both can produce a complex, yet coherent mapping from input to output, this is like saying that birds and aircraft are similar because both fly and have wings.
- quotemstr 8y agoYet both birds and aircraft can stall. (Granted, birds can get out of stalls much more easily than aircraft can.) While many analogies between natural neural systems and various ML arrangements are false, not all of them are.
- noobermin 8y agoThe analogy is like "both a human and a baseball launcher can throw a ball, so they must be the same." People really need to let go of the terminology around machine learning that leads people to really believing that it's close to actual cognition.
- pas 8y agoBut what is cognition? People should reevaluate their views on how it works. The brain is very complex and a very big ball of highly specialized circuitry (and the software running on those and managing various aspects and parts of it and them). The similarities are at the same time very striking but the differences are just as drastic in number of layers and other size related parameters our various brain components have. Yes, AlphaGoZero is not going to learn to cook and sing and dance, but it beats our gaming component in a lot of areas. Of course a synapse is not really a memristor or a ReLU, but at the same time not that far off either. Similarly a biological neuron is not just a simple backpropagatin integrator like a perceptron, but it works very similarly. And just as we don't really know how all the representations work in an ANN, we don't exactly know the biological aspects of memory/learning/seeing work in our brains, yet we have made enormous progress on both. See all the demos/visualizations on how various layers encode in deep nets and look at the data about the brain's visual cortex and the V1 circuit, look at the gene spliced mouse studies where memory encoding is studied (and sometimes only one neuron encodes for a memory/face/concept - just as with deep nets).
- virmundi 8y agoI’ve raised the issue of pseudo code and missing data multiple times in my masters work. My contention is that public funding should require public access to such information for external verifying. I’ve been shutdown on every front. People say that the private work needs to remain private in order to ensure future research monitization from grants and IP sales.
- hyperion2010 8y agoI was talking to a friend who moved from the Google Brain team to the DeepMind team and he said point blank that no one was going to reproduce their work and that they were not going to reproduce other groups' work. Everyone has their own research agenda and the cost of the hardware and compute time needed to process the data is available to a tiny number of research groups around the world.
- BlackFly 8y ago> the cost of the hardware and compute time needed to process the data is available to a tiny number of research groups You'd think this would make them more interested in replication, since otherwise the likelihood that they are actually doing something of value is uncertain. Science and research is full of all kinds of false starts, you don't increase your speed by blinding yourself. Falsification is the very definition of science.
- camelCaseOfBeer 8y agoResearchers can get published for interesting applications of ML algoritms treating them by and large as a black-box. I walked around a ML conference consisting of mostly grad students and found a good 30% of the presenters likely used the same data for training and testing.
- MaxBarraclough 8y ago> 6% included code for the papers' algorithms This is completely ridiculous. In pure mathematics, everything you need is right there in the paper. If instead you're in the physical sciences, you obviously can't include your lab in the paper. For software, it's perfectly possible to include the lab in the paper, so to speak, and there's no excuse for depriving the rest of us of access. Not interested in releasing the source? Ok, just don't brand your project as 'computer science'. Including full source-code (and ideally data-sets, unless there's good reason this can't be done) should be a basic requirement for serious publications. It's disappointing that this isn't (yet?) the norm. </rant>
- deleted 8y ago[deleted]
- deleted 8y ago[deleted]
- ianai 8y agoAgreed, at least pseudo-code can be included.
- MaxBarraclough 8y agoI have to strongly disagree here. The source code itself - and not some proxy for it - must be made available. Pseudocode is just another description of the approach the authors took. It doesn't allow another researcher to carefully recreate the exact experiment that the authors ran. Software is famously hard to get right. This can be exacerbated by certain research problems where you don't know the result to expect (modelling, say). If you're going to publish results, the source you used should be available for inspection, for similar reasons to why mathematicians have to publish the proof, not just the conclusion plus a promise. The upsides here are many: protecting science against software bugs, easier replication of experiments, better protection against academic fraud, and helping other researchers extend the work. (I realise this last point may be at odds with toxic academic careerism. All the more reason for the publication process to insist on it.)
- 8y ago
- sah2ed 8y ago(A small note on your interchangeable use of reproducibility and replicability in research because they mean slightly different things. reproducible: independently achieving the same result using existing data; replicable: independently achieving the same result using new data.)