3 ms·
Curious (possibly naive) question: isn't there a fundamental difference between the goals behind creating models with ML vs the "old-fashioned" way? That is, in
by willj 8y ago
Curious (possibly naive) question: isn't there a fundamental difference between the goals behind creating models with ML vs the "old-fashioned" way? That is, in modern ML applications, you're creating a model with dozens/hundreds of potential variables, without a hypothesis of how they relate or contribute to the target (other than that they might, hence your including them in the modeling process). You're using the model for predictions more than for explainability (though there is work ongoing into improving explainability, but it seems kind of post hoc to me). And there's an expectation that you will retrain, or at least tune, the model as its predictive accuracy decays over time.
By contrast, traditionally in science you're coming in with a hypothesis ahead of time about what variables predict what target. The goal is to come up with a model that is consistent with your hypothesis (and possibly some existing theory), and which can be applied generally, and which should need no tuning. For example, the very simple model for Beer's Law-- absorbance vs concentration. That is a law that will apply in every other circumstance, but if modern ML methods had been applied, the scientist might have chosen the model with a slightly better score but which includes nonsense variables in addition to concentration.
All that to say, it seems to me the problem stems from scientists' lack of hypotheses at the outset of a project, and/or the understandable desire to get the best bang for their buck out of an experiment by measuring dozens of variables at once and hoping the magic of ML can find a hypothesis for them.
Hope that made sense.
- fock 8y agoI think you got the point. A lot of people don't seem to realize that ML might be great for finding patterns but will never yield scientific knowledge in the sense of cause-reaction sense. Unfortunately everyone thinks he can use it for finding "new stuff" and so in my field they "predict material properties", etc. using ML fed with data where every review about the physics tells you that the algorithms they use for extracting that data are domain-specific and might yield results different on the order of magnitudes. But nobody cares; take some SW off the net, which claims to be able to extract what you want, run it, train your ML, publish your results.
- ppod 8y agoWhat method would you use to yield scientific knowledge in the sense of cause-reaction? Many important processes really do have large numbers of causal factors that interact non-linearly. If we want to try to learn about that, some statistical method that deals with many parameters will be needed. Such models are generally referred to as "Machine Learning". Their generalisation or causal inference properties are particular to each implementation and identification strategy, but you can't just say "ML will never yield scientific knowledge".
- fock 8y agoA tool called Mathematics which can exactly describe this interactions. And if those processes have a lot of variables, a ML-model might certainly be useful, but it will never be generally applicable! This probably also contributes to "scientific knowledge" but it's not the same as scientific facts (or whatever you call universally transferable results).
- ppod 8y agoYou're building your mathematical model based on the knowledge you have, which is from the data you have, and there is still the same risk that your theory won't generalize to new observations.
- hannob 8y ago> By contrast, traditionally in science you're coming in with a hypothesis ahead of time about what variables predict what target. That is an idealistic view of what science should be, it's not what happens in the real world. HARKing ("Hypothetizing after the results are known") was a thing before ML was cool. But ML is amplifying that, it's a more effective tool to perform bad science.
- olooney 8y agoI would guess both you and author of the article have in mind something like gene expression[GE] in bioinformatics. Thousands of markers, but only hundreds of examples, and the researchers are using some automated feature selection approach[LB] to pick genes that predict some disease. [GE]: https://en.wikipedia.org/wiki/Machine_learning_in_bioinformatics#Microarrays https://en.wikipedia.org/wiki/Machine_learning_in_bioinforma... [LB]: https://www.quora.com/How-is-Lasso-method-used-in-bioinformatics-and-systems-biology https://www.quora.com/How-is-Lasso-method-used-in-bioinforma... Obviously if this were done sloppily it would be a huge problem and could produce a ton of false positives. But that's not actually what happens. The idea that ML practitioners just fit crazy complicated models to data and blindly believe whatever the model fits seems to be a common stereotype but is completely inaccurate. We are acutely aware that powerful models can overfit all to easily and spend perhaps the majority of our time understanding and fighting this exact phenomena. Because we tend to work with models for which few closed-form analytic theorems exist, we tend to do this empirically but no less rigorously. In fact, we tend to be more scientific and rely on fewer assumptions than classical statistics. The dominant paradigm is empirical risk minimization, sometimes called structural risk minimization[SRM], especially when complexity is being penalized. The idea is to acknowledge that models are always fit to one particular sample from the population but that the goal is to generalize to the full population. We can never truly evaluate a model on a whole population, but we can form an empirical estimate for how well our model will do by taking a new sample from the population (not used for fitting/training) and evaluating model performance on this new sample. Computational learning theories such as VC Theory[VC] and Probably Approximately Correct Learning[PAC] provide theorems that give bounds on how tight these empirical bounds are. For example, VC Theory and Hoeffding's Inequality[HI] can give us an upper bound on how large the gap between "true" performance and this empirical estimate is for a binary classifier in terms of the number of observations used to measure performance and the "VC Dimension" (roughly the number of parameters) of the model. A typical SRM workflow would be to divide a data set up into "training," "validation," and "test" sets, fit a set of candidate models to the training set, estimate their performance from the validation set, select the best based on validation set performance[MS], then evaluate the final model performance from the test set. This procedure can be used on arbitrary models to demonstrate the validity of fit models. For example, a model which is just randomly picking 5 genes based on noise in the training set is extremely unlikely to perform better than chance on the final test set. [SRM]: http://www.svms.org/srm/ http://www.svms.org/srm/ [VC]: https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_theory https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_th... [PAC]: https://en.wikipedia.org/wiki/Probably_approximately_correct_learning https://en.wikipedia.org/wiki/Probably_approximately_correct... [HI]: https://people.cs.umass.edu/~domke/courses/sml2010/10theory.pdf https://people.cs.umass.edu/~domke/courses/sml2010/10theory.... [MS]: https://en.wikipedia.org/wiki/Model_selection https://en.wikipedia.org/wiki/Model_selection Not every machine learning practitioner is familiar with VC Theory or PAC, but almost everyone uses the practical tools[CV] and language[BV] that arose from SRM. If you're following Andrew Ng's or Max Kuhn's advice[NG][MK] on "best practices" you are in fact benefiting from VC Theory although you may never have heard of it. [VC]: https://en.wikipedia.org/wiki/Cross-validation_(statistics) https://en.wikipedia.org/wiki/Cross-validation_(statistics) [BV]: https://en.wikipedia.org/wiki/Bias%E2%80%93variance_tradeoff https://en.wikipedia.org/wiki/Bias%E2%80%93variance_tradeoff [NG]: https://www.youtube.com/playlist?list=PLA89DCFA6ADACE599 https://www.youtube.com/playlist?list=PLA89DCFA6ADACE599 [MK]: http://appliedpredictivemodeling.com/ http://appliedpredictivemodeling.com/ So that's my answer to the question of validity: ML researchers use different techniques, but their techniques have equally good theoretical foundations but make very few assumptions and are very robust in practice. If researchers aren't using these techniques, or abusing them, it's not because ML is unsatisfactory or broken, but because of the same perverse incentives we see everywhere in academia. There's another criticism floating around that ML models are "black boxes", useful only for prediction and totally opaque. This is only true because non-linear things are harder to understand, and to the extent to which it is true, it is equally true of classical models. A linear model with lots of quadratic and interaction terms, or a model on stratified bands, or a hierarchical model, can be just as hard to interpret. A properly regularized ML model only fits a crazy non-linear boundary when the data themselves require it. A classical model fit to the same data will either have to exhibit the same non-linearity or will be badly wrong. A lot of researcher papers are wrong because someone fit a straight line to curved data! I also think the "total opaque black box" meme is overstated. We can often understand even very complex models to some degree with a little effort. A basic technique is to run k-means with high k, say, 100, to select a number of "representative" examples from your training set and look at the model's predictions for each. It's also incredibly instructive just to look at a sample of 100 examples the model got wrong. One way to understand a non-linear response surface is by focusing in on different regions where the behavior is locally linear and trying perturbations[LIME]. There are also ML models which do fit easy to understand models[MARS]. It's also usually possible to visualize the low level features[DFV]. [LIME]: https://www.oreilly.com/learning/introduction-to-local-interpretable-model-agnostic-explanations-lime https://www.oreilly.com/learning/introduction-to-local-inter... [MARS]: https://en.wikipedia.org/wiki/Multivariate_adaptive_regression_splines https://en.wikipedia.org/wiki/Multivariate_adaptive_regressi... [DFV]: https://distill.pub/2017/feature-visualization/ https://distill.pub/2017/feature-visualization/
- analog31 8y agoFor decades, folks in the area of industrial quality control have used methods called "design of experiments" (DOE), that could be loosely described as fitting minimalistic experimental data sets to arbitrary functions (typically low order multivariate Taylor polynomials). The tools are usually packaged for use by people who don't have the math background to understand the underpinnings. In the results of DOE's, I've seen everything that is now generalized as the "replication crisis." Everything I've learned about ML so far (granted not a huge amount) invokes DOE -- fitting data sets to arbitrary functions whose form is more flexible than a Taylor polynomial but otherwise cut from the same cloth. I've seen exactly the same problem as with ML, but 30 years ago: It can help you optimize a process that you don't understand, turning it into a better process that you also don't understand. But it can't tell you how something works. DOE seems to be a microcosm of ML, with all of the pitfalls such as overfitting and underfitting.