6 ms·
I am increasingly worried with people applying ML in everything without any rigour. Statical inference generally only works well in very specific conditions:
by carlosf 6y ago
I am increasingly worried with people applying ML in everything without any rigour.
Statical inference generally only works well in very specific conditions:
1 - You know the distribution of the phenomenon under study (or make an explicit assumption and assume the risk of being wrong)
2 - Using (1), you calculate how much data you need so you get an estimation error below x%
Even though most ML models are essentially statistics and have all the same limitations (issues with convergence, fat tailed distributions, etc...) it seems the industry standard is to pretend none of that exists and hope for the best.
IMO the best moneymaking opportunities in the decade will involve exploiting unsecured IOT devices and naive ML models, we will have plenty of those.
- CabSauce 6y agoWait until you find out low many studies have been published in medical journals with serious statistical flaws.
- iagovar 6y agoML looks (for many peole) like a way to circunvent your grumpy statiscian saying that the underlying data is worthless and/or you should focus on getting the data pipeline done properly for a logit model on your churn rate.
- analog31 6y ago"Scientist free science," -- being able to optimize systems without understanding them, has been a dream of the business world since the dawn of time. There's always been a market for cookbook recipes that automate the collection of data, and interpretation of results. Before ML, there were "design of experiments," and "statistical quality control."
- carlmr 6y ago>Before ML, there were "design of experiments," and "statistical quality control." Statistical quality control, at least the way I know it, is very useful in finding problems in your process. I'm also not sure how this fits with your premise. It's about optimizing systems by first finding out where to look, and then looking there in detail with expert knowledge, i.e. deep understanding of your system.
- analog31 6y agoI'm definitely with you there, but I've also seen the side of it where it turns into a cargo cult and runs headlong into the replication crisis. Perhaps the good thing is that as the new things gain popular attention, the old techniques such as SPC are under less pressure to support success theater, and revert to being actual useful, solid tools.
- astrophysician 6y agoI agree -- as ML becomes increasingly easy to be applied by non-experts or people without a heavy math/stats background, I've seen an increasing volume of arguments against the data science profession (someone the other day called DS the "gate-keepers") but: there be dragons. Anyone can use SOTA deep learning models today, but in my experience, it's more important to understand the answer to "what are the shortcomings/consequences of using a particular method to solve this problem?" "what is (or could be) biases in this dataset?", etc. It requires a non-trivial understanding of the underlying methodology and statistics to reliably answer these questions (or at least worry about them). Can you apply deep reinforcement learning to your problem? Maybe. Should you? Well, it depends, and you should understand the pros and cons, which requires more than just the knowledge of how to make API calls. There are consequences to misusing ML/AI, and they may not even be obvious from offline testing and cross validation.
- erichahn 6y agoIsn't the point of ML exactly that you don't know the underlying distribution? How is this ever assumed in any way? ML is not parametric statistics.
- nonameiguess 6y ago(Some) ML is non-parametric, but there are always some questions you need to be able to answer about your data. At bare minimum, is the generating process ergodic, what is the error of your measurement procedure, how representative of the true underlying distribution is your sampling procedure? All use of data should start with some exploratory analysis before you ever get to the modeling stage. Once you have a model, at minimum understand how to tune for the tradeoffs of different types of error and don't naively optimize for pure accuracy. At the obvious extremes, if you're trying to prevent nuclear attack, false negatives are much more costly than false positives, if you're trying to figure out whether to execute someone for murder, false positives are much more costly than false negatives. Understand the relative costs of different types of error for whatever you're trying to predict and proceed accordingly.
- contravariant 6y agoWell, all optimization problems are equivalent to a maximum likelihood estimate for a corresponding probability distribution so you may make more implicit assumptions than you think. Typical ML methods just have a huge distribution space that can fit almost anything from which they pick just 1 option. This has two downsides: Since your distribution space is several times too large by design you lose the ability to say anything useful about the accuracy of your estimate, other than that it is not the only option by far. Since you must pick 1 option from your parameter space you may miss slightly less likely explanations that may still have huge consequences, which means your models tend to end up overconfident.
- erichahn 6y agoI mean yes, there is parametric ML (maximum likelihood, MAP, GMMs, ...) and there is non-parametric ML (everything neural network, SVM, GBM, random forrests, ...). I'd argue that the latter had bigger success in the past since the prior on the data distribution is usually wrong in real life. Think about a prior for image data distributions or the same in nlp. Forget about it.
- currymj 6y agoi think this actually gets at what makes applied ML distinct from statistics as a practice, even though there is a ton of overlap. statisticians make assumptions 1 and 2, and think of themselves as trying to find the "correct" parameters of their model. people doing applied ML typically assume they don't know 1 (although they might implicitly make some weak assumptions like sub-gaussian to avoid fat tails, etc.) and also typically don't care about being able to do 2. and they don't care about their parameters; in a sense to an ML practitioner, every parameter is a nuisance parameter. instead you assume you have some reliable way of evaluating performance on the task you care about -- usually measuring performance on an unseen test set. as long as this is actually reliable, then things are fine. but you are right that in the face of a shifting distribution or an adversary crafting bad inputs, ML models can break down -- but there is actually a lot of research on ways to deal with this, which will hopefully reach industry sooner rather than later.
- RobinL 6y agoYes - this is pretty much exactly how I explain the difference between machine learning and statistics. Despite using similar models, the expertise required for 'doing statistics' (statistical inference) is actually very different from machine learning. Machine learning fits into the 'hacker mentality' well - try stuff out see what works. To do statistical inference effectively, you really do need to spend time learning the theory. They both require deep skills - but the skills are surprisingly different considering it's often the same underlying model.
- nickforr 6y agoBut without some statistical knowledge, isn’t there a risk of a lack of understanding about the robustness of “what works”?
- RobinL 6y agoyeah, agreed - a good understanding of the model's statistical assumptions can often help you make the model more robust and also give you ideas for what types of feature engineering are likely to work.
- kvathupo 6y agoAs currymj commented, this isn't accurate for ML, only for classical statistics. In ML (or more specifically deep learning), we make no distribution-based assumptions, other than the fundamental assumption that our training data is "distributed like" our test data. Thus, there aren't issues with fat-tailed distributions since we make no such normality assumptions. Indeed, with the use of autoencoders, we don't assume a single distribution, but rather a stochastic process. I suppose you could say statistics is less "empirical" than ML in the sense that it is axiom-based, whether that is a normality assumption of predictions about a regression line or stock prices following a Wiener process. By contrast, ML is less rationalist by simply reflecting data.
- mochomocha 6y agoI agree that ML tends to put weaker assumptions on the data than classical statistics and that it's a good thing. However most ML certainly makes distributional assumptions - they are just weaker. When you're learning a huge deep net with an L2 loss on a regression task, you have a parametric conditional gaussian distribution under the hood. It's not because it's overparametrized that there's no distributional assumption. Vanilla autoencoders are also working under a multivariate gaussian setup as well. Most classifiers are trained under a multinomial distribution assumption etc. And fat-tailed distributions are definitely a thing. It's just less of a concern for the mainstream CV problems on which people apply DL.
- peytn 6y agoI dunno, there are definitely distribution-based assumptions—good luck working with skewed data. Most old-school techniques are kinda additive, so nobody's really been assuming a single distribution for practical applications. Current ML techniques just work well for the kinds of problems people are applying them to, which is kind of a tautology. We should definitely seek to understand the theory behind stuff like dropout and not consider our lack of understanding a strength.
- fractionalhare 6y ago> In ML (or more specifically deep learning), we make no distribution-based assumptions, other than the fundamental assumption that our training data is "distributed like" our test data. Okay, so that's about the same as classical statistics. You're just waiving the requirement to know what the distribution is. You are still assuming there exists a distribution and that it holds in the future when you apply the model. Sure you may not be trying to estimate parameters of a distribution, but it is still there and all standard statistical caveats still apply. > Indeed, with the use of autoencoders, we don't assume a single distribution, but rather a stochastic process. Classical statistics frequently makes use of multiple distrutions and stochastic processes.
- rademacher 6y agoThe problem is high dimensions knowing the distribution or even characterizing it fully with data is incredibly difficult (curse of dimensionality). I think the real assumption in ML is just that there is some low dimensional space that characterizes the data well and ML algorithms find these directions where the data is constant.
- sfink 6y agoPersonally, I think the main problem with ML is simpler: it works well for interpolation, and is crap for extrapolation. If the outputs you want are well within the bounds of your training data set, ML can do wonders. If they aren't, it'll tell you that in 20 years everyone will be having -0.2 children and all the other species on the planet will start having to birth human babies just so they can be thrown into the smoking pit of bad statistical analysis.
- carlosf 6y agoI agree, but that's equivalent to my original claim. Being bad at extrapolation is a consequence of assuming all training data can describe your phenomena distribution and being wrong.
- sidpatil 6y ago> If they aren't, it'll tell you that in 20 years everyone will be having -0.2 children and all the other species on the planet will start having to birth human babies just so they can be thrown into the smoking pit of bad statistical analysis. https://xkcd.com/605/ https://xkcd.com/605/
- clircle 6y agoOutside of simple time series, I'm not aware of any good way to extrapolate.
- rsfern 6y agoOne way to extrapolate is to use a mechanistic or semi-mechanistic model. The recent advances in neural differential equations are a really cool example of this
- cambalache 6y ago> You know the distribution of the phenomenon under study If you know the distribution of the phenomenon under study you dont need ML, that is what probability is for. > or make an explicit assumption and assume the risk of being wrong No.You have the Bias/Variance tradeoff here.You can make an explicit assumption about your model or not. > Using (1), you calculate how much data you need so you get an estimation error below x% This is extremely complicated for anything except the most trivial toy examples, probably not solvable at all and definitely not the way biological intelligent systems (aka some humans) do it.
- bigbillheck 6y ago> 1 - You know the distribution of the phenomenon under study (or make an explicit assumption and assume the risk of being wrong) Nonparametric methods say 'hi'.
- boilerupnc 6y ago[Disclosure: I'm an IBMer - not involved with this work] With regard to exploitation, IBM research has done some interesting work in the form of an open source "Adversarial Robustness Toolbox" [0]. "The open source Adversarial Robustness Toolbox provides tools that enable developers and researchers to evaluate and defend machine learning models and applications against the adversarial threats of evasion, poisoning, extraction, and inference." It's fascinating to think through how to design the 2nd and 3rd order side-effects using targeted data poisoning to achieve a specific outcome. Interestingly, poisoning could be to force a specific outcome for a one-time gain (e.g. feed data in a way to ultimately trigger an action that elicits some gain/harm) or to alter the outcomes over a longer time horizon (e.g. Teach the bot to behave in a socially unacceptable way) [0] https://art360.mybluemix.net/ https://art360.mybluemix.net/