4 ms·
Modeling Uncertainty with PyTorch
- gillesjacobs 5y agoThe field of ML is largely focused on just getting predictions with fancy models. Estimating the uncertainty, unexpectedness and perplexity of specific predictions is highly underappreciated in common practice. Even though it is highly economically valuable to be able to tell to what extent you can trust a prediction, the modelling of uncertainty of ML pipelines remains an academic affair in my experience.
- NeedMoreTime4Me 5y agoYou are definitely right; there are numerous classic applications (i.e. outside of the cutting-edge CV/NLP stuff) that could greatly benefit from such a measure. The question is: Why don’t people use these models? While Bayesian Neural Networks might be tricky to deploy & debug for some people, Gaussian Processes etc. are readily available in sklearn and other implementations. My theory: most people do not learn these methods in their „Introduction to Machine Learning“ classes. Or is it lacking scalability in practice?
- disgruntledphd2 5y agoIt takes more compute, and the errors from badly chosen data vastly outweigh the uncertainties associated with your parameter estimate. To be fair, I suspect lots of people do this, but for whatever reason nobody talks about it.
- b3kart 5y agoThey often don’t scale, they are tricky to implement in frameworks that people are familiar with, but, most importantly, they make crude approximations meaning after all this effort they often don’t beat simple baselines like bootstrap. It’s an exciting area of research though.
- shakow 5y ago> Or is it lacking scalability in practice? Only speaking from my own little perspective in bioinformatics, lack of scalability above all else, both for BNNs and GPs. Sure, the library support could be better, but that was not the main hurdle, more of a friction.
- NeedMoreTime4Me 5y agoDo you have an anecdotal guess on the scalability barrier maybe? Like does it take too long with more than 10,000 data points having 100 features? Just to get a feel.
- shakow 5y agoPlease don't quote me on that, as it was academic work in a given language and a given library and might not be representative of the whole ecosystem. But in a nutshell, on OK-ish CPUs (Xeons a few generations old), we started seeing problems past a few thousands points with a few dozens features. And not only was the training slow, but also the inference: as we used the whole sampled chain of the weights distributions parameters, not only was memory consumption a sight to behold, but inference time quickly grew through the roof when subsampling was not used. And all that was on standard NNs, so no complexity added by e.g. convolution layers.
- rsfern 5y agoThe main bottleneck in GP models is the inversion of an NxN covariance matrix, so training with the most straightforward algorithm has cubic complexity (and quadratic memory complexity). 10k instance is what I’ve seen as the limit of tractability. The input dimensionally doesn’t necessarily matter since it’s kernel method, but if you have many features and want to do feature selection or optimize parameters things can really stack up. There are scalable approximate inference algorithms, and pretty good library support (gpflow, gpytorch, etc), but it seems like they are not widely known, and there are definitely tradeoffs to consider among the different methods.
- marbletimes 5y agoWhen I was in academia, I used to fit highly sophisticated models (think many-parameters, multi-level non-linear mixed effect models) who were giving not only point estimate but also confidence and predictive intervals ("please explain to me the difference between the two" is one of my favorite interview questions and I still have not heard a correct answer). When I tried to bring an "uncertainty mindset" over when I moved to industry, I found that (1) most DS/ML scientists use ML models that typically don't provide an easy way to estimate uncertainty intervals, (2) in the industry I was in (media) people who make decisions and use model prediction as one of the input for their decision-making are typically not very quantitative and an uncertainty interval, rather than give strength to their process, would confuse them more than anything else: they want a "more or less" estimate, more than a "more or less plus something more and something less" estimate. (3) When services are customer-facing (see ride-sharing) providing an uncertainty interval (your car will arrive between 9 and 15 minutes) would anchor the customer to the lower estimate (they do for the price of rides book in advance, and they need to do it, but they are often way off). So for many ML applications, an uncertainty interval that nobody internally or externally would base their decision upon is just a nuisance.
- joconde 5y agoWhat do "multi-level" and "mixed effects" mean? There are tons of non-linear models with lots of parameters, but I've never heard these other terms.
- canjobear 5y agohttps://en.wikipedia.org/wiki/Nonlinear_mixed-effects_model https://en.wikipedia.org/wiki/Nonlinear_mixed-effects_model
- code_biologist 5y agoGreat answer. It prompts a bunch of followup questions! most DS/ML scientists use ML models that typically don't provide an easy way to estimate uncertainty intervals Not an DS/ML scientist but a data engineer. The models I've used have been pretty much "slap it into XGBoost with k-fold CV, call it done" — an easy black box. Is there any model or approach you like to estimate uncertainty with similar ease? I've seen uncertainty interval / quantile regression done using XGBoost, but it isn't out of the box. I've also been trying to learn some Bayesian modeling, but definitely don't feel handy enough to apply it to random problems needing quick answers at work.
- math_dandy 5y agoUncertainty estimates in traditional parametric statistics are facilitated by strong assumptions on the distribution of the data being analyzed. In traditional nonparametric statistics, uncertainty estimates are obtained by a process called bootstrapping. But there's a trade-off. There's no free lunch!) If you want to eschew strong distributional hypotheses, you need to pay for it with more data and more compute. The "more compute" typically involves fitting variants of the model in question to many subsets of the original dataset. In deep learning applications in which each fit of the model is extremely expensive, this is impractical.