13 ms·
I think a lot of statisticians and machine learners remain to be convinced that's there's much payoff available from trying to do efficient statistical inferenc
by mjw 12y ago
I think a lot of statisticians and machine learners remain to be convinced that's there's much payoff available from trying to do efficient statistical inference in such a general setting. As the article warns, it's inherently really hard in its full generality, and I don't think anyone expects a silver bullet. It seems likely that the most general probabilistic programming tools will be strongest on:
* Problems with few parameters and/or few data (I was going to say toy problems, but there are sometimes important and interesting problems of this nature)
* Problems where the generative model is so complicated that you have no hope of doing any better than this and turn to it as a last resort. A bit like in combinatorial optimisation where you just say "gah, let's throw it at a SAT solver!".
(Perhaps that's not a bad analogy actually. If they can get to the point where SAT solvers are now, that would actually not be a bad proposition.)
* In particular, problems where the generative model is complicated but the complicated part of it is largely deterministic -- perhaps some kind of non-linear inverse problem where there's some simple additive observation noise tacked onto the end, for example.
What I do fear about, is the suggestion that people can just start building fiendishly complicated hierarchical Bayesian models using these things and get valid, useful, robust, interpretable inferences from them without much in the way of statistical training. I suspect even a lot of statisticians would be a bit scared of this sort of thing. Make sure you really read up on things like model checking and sensitivity analysis, that you know something about the trade-offs of different model structures and priors etc. And that's before you start to worry about the frequentist properties and failure cases of any approximate inference procedure which is magically derived for you.
Statisticians tend to favour simpler parsimonious models, not only for computational convenience but because it's easier to reason about them, understand and check their assumptions, understand their failure cases and so on.
I wish these guys lots of luck though, it is a really interesting area and the computer scientist in me really wants them to succeed!
- mturmon 12y agoI generally agree, but I feel like there is less skepticism than you portray among the ML/Stats community about probabilistic modeling languages (hence the DARPA call for white papers). Universal inference for well-posed models is ambitious, but a lot of people have worked hard, with significant success, to make it happen. The Geman and Geman paper from 1984 was directed at general procedures for inference using sampling, as was the work of the Brown group (Grenander et al.) throughout the 80s and 90s. BUGS (http://www.mrc-bsu.cam.ac.uk/software/bugs/ http://www.mrc-bsu.cam.ac.uk/software/bugs/) was specifically directed at that goal, and has been highly successful in real applications since the mid-1990s. Other more recent software, like Stan (http://mc-stan.org http://mc-stan.org) is also targeted at this approach. Looking at this from a different angle, the published work of many ML researchers has been directed at unifying common models with the clear view of producing software. I'm thinking about Michael Jordan's "plate" notation (http://en.wikipedia.org/wiki/Plate_notation http://en.wikipedia.org/wiki/Plate_notation) and the various unifying reviews of multi-component models and inference algorithms for time series (e.g., http://dl.acm.org/citation.cfm?id=309396 http://dl.acm.org/citation.cfm?id=309396), and the corresponding reviews for the EM algorithm done by Meng, van Dyk, and others (e.g., http://www.stat.harvard.edu/Faculty_Content/meng/JCGS01.pdf http://www.stat.harvard.edu/Faculty_Content/meng/JCGS01.pdf).
- mjw 12y agoOK, fair points. Perhaps think that came across a bit more negative than I intended it to. I think there's a lot of value (and a lot of buy-in) for probabilistic modelling languages, which impose some constraints in order to get efficient inference, which help you meet those constraints and don't provide a false promise of generality. And tonnes of value in research which looks for nice unifying formalisms to enable this sort of thing. Also in nice formalisms for inference in general, so perhaps that will be a useful side-effect of this work. It's the idea of full-blown probabilistic programming, where you have unconstrained turing-complete non-deterministic programming language and inference just works, where I've seen a bit more in the way of healthy skepticism. Of the "this will be nice if it ever arrives, but in the meantime I'll be getting on with doing statistics, which it doesn't obviate the need for". Similar to a computer scientist and the ultimate "sufficiently smart" compiler I suppose.
- tlarkworthy 12y agoI disagree. People who know how to build complex bayesian models will be over the moon they don't have to piece a complete system together using matrices. I see no reason for such a system not to scale, essentially the underlying math is elementary, but each statistical package at the moment has a kind of lock in though lack of interoperability. Ideally you want to only do MCMC as a last resort for part of the reasoning, but once you are in WIN bugs you end up doing all the basic stuff with MCMC too, as its too much of a pain to try and tie it to a more efficient system for the bits you can reason with efficiently. I see potential in models that are partially MCMC and partially analytical. I see even more potential in embedding those kinds of systems within larger smart systems. Some probabilistic standardisation would be lapped up by the applied community.
- nashequilibrium 12y agoI am enjoying going through this "Probabilistic Models of Cognition":https://probmods.org/generative-models.html https://probmods.org/generative-models.html .Even though i am a python guy but the fact that it is written using a functional probabilistic programming language called Church, really makes it easy to follow along.
- tlarkworthy 12y agoyeah looks like nice externals, but yet again, all the inference is with MCMC :s MCMC is the most general approach, but also the most inefficient, so its got limited appeal in crunching live numbers coming out a physical system. Some problems are intractable and its the only way, but really you want to limit application of MCMC as little as possible, so if you use church you now have an interoperability problem with binding (LISP) to some ugly but efficient FORTRAN system or some such. Exactly why I think think this initiative will be highly welcome!
- mjw 12y agoSo you're probably aware of this, but I think it might be interesting to try and articulate why I fear extending fully-Bayesian inference algorithms beyond MCMC could be a challenge, from experience with ML-focused Bayesian models and larger datasets: Generally to make inference for these kinds of models fast, you have to make it approximate. MCMC has the nice property that, while it's an approximate method, it becomes exact in the limit of infinite samples. Meaning that when applying it in a general setting, the compiler doesn't have to make hard and final decisions about where and how to introduce approximations -- it produces MCMC code and you decide how long you can be bothered to wait for it to converge and how accurate you want the answer to be, based on various diagnostics. Most other non-MCMC-based approximate inference methods (MAP, EM, Variational Bayes, EP, various hybrids of these and MCMC...) don't converge to the exact answer, they converge to an approximate answer which remains inexact no matter how long you run the algorithm for. Different approximations have different strengths and weaknesses, and the best choice may depend on the model, the data, what you actually want to do with the posterior (a mode, a mean, an expected loss, predictive means, extreme quantiles?), and what frequentist properties you want for the results of the inference, especially given that this is no longer pure Bayesian inference but some messy approximation to it. Often the only practical way to decide will be to try a bunch of different things and see which does best on the application (or best predicts held-out data), even armed with full human intuition. A compiler doesn't really have a chance here. In short, it's not going to be easily to automate fully because it's not something you can decide on formal grounds alone. You have to decide how and where you're willing to approximate, and you have to understand the approximations used and check the resulting approximated posterior against data in order to know whether to trust the results and how to interpret them.