4 ms·
this stuff existed for decades and it hit a wall. sampling is slow. general AI needs structured prediction. yes, graphical models can do that (HMM, CRF etc.)
by botexpert 10y ago
this stuff existed for decades and it hit a wall.
sampling is slow.
general AI needs structured prediction. yes, graphical models can do that (HMM, CRF etc.) but inference starts to get slow and special case implementations are required for different domains. [1]
I see no way for someone getting automatic inference and training for [1] with a probabilistic programming language.
given the new deepmind paper on discovering shortest path algorithms, it's quite clear that structured predicition assisted by deep networks works quite well (this was demonstrated by a vast array of work) and graphical models represented by probabilistic programming languages are far away from being that successful.
[1]: http://www.philkr.net/home/densecrf http://www.philkr.net/home/densecrf
- eli_gottlieb 10y ago>this stuff existed for decades and it hit a wall. >sampling is slow. Right. Sampling is slow. That's why automated variational inference is a very active field of research these days: instead of approximating the posterior by sampling, you approximate it with an optimization problem whose gradient-descent provides a bound on the posterior probabilities across the parameter space. All the work on training deep neural networks has made our hardware and software very efficient at solving optimization problems.
- twiecki 10y agoI don't think it has to be either or. For example, you can implement a Deep Net in a probabilistic programming framework (http://twiecki.github.io/blog/2016/07/05/bayesian-deep-learning/ http://twiecki.github.io/blog/2016/07/05/bayesian-deep-learn...) like PyMC3. The inference here is not using sampling but rather ADVI (http://pymc-devs.github.io/pymc3/api.html#advi http://pymc-devs.github.io/pymc3/api.html#advi) which is almost as general but much much faster, and can be run on sub-samples of the data (similar to stochastic gradient descent used in deep learning). Once we bridge these two domains we can get the best of both worlds, like a deep net HMM.