5 ms·
hi. i'm one of the Stan devs. (i work on variational inference: ADVI). happy to answer any questions here.
by proditus 11y ago
hi. i'm one of the Stan devs. (i work on variational inference: ADVI).
happy to answer any questions here.
- chmullig 11y agoYay Stan! So what's the deal with BBVI? I heard it's integrated in the newest Stan, but can I use it? What for?
- proditus 11y agoglad to see the enthusiasm! ADVI [1] is a variant of BBVI [2] where we fully leverage all of the amazing things that Stan has to offer (like automatic differentiation and automatic transformations of constrained parameters). you can use ADVI to get an approximation to the Bayesian posterior. the advantage of using ADVI over sampling is that ADVI is typically faster for large models (both in terms of # of parameters and # of data observations). ADVI also a bit better at handling models with multi-modal posteriors, such as mixture models. ADVI is currently in cmdStan (but not in RStan or PyStan). we're continuing to make the algorithm more robust. [1] http://arxiv.org/abs/1506.03431 http://arxiv.org/abs/1506.03431 [2] http://www.jmlr.org/proceedings/papers/v33/ranganath14.pdf http://www.jmlr.org/proceedings/papers/v33/ranganath14.pdf
- marmaduke 11y agohi, funny this came up on HN; I'm just trying to get a handle on fitting a dynamical systems model of neural activity propagation with Stan. What in general are the disadvantages of variational inference vs full MCMC? Are there in general significant advantages besides the speed up?
- proditus 11y agothat's a really good question. some aspects of VI vs MCMC are areas of active research. so it's tough to respond succinctly, but i'll try. the key disadvantages of VI (particularly ADVI) are: 1. mean-field variational inference cannot model posterior correlations. so if you expect your model + dataset to give a "skewed" posterior, then mean-field variational inference will have a difficult time describing such a posterior. (it will under estimate marginal variances.) 2. full-rank variational inference can model posterior correlations. but it can become too expensive for big models. there is a lot of great research coming up in this vein, such as [1,2]. 3. in either case, the version of variational inference we have in Stan (ADVI) uses a normal approximation in a transformed parameter space. thus, there is an additional mismatch of the shape of the variational posterior to the full MCMC posterior. in terms of advantages: 1. variational inference is (in general) a non-convex optimization problem. so it's easy to know when we've converged to a local optimum. convergence in MCMC is a bit more tricky to assess. 2. if your model has a multi-modal posterior, then variational inference will focus on just one of the modes. this is sometimes desirable as MCMC techniques might end up jumping around all of the modes and producing poor samples. this is just the tip of the iceberg. but i hope it helps! [1] http://arxiv.org/pdf/1506.03159.pdf http://arxiv.org/pdf/1506.03159.pdf [2] http://arxiv.org/pdf/1502.07685.pdf http://arxiv.org/pdf/1502.07685.pdf