6 ms·
The number one thing in my opinion is that Stan's algorithm(s) for drawing from a posterior distribution produce samples that have much less dependence among ad
by bengHN 11y ago
The number one thing in my opinion is that Stan's algorithm(s) for drawing from a posterior distribution produce samples that have much less dependence among adjacent draws than other simpler MCMC algorithms like Metropolis-Hastings or Gibbs samplers. Since it is often difficult to answer the question "How much dependence is too much in practice?", it is prudent to use the algorithm that yields the least dependence because the effective sample size (from the posterior distribution) per unit of wall time will usually be greater. PyMC3 has started to incorporate some of Stan's algorithms, although their implementations are not as far along.
- wlievens 11y agoI know some of those words...
- zmmmmm 11y agoForget about the statistics and MCMC and all that. Imagine you have a function that is a black box and returns a floating point number. You can't see the implementation, but you want to find its maximum. How do you do it? If you're like me you'd probably start with some kind of "grid search" by giving it evenly spaced parameters. Then you would evaluate it at each one and do some kind of gradient descent / ascent type algorithm. This will work, but in a high dimensional space (ie. a function with many arguments) you have no chance of covering even a tiny fraction of the parameter space. The key is, you don't want to waste time evaluating a set of parameters if it doesn't tell you anything new. ie. if you get back a similar answer to what you already had. This is what the GP means (I think) by trying to avoid "dependence between adjacent draws". I could be wrong about all this, I'm trying to learn it too.
- wlievens 11y agoThanks, that clarifies it a bit.
- jsalvatier 11y agoI think we're (PyMC3) at parity in terms of algorithms actually.