3 ms·
EBMs actually associate an "energy" to each point of the input distribution which then defines a probability distribution through the Boltzmann Distribution. It
by yilundu 8y ago
EBMs actually associate an "energy" to each point of the input distribution which then defines a probability distribution through the Boltzmann Distribution. It's true that Langevin dynamics get stuck at low-probability modes and it would be worth drying with an adaptive version of HMC. However, since we initialize chains with a random prior distribution, each individual chain is individually likely to hit any mode so all modes are likely to be explored.
- lenticular 8y ago>EBMs actually associate an "energy" to each point of the input distribution which then defines a probability distribution through the Boltzmann Distribution. Yes, this is precisely what MCMC methods do as well. Every posterior distribution is a Boltzmann distribution for some energy function. >However, since we initialize chains with a random prior distribution, each individual chain is individually likely to hit any mode so all modes are likely to be explored. This is also a pretty standard technique in MCMC. But most high-dimension Bayesian models have a huge amount of modes that that cannot be explored in a reasonable number of samples/chains.
- yilundu 8y agoRight and this arguably is especially the case for high dimensional image datasets. Yet despite this case, we are able to train models on these high dimensional dataset through MCMC (with some tricks) with good likelihood, indicating that standard MCMC technique can actually scale up to very high dimensional multi-modal situations which was previously thought to be computationally intractable.