5 ms·
If you read the paper https://arxiv.org/abs/1705.09558 https://arxiv.org/abs/1705.09558, in section 2.1, it defines two conditional posteriors for the parameter
by xcodevn 9y ago
If you read the paper https://arxiv.org/abs/1705.09558 https://arxiv.org/abs/1705.09558, in section 2.1, it defines two conditional posteriors for the parameters of the generator and discriminator. Then, classical GAN is just a maximum likelihood estimate (or, a MAP estimate with uniform prior) of the parameters.
Next, the paper said: how about sampling the whole posterior distribution instead of finding only one point of maximum likelihood as classical GANs do. This is the point where the famous Markov chain Monte Carlo (MCMC) algorithms become useful. They use something called stochastic gradient Hamiltonian Monte Carlo, basically, it is a random walk algorithm, at each step, you follow a noisy gradient, as a result, you converge to the posterior distribution instead of a local minima as gradient descent does.
The paper claims that sampling the whole posterior helps to resolve problems with classical GANs. IMHO, this isn't a surprise claim, this is exactly what is good about Bayesian statistics.
- carbocation 9y agoIs it fair to presume that we'll next see a variational inference approach to this, with faster but slightly less optimal results?