3 ms·
Yes, although the particular technique used here (variational autoencoder, or VAE) doesn't benefit from increasing the code space due to a particular penalty (K
by kastnerkyle 10y ago
Yes, although the particular technique used here (variational autoencoder, or VAE) doesn't benefit from increasing the code space due to a particular penalty (KL divergence against a fixed N(0, 1) Gaussian prior) during training. So even by making it go to 1000 digits in the code space, the model will probably still only choose to use 5 or 10 dimensions unless the KL penalty is relaxed, making it closer to a standard autoencoder.
An autoencoder with a much larger code space could be an improvement, or some newer works such as Adversarially Learned Inference / Adversarial Feature Learning [0, 1] or real NVP [2, 3] could probably do a much better job at the task, at the cost of increased computation.
Also something like inpainting parts of the frames with pixel RNN [4] would be interesting.
[0] http://arxiv.org/abs/1606.00704 http://arxiv.org/abs/1606.00704
[1] http://arxiv.org/abs/1605.09782 http://arxiv.org/abs/1605.09782
[2] http://www-etud.iro.umontreal.ca/~dinhlaur/real_nvp_visual/ http://www-etud.iro.umontreal.ca/~dinhlaur/real_nvp_visual/
[3] http://arxiv.org/abs/1605.08803 http://arxiv.org/abs/1605.08803
[4] http://arxiv.org/abs/1601.06759 http://arxiv.org/abs/1601.06759
- vintermann 10y agoGiven that Adversarially Learned Inference came out yesterday, the best approach may be to just wait a couple of months and see what the state of the art is then.
- kastnerkyle 10y agoSure - that is always an option especially in deep learning right now. But this current crop of models (counting in DCGAN and LAPGAN/Eyescream) has really made a leap in my eyes from before "oh cool generative model" to "are these thumbnails real?". They are really generating a lot of cohesive global structure, which is pretty awesome!