4 ms·
Unfortunately this talk is kind of dated already. Most people don't stack RBMs or autoencoders to pretrain the weights anymore. If you use dropout with rectifie
by rudyl313 11y ago
Unfortunately this talk is kind of dated already. Most people don't stack RBMs or autoencoders to pretrain the weights anymore. If you use dropout with rectified linear units, you don't have to pretrain, even for large architectures.
- bradneuberg 11y agoIt's not just ReLUs that have helped, its also better random initialization before starting training, such as using Xavier initialization (http://andyljones.tumblr.com/post/110998971763/an-explanation-of-xavier-initialization http://andyljones.tumblr.com/post/110998971763/an-explanatio...) Also, batch normalization helps with convergence as well. In addition, LSTMs work when dealing with recurrent neural nets.