3 ms·
From reading the abstract it seems that they are claiming that introducing some randomness into your gradient of weight changes allows for the quicker convergen
by bearzoo 12y ago
From reading the abstract it seems that they are claiming that introducing some randomness into your gradient of weight changes allows for the quicker convergence of solution - I did not read the paper. I also don't exactly understand why it works - it sounds like they are claiming traditional back prop has room for improvement.
- Houshalter 12y agoThat's a very old strategy called jittering (also see stochastic gradient descent.) This is something entirely different. They are not doing regular backpropagation at all, but somehow using neurons to learn how to backpropagate values. I haven't read the paper yet, just read their slides earlier, so that might not be correct.