3 ms·
Super cool stuff in this paper. At its heart, this is a new training architecture that allows parameter weights to be updated faster in a distributed setting.
by nicklo 10y ago
Super cool stuff in this paper.
At its heart, this is a new training architecture that allows parameter weights to be updated faster in a distributed setting.
The speed-up happens like so: instead of waiting for the full error gradient to propagate through the entire model, nodes can calculate the local gradient immediately and estimate the rest of it.
The full gradient does eventually get propagated, and it is used to fine-tune the estimator, which is a mini-neural net in itself.
Its amazing that this works, and the implications that full back-prop may not always be needed shakes up a lot of assumptions about training deep nets. This paper also continues the trend of this year of using neural nets as estimators/tools to improve the training of other neural nets. (I'm looking at you GANs).
Overall, excited to see where this goes as other researchers explore the possibilities when you throw the back-prop assumption out.