3 ms·
Maybe the additional parameters give the entire network more "leeway" to find such subnetwork structures, i.e. ease gradient descent by smoothing the loss lands
by longtom 7y ago
Maybe the additional parameters give the entire network more "leeway" to find such subnetwork structures, i.e. ease gradient descent by smoothing the loss landscape?
- Akababa 7y agoThat's intuitive but doesn't support the result of the lucky subnetwork (once found and re-initialized) training faster and outperforming the original.
- longtom 7y agoThis does not seem to be a contradiction. Once you are in the right region of solution space training is expected to be faster and easier. Re-initialization could have a regularizing effect, explaining the better performance.
- Akababa 7y agoThey re-use the same initialization, so it appears that the initial weights are inherently coupled with the nonzero structure.