3 ms·
Yes they do (arxiv is a place for scientific papers not press releases). I've only skimmed it, but the paper introduce an adaptive way to clip gradients. Meanin
by belgian_guy 6y ago
Yes they do (arxiv is a place for scientific papers not press releases). I've only skimmed it, but the paper introduce an adaptive way to clip gradients. Meaning that if the ratio of the gradient norm to weight norm surpasses a certain threshold, they clip it. This stabilizes learning and seems to avoid the need for batch normalization. Seems quite promising imo and something that could stick (I'm quite happy if we could finally do away with batchnorm).
- eggie5 6y agoYou missed a big part: they did a big NAS run to make it work.
- rsfern 6y agoWhere did you see that they used NAS? Their preliminary results show it works even for the baseline model they did a lot of manual hyperparameter optimization, and spend a fair amount of time unpacking the rationale for their choices, including a negative results section (!) in the appendix