3 ms·
Interesting paper that looks at the generalization abilities of deep CNN architectures. They begin by highlighting the generalization gap or the difference betw
by eggie5 9y ago
Interesting paper that looks at the generalization abilities of deep CNN architectures. They begin by highlighting the generalization gap or the difference between the learning curves of the test and transit and how typically, in deep CNN architectures, the gap is relatively small. They then go on to hight how that this small generalisation gap is often attributed to the fact that that deep CNNs learn high-level semantic meaning. They then counter that common notion by highlighting the recent and popular research in Adversarial Examples, and note the high sensitivity to adverbial pertubtations. If a CNN is learning semantic meaning, then why does adding static to the image break it? Also, related is the recent research by Zhang, Chiyuan, et al. “Understanding deep learning requires rethinking generalization.” in which they show that CNN arch. can perfectly fit random labels which leads us to more down the path that generalization capabilities of CNNs are currently unknown to the community. They then go on to introduce their experiment to try and isolate what CNNs are doing.
http://www.eggie5.com/129-Paper-Review-Measuring-the-tendency-of-CNNs-to-Learn-Surface-Statistical-Regularities http://www.eggie5.com/129-Paper-Review-Measuring-the-tendenc...