3 ms·
> How well does a classic deep net architecture like AlexNet or VGG19 classify on a standard dataset such as CIFAR-10 when its “width”— namely, number of channe
by PartiallyTyped 3y ago
> How well does a classic deep net architecture like AlexNet or VGG19 classify on a
standard dataset such as CIFAR-10 when its “width”— namely, number of channels
in convolutional layers, and number of nodes in fully-connected internal layers —
is allowed to increase to infinity? Such questions have come to the forefront in the
quest to theoretically understand deep learning and its mysteries about optimization
and generalization. They also connect deep learning to notions such as Gaussian
processes and kernels. A recent paper [Jacot et al., 2018] introduced the Neural
Tangent Kernel (NTK) which captures the behavior of fully-connected deep nets in
the infinite width limit trained by gradient descent; this object was implicit in some
other recent papers. An attraction of such ideas is that a pure kernel-based method
is used to capture the power of a fully-trained deep net of infinite width.
https://openreview.net/pdf?id=rkl4aESeUH https://openreview.net/pdf?id=rkl4aESeUH, https://github.com/google/neural-tangents https://github.com/google/neural-tangents
> It has long been known that a single-layer fully-connected neural network with an i.i.d. prior over its parameters is equivalent to a Gaussian process (GP), in the limit of infinite network width.
https://arxiv.org/abs/1711.00165 https://arxiv.org/abs/1711.00165
And of course, one needs to look back at SVMs applying a kernel function and separating with a line, which looks a lot like an ANN with a single hidden layer followed by a linear mapping.
https://stats.stackexchange.com/questions/238635/kernel-methods-how-do-the-infinite-dimensions-arise https://stats.stackexchange.com/questions/238635/kernel-meth...