4 ms·
Thanks for your [2] link paper, it's an interesting one. If you think about it, the high-level idea makes some intuitive sense. Since NNs are known to be "equi
by medlyyy 6y ago
Thanks for your [2] link paper, it's an interesting one.
If you think about it, the high-level idea makes some intuitive sense. Since NNs are known to be "equivalent" to kernel methods - which are, to oversimplify, essentially nearest-neighbour + a similarity function - then the ability of NNs to memorise specific training examples can be analogised to adding another "neighbour" to interpolate from. So maybe it's not too surprising that NNs which can do this can have better generalisation performance. (Although it is still surprising, since I certainly wouldn't predict it a priori!)
Really, what is the difference between a "feature" and "memorising a data point"? It seems only a matter of scale - how much of the input space is "relevant" to the learned representation.
- robbmorganf 6y ago> Since NNs are known to be "equivalent" to kernel methods I remember reading this somewhere, but I can't find the exact paper. Do you think you could link me to it? Also, I seem to recall the paper didn't provide a method to actually construct the equivalent kernel. Do you know of any work since then that actually constructs an equivalent kernel method?
- medlyyy 6y agoI believe it's this: https://arxiv.org/abs/1711.00165 https://arxiv.org/abs/1711.00165 And then the NTK paper: https://arxiv.org/abs/1806.07572 https://arxiv.org/abs/1806.07572 There is also this which seems to have been produced in parallel, closely related to the first but not exactly the same: https://arxiv.org/abs/1804.11271 https://arxiv.org/abs/1804.11271