6 ms·
As a PhD student who sort of burned out on this type of research, I agree that the complexity of Neural Networks as a mathematical construct makes them very dif
by fogof 5y ago
As a PhD student who sort of burned out on this type of research, I agree that the complexity of Neural Networks as a mathematical construct makes them very difficult to analyze. This might also have to do with Deep learning theory being a subset of learning theory which is subject to "No Free Lunch" [1], which means that you always have to be very careful not to try to prove something that turns out to be impossible.
That being said, research on the Kernel regime is one of the very cool ideas, in my opinion, to gain traction in this field in the past few years. To summarize: "If you make a neural network wide enough, it gains the power to control its output on each individual input separately, and will begin to fit its training data perfectly". Of course, the real pleasure is in understanding all the mathematical details of this statement!
[1] : https://en.wikipedia.org/wiki/No_free_lunch_theorem https://en.wikipedia.org/wiki/No_free_lunch_theorem
- joe_the_user 5y agoI got my master's years ago so now I'm a strict amateur. That said, I don't think the "No free lunch theorem" is very "interesting". It's nearly tautological that no approximation method works for "any" function. The set of predictable/interesting/useful/"real-world" functions is going to have measure 0 compared to white noise so "any function" will basically look like white noise and can't be predicted. Approximating functions/sequences with vanishingly low Kolmogorov complexity is more interesting, impossible in general by Godel's theorem but what's the case "on average"? (depends on the choice process and so ill-defined but defining might be interesting). The kernel regime stuff looks interesting but I don't know it's relation to wide networks. Neural networks "tend to generalize well in the real world". That's a pretty fuzzy statement imo since "real world" is hardly defined but it's still what people experience and it's more useful to provide a more precise model where this works rather than a model where this doesn't work. Also, there's good theory on deep networks as universal well as theories of wide/shallow networks [1]. [1]: https://arxiv.org/abs/1901.02220 https://arxiv.org/abs/1901.02220
- roenxi 5y ago> Neural networks "tend to generalize well in the real world". I've always interpreted that as "we've found an algorithm that could, given a foreseeable amount of computing power and maybe some tweaks, simulate human decision making". It isn't so much that neural networks can approximate the real world as they can approximate human perception of the real world.
- joe_the_user 5y agoWell, I quote the statement to show how vague it is, among things. Neural networks are "universal approximators" in that they work as well as virtually any previous approximation method. So given big snapshot of input data and human judgement on it, they can approximate that. They can also approximate a snapshot of some input-output pairs not produced by human but having patterns (solutions to differential equations, for example). So, they can approximate what humans do in a given domain. But there's no reason to think they're acting in the same way as humans and I'd say very few people seriously working on ML believe that.
- make3 5y agoThis intuition is very dangerous and leads to huge misconceptions about deep neural nets. Neural nets don't learn anything like us, and they don't reproduce our functions. We build on massive amounts of general symbolic knowledge, and can zero shot tasks (without explicit examples) easily. Neural networks really should be seen as just giant random functions that you progressively modify in tiny ways until they fit your data. As parent says, we've just been lucky or good at constraining these functions in a way that they can only learn useful functions (ie convnets) or that they somehow learn these more quickly
- roenxi 5y agoHumans certainly do not build on massive amounts of symbolic knowledge because we are absolutely terrible at symbolic knowledge. Reliably reasoning through a basic logical argument is a specialist skill. Even reviewing evidence before making decisions is uncommon, most humans operate on a look -> assess -> do model where the tricky bit is well approximated by a neural net. Which is why neural nets seem to be so good at real-world tasks. It is completely plausible that when neural nets get scaled up to something approaching human-brain numbers of connections they will well approximate a human brain or be a few tweaks away. Obviously it won't be knowable until state of the art gets there, but there is no reason to think human intelligence is going to be complicated. It is one evolutionary step up from some pretty basic animals.
- Cacti 5y agoNFL theorems aren’t an argument about noise, they’re an argument about the uncountability of real numbers. NFL states that over all problems any optimization method performs equally poorly to any other, or equivalently, _that if an optimization method does well on some problems, it must do equally poorly on some other problems_, and those others aren’t necessarily noise, they could be anything. The problem is you don’t know which problems it is going to do poorly on in advance. You hope it does poorly on noise or on problems that you don’t care about, but you can’t tell. That is a very different statement than what you’re saying, and it’s as equally non-trivial as Godels and Turings statements in decidability.