3 ms·
You're talking about second order optimization. The full-fat version is too slow (the Hessian, the matrix of second derivatives, is huge for DL models), the opt
by beecafe 4y ago
You're talking about second order optimization. The full-fat version is too slow (the Hessian, the matrix of second derivatives, is huge for DL models), the optimized approximations don't seem to provide a meaningful benefit over first order methods like SGD (or more commonly, Adam).
A more promising path is to regularize the loss function so that SGD can optimize it quickly, rather than switching to a second order method.
- p1esk 4y agoYou're talking about second order optimization No. I'm asking for a new kind of math for ML, inspired by the paper in this post. I would like to treat a dataset as a curve, which maps to another curve (loss landscape), under the model architecture constraints, and then find the minima on that landscape in one shot. There might be some kind of interpolation possible which has nothing to do with SGD optimization (of any order). When I learn something new, for example a difficult topic like functional programming, or signal processing, the process feels nothing like SGD. At first I gather some background information, motivation, goals, etc. Then I unpack the concept/formula, and try to understand how it works, how the pieces fit together. Could be bottom up, or top-down. The information is accumulated until it reaches critical mass, and then boom - I get it. That stage happens quickly, it's an "aha" moment. This learning process involves both statistical pattern matching, and some other mechanism, perhaps referring to a rule based system like a decision tree, or accessing some facts in my long term memory database. I have no idea how a brain does it, but clearly I don't need a thousand examples to get Fourier transform, or monads, or whatever it is I'm learning about. Often just one example is sufficient to form new patterns and rules in my mind. I'm guessing this is because we are not randomly probing the loss landscape in the dark. We might be interpolating it, filling in the blanks. I do, however, need a thousand examples/repetitions to learn a new motor skill, like playing tennis or a piano, or even learning a new language - training my ear for many hours until I start recognizing the right sound patterns, so there are types of learning I do which might be similar to SGD. But most mental concepts are not learned like that.