3 ms·
This is primarily in regard to deep neural nets, where second-order methods are too expensive (O(p^2) in parameters, where p is on the order of millions). I'm
by highd 10y ago
This is primarily in regard to deep neural nets, where second-order methods are too expensive (O(p^2) in parameters, where p is on the order of millions).
I'm not sure conjugate gradient methods are used much, due to the non-convexity of the merit functions, and hard constraints aren't used much either.
- partykid92 10y agoquasi newton methods are not square in the dimension of the problem (think limited memory L-BFGS), and can be run in linear time. In my experience, however, they're 2-3 times slower than regular methods like ADAM.