3 ms·
"Cheap iterations" led me to expect a 1st order method, not requiring computing the matrix of 2nd derivatives, which is the bottleneck in large systems. But thi
by rwilson4 5y ago
"Cheap iterations" led me to expect a 1st order method, not requiring computing the matrix of 2nd derivatives, which is the bottleneck in large systems. But this does indeed rely on the Hessian, so I'm not sure what makes the iterations "cheap".
- MauranKilom 5y agoIf you don't want to compute the Hessian, you're at the wrong address with Newton's method. Quasi-Newton methods is where you want to look.
- rwilson4 5y agoSure, or Nesterov's accelerated gradient descent or something. I'm just not sure what the authors mean by "cheap iterations". Update: your other comment clarified this for me. It's cheap relative to other methods that work with non-self-concordant objectives.