3 ms·
It's a little easier to see what's happening if you fully write out the central flow: -1/η * dw/dt = ∇L - ∇S * ⟨∇L, ∇S⟩/‖∇S‖² We're projecting the loss gr
by programjames 1y ago
It's a little easier to see what's happening if you fully write out the central flow:
-1/η * dw/dt = ∇L - ∇S * ⟨∇L, ∇S⟩/‖∇S‖²
We're projecting the loss gradient onto the sharpness gradient, and subtracting it off. If you didn't read the article, the sharpness S is the sum of the eigenvalues of the Hessian of the loss that are larger than 2/η, a measure of how unstable the learning dynamics are.
This is almost Sobolev preconditioning:
-1/η * dw/dt = ∇L - ∇S = ∇(I - Δ)L
where this time S is the sum of all the eigenvalues (so, the Laplacian of L).
- lcnielsen 1y agoYeah, I did a lot of traditional optimization problems during my Ph. D., this type of expression pops up all the time with higher-order gradient-based methods. You rescale or otherwise adjust the gradient based on some system-characteristic eigenvalues to promote convergence without overshooting too much.