2 ms·
I completely agree, we shouldn't be depending on our optimizers do some approximate Bayesian inference- an optimizer should optimize only. However, I think it'
by Straw 7y ago
I completely agree, we shouldn't be depending on our optimizers do some approximate Bayesian inference- an optimizer should optimize only.
However, I think it's a different effect- even purely in terms of optimizing the training loss, on a quadratic (with noisy gradients), the short-horizon bias effect exists.