3 ms·
> Our sharp rate depends on a key observation — although we don’t know the shape of the stuck region, we know it is very thin. Oh... really? :) (After 12 year
by cool_username 9y ago
> Our sharp rate depends on a key observation — although we don’t know the shape of the stuck region, we know it is very thin.
Oh... really? :)
(After 12 years I finally get an excuse to show a fun side project I coauthored during my PhD...)
http://graemebell.net/pubs/taros05-bl-embedded-preprint.pdf http://graemebell.net/pubs/taros05-bl-embedded-preprint.pdf
Check out Figure 5 / Section 3.4
The rest of the paper is an introduction to why saddle points can be surprisingly problematic for people using potential fields (neural nets, game/robot navigation). Hope someone finds it interesting.
- Eridrus 9y agoThese results seem important for nonconvex optimization in general, but for ML applications where we usually use stochastic/batch gradient descent, I wonder if the the stochasticity adds enough perturbation for this to not really be that useful.
- Cacti 9y agoWell, in general, no, there is no perturbation method large (or good) enough to get out of saddle points, including via stochastic gradient descent. It might work in a particularly specific problem, but not in general.
- Eridrus 9y agoFine, but I mostly care about ML applications, where I'm wondering if this is expected to help at all.