3 ms·
Their mathematical analysis is for sure beyond my understanding, but on some level, I'm surprised this is surprising. While baby steps may be the norm in gradi
by jonnycat 3y ago
Their mathematical analysis is for sure beyond my understanding, but on some level, I'm surprised this is surprising. While baby steps may be the norm in gradient descent, there are plenty of there optimization techniques that embrace larger step sizes. Some, like simulated annealing, explicitly model changes in step size over time. Others, like genetic algorithms with crossover let the balance between exploration (big steps) vs exploitation (small steps) emerge more organically.
In fact, applying the idea of "no free lunch in search and optimization" (https://en.wikipedia.org/wiki/No_free_lunch_in_search_and_optimization https://en.wikipedia.org/wiki/No_free_lunch_in_search_and_op...), you'd kind of expect this result: sometimes big steps good, sometimes big steps bad.
- marmakoide 3y agoThere are fairly old takes on that idea. For example, the evolutionary algorithm community knew the relevance of using steps size drawn from a Cauchy distribution, for example "Evolutionary Programming Made Faster", Xin Yao et al., 1999.