4 ms·
OP doesn't weight the size of each technology by adoption -- if they did that then yeah there are some centroids. But I think their point is correct that parts
by kajecounterhack 5y ago
OP doesn't weight the size of each technology by adoption -- if they did that then yeah there are some centroids. But I think their point is correct that parts of the ecosystem remain quite fragmented.
> For example, when I first got into ML the ADAM optimizer was the new big thing, since then hundreds of 'better' optimizers have been published. This paper from August '21 shows that most of that is overblown and no optimizer consistently outperforms ADAM: https://arxiv.org/pdf/2007.01547.pdf https://arxiv.org/pdf/2007.01547.pdf
FWIW this paper acknowledges that optimizer is application-dependent, yet draws its conclusions from < 10 canned datasets (mostly classification). "no optimizer consistently outperforms ADAM" is a misleading statement -- I think a more correct one is "no optimizer consistently outperforms ADAM on {MNIST, CIFAR-10, SVHN}". If you do work on much larger and more novel architectures / datasets, I think you'll find the advice to "just use adam" to be insufficient. At least at the orgs I've worked at, folks often do a (sometimes smart) param sweep to find best optimizer / settings << if you're lucky enough to have infra that supports this :)
That aside and more generally, there certainly are commonly accepted best practices in ML. But that doesn't mean ML organizations all work the same way, and the OP's point is that the dust has not settled on what the best tools are to do X.