6 ms·
Gradient descent type optimization is far from "trying random things until one of them works without really knowing why". You can calculate all partial derivati
by boyka 3y ago
Gradient descent type optimization is far from "trying random things until one of them works without really knowing why". You can calculate all partial derivatives and understand the impact.
- deleted 3y ago[deleted]
- varispeed 3y agoSure you can, but you can also throw things in randomly and see what works then build theory around it.
- quickthrower2 3y agoI think it is a freaking miracle it even works. I understand how it works for say 100 parameter linear regression, but that it would work for billions of parameters (billion dimensional space) by nudging the parameters each a little (based on purely it's impact on the loss, assuming everything stays the same), is not obvious to me. It is a kind of magic. Regarding randomness, the initialization of the weights is random, and if they use dropouts that is random too, plus the order in which to process the text might be random.
- rmnclmnt 3y agoWhat’s even crazier is the fact that the underlying numerical math solution was exposed early 1800’s by Gauss!
- bravura 3y agoTo be honest, there’s always been a tension in machine learning between the: “we won’t do it unless the theory is complete and sound” versus the “we don’t understand this but it works much better consistently so we do it” crowd. In the 90s and even 00s, the theory first crowd was mainstream and the empirical first crowd was considered fringe. Very fringe. Personally I appreciated it when LeCun was like: “you can’t find the solution if you only search where the lamplight is shining.” Or other early deep learning practitioners note that ML theory is usually so far disconnected from practice in terms of tightness and bounds that you might as well ignore pure theory completely. Anyway, it wasn’t until deep learning methods really smashed benchmarks across the board did people give in to the black magic / alchemy driven approaches of empiricism based upon intuition and bias developed through long-held experience.
- deleted 3y ago[deleted]
- creata 3y agoTo a layperson like me, the "try random things" part of machine learning doesn't seem to lie in optimizing the parameters, but in designing the model.