5 ms·
A lot of effort went into these efforts to understand neural networks in terms or kernels or SVMs. To the best of my knowledge, these efforts have not inspired
by erostrate 5y ago
A lot of effort went into these efforts to understand neural networks in terms or kernels or SVMs. To the best of my knowledge, these efforts have not inspired useful new architectures, nor have they made useful experimental predictions on real neural nets, nor have they had any significant impact on important machine learning benchmarks.
I think some researchers are refusing to accept the idea that machine learning is very much an experimental science today, and the (very cool) mathematics of kernels, SVMs, empirical risk minimization, bayesian statistics, etc. are simply no longer useful in the large scale regime.
My prediction is that there will be a useful "deep learning theory" in the future, but it will look a lot more like physics (such as Kaplan's scaling laws) than early 21st century machine learning mathematics/statistics.
- curiousgal 5y agoThis is ironic because the exact same thing could have been said about Neural Networks back in the 90s. Just because there aren't any practical applications now of "mathy ML", it does not mean it's a dead end.
- civilized 5y agoI remember how uncool neural nets were back then! Kernel SVM was so hot, but LeCun said it was nothing more than template matching with fancy math. And he was right...
- ThereIsNoWorry 5y agoNeural networks are very non-linear. The blessing and curse of non-linearity is its usefulness while escaping what mathematical logic inherently is able to model. So yes, I strongly assume there will never be a mathematically sound theory that predicts deep learning in any meaningful way. Same problem with everything non-linear. From two-body to chaotic systems, biological interactions to the stock market. All just wishful thinking until Mathematicians give up and instead continue to play around in their well-defined esoteric spaces.
- gmadsen 5y agoyou don't seem to have a background in mathematics. A huge(and growing) body of work exists for every example you gave. Obviously mathematicians are interested in nonlinear systems
- goatlover 5y agoSo you're saying Hari Seldon's psychohistory is BS?
- ravi-delia 5y agoI think you conflate several ideas of understanding here. We don't have good ways of predicting chaotic systems, but we often know a lot about them. Just knowing they're highly sensitive would be hard to confirm without math.
- bcaine 5y agoWhile I sort of agree that machine learning will end up as an experimental science, it's way, way too early to say whether the theory relating deep learning to kernel methods (e.g. Neural Tangent Kernels) will be useful or not. As an example, just last week a (huge) paper [1] was put on arXiv that used these theoretical methods to analyze a bunch of common architecture building blocks (skip connections, normalization, etc), and then applied their theoretical findings to figure out how to train Resnet like models in similar training time without these seemingly "required" building blocks. Deep Learning is still in its infancy in many ways, and this type of research takes time, slowly building on successive results. [1] https://arxiv.org/abs/2110.01765 https://arxiv.org/abs/2110.01765
- mellavora 5y agoWhen you wrote 'huge', I thought you meant huge potential impact; I wasn't expecting 172 pages. team behind the paper is Deepmind/Google. It is probably worth a read.
- ChrisLomont 5y ago>I think some researchers are refusing to accept the idea that machine learning is very much an experimental science today, and the (very cool) mathematics of kernels, SVMs, empirical risk minimization, bayesian statistics, etc. are simply no longer useful in the large scale regime. The same statement could have been said of neural networks for a decades, but researchers poking at corners eventually found methods to turn them into the useful tools they are now. After all, those methods you downplay were created since neural networks were not useful at that time, whereas many of these were. I'd not poo poo what researchers decide to poke at. Pretty much every breakthrough is people poking at the edges of understanding. If solutions or steps were straightforward, then it would be engineering, not research. >My prediction is that t Mine is that neural networks as we understand them now get replaced by much more solid methods, based on the principles from scientific machine learning, where sophisticated differentiable models that are designed to mimic the problem space get tuned. After all, even current neural networks are heading that direction. Neural networks are simply too simplistic to capture lots of the complexity that problems demand (hence the current move past them in many domains). Gluing linear functions together ad-hoc is simply a low level approximation to what can be developed using centuries of powerful mathematics to make models.
- mrtranscendence 5y ago> Mine is that neural networks as we understand them now get replaced by much more solid methods, based on the principles from scientific machine learning, where sophisticated differentiable models that are designed to mimic the problem space get tuned. Maybe so. But I suspect that scientific machine learning would be difficult to grasp by most working data scientists and ML users who don't have graduate-level training in a quantitative science or engineering, which implies an uphill battle for adoption. You can do a lot with ML and deep learning without to construct a sophisticated mathematical model of your problem space. (I took one class in differential equations in 2004, touched up on that in a boot camp before starting a PhD in econ, and have probably forgotten everything I ever learned about the topic ever since. Why yes, I feel mathematically inadequate sometimes.)
- ChrisLomont 5y ago
- 6gvONxR4sf7o 5y ago> I think some researchers are refusing to accept the idea that machine learning is very much an experimental science today… This is like criticizing the physicists working on the principles underlying steam engines during the days when people were making empirical advances in building steam engines, but nobody had figured out all the core principles yet. Of course they understand that it’s an empirical science today. That’s exactly why they’re doing the work they’re doing.
- civilized 5y ago> My prediction is that there will be a useful "deep learning theory" in the future, but it will look a lot more like physics (such as Kaplan's scaling laws) than early 21st century machine learning mathematics/statistics. I respectfully but completely disagree. All the enabling technologies of deep learning come from machine learning and statistical learning theory. Stochastic gradient descent, regularization, dimension reduction, bootstrap, bagging, boosting: these techniques remain the fundamental tools in the deep learning toolbox, and a constant source of inspiration for the latest innovations. Physics has done next to nothing for deep learning in comparison. It's as marginal as the kernel SVM stuff. Just fiddling around with the same old stat mech / network theory / power law stuff the Complex Systems types been doing for the last few decades. Stat mech works great for materials because materials are relatively simple things and we have relatively simple questions about them. We want to know how they respond to electricity, magnetism, heat, pressure, etc. When you turn to elaborate gadgets for machine translation or image recognition, sure, you can ask similar questions, but the questions just aren't as interesting. You'll get plenty of plots and histograms and power laws out of it, but it's all going to be very superficial. It's not going to tell you how the gadget works. All that said, there's no way to be sure where the big advances in deep learning theory are going to come from. Your guess might be as good as mine.