4 ms·
This is something I am interested and considering as a possible future for my career, coming from a math background and starting to dip my toes into ML. One thi
by orange_tee 6y ago
This is something I am interested and considering as a possible future for my career, coming from a math background and starting to dip my toes into ML.
One thing I am afraid of is how useful is it actually going to be to create a rigorous mathematical framework?
I am afraid it might end up like mathematical physics, where they are almost a century behind theoretical and experimental physics and playing catch-up.
What does anyone else think?
- bigdict 6y ago> This is something I am interested and considering as a possible future for my career, coming from a math background and starting to dip my toes into ML. I'm in the exact same position. I think mathematical explanations and (substantiated) intuition for the things that practitioners discover would be very useful. Maybe an all explaining grand theory of deep learning is possible.
- orange3xchicken 6y agoimo, instead of developing new theory entirely from the ground up, its more useful to address & work in the context of understanding longstanding open problems/phenomenon. Theoretical insight and the framework should follow. e.g. Robustness & expressiveness Memorization & catastrophic forgetting Ensembling & ntk & optimization Double descent Pruning
- dr_dshiv 6y agoRelate it to free energy minimization, that's very hot right now. Read Smolensky's 1986 article in PDP on the harmonium, it was the first restricted boltzmann machine.
- canjobear 6y agoDon't most neural networks already do free energy minimization? Just about any information theoretic objective can be interpreted that way...
- dr_dshiv 6y agoYes, that's arguably the dominant paradigm. It's interesting because of the relationship to thermodynamics -- again, check out Smolensky's paper in PDP, he was a postdoc (along with Geoff Hinton) at UCSD's cognitive science dept, run by Don Norman. https://apps.dtic.mil/sti/pdfs/ADA620727.pdf https://apps.dtic.mil/sti/pdfs/ADA620727.pdf
- eli_gottlieb 6y ago>Relate it to free energy minimization, that's very hot right now. It is? I didn't see that buzzword very much at NeurIPS this year.
- GregarianChild 6y agoIn what sense are neural nets not rigorous? The whole pipeline from Pytorch or TensorFlow and Python to LLVM, to GPUs or TPUs is absolutely rigorous. Much more rigorous, in fact, than normal, hand-written mathematics, as you find it in e.g. a typical Annals of Mathematics publication, or mathematical textbook! I think what you really have in mind is a simple model of modern deep learning that is not fully accurate, but still useful! Let me argue by analogy. You are looking for something that is to deep learning what the lambda-calculus is to the Haskell compiler. One of the main simplifications in programming language theory is replacing finite precision arithmetic (which is painfully complex) with mathematical integers and real numbers (which are much simpler). Would a theory of deep learning based on mathematical reals be valuable in a theory of deep learning? The stunning success of floating point formats like bfloat16 [1] suggests otherwise, since arithmetic precision in deep learning is closely connected to important learning phenomena such as overfitting and regularisation. I am tempted to be provocative and say that you are really looking for less rigour! [1] https://en.wikipedia.org/wiki/Bfloat16_floating-point_format https://en.wikipedia.org/wiki/Bfloat16_floating-point_format
- avrionov 6y agoNot the OP, but I've built several ML pipelines in the last few years. Even if every step in the process very rigorous and using solid software, there are still challenges. ML models are very difficult (or impossible to test). The usual testing approaches don't work at all for ML features.
- b3kart 6y agoWe don't know which model architecture will work for which problem, and why it would/wouldn't. We experiment until we find something that works, and can sometimes try to guess why it did. But none of this knowledge is formalized in a way that can reliably predict performance in future problems. We engineer solutions to problems, but don't build a rigorous body of knowledge to help us in future problems.
- glial 6y agoWell put. I would say that often even the problems aren’t particularly well defined.
- konjin 6y agoYou're 5 years too late to try and break into ML.
- enriquto 6y ago> This is something I am interested and considering as a possible future for my career, coming from a math background and starting to dip my toes into ML. This is a great time and place to be! Neural networks are the twenty-first century Fourier series. It's just that we don't yet understand them. We can easily run them (synthesis) but we are missing the analysis. There's a lot of math to do here.
- joe_the_user 6y agoOne thing to consider is the distinction between deep neural networks as mathematical objects and machine learning as currently practiced. Lately, there have been quite a few theories on neural networks as ideal nonlinear approximators [1]. Similarly, people have shown many ways that gradient descent can tend to reach a global maximum of regularized curve-closeness[2]. Which is to say, if your development cycle is: gather-data, train, test, deploy, we know this approximates the data almost ideally; you can't really do much better than a deep network. But we know in practice, when deployed, that deep neural networks actually have many limitations (compared to our intuitions or compared to human performance, etc). There are some obvious explanations. Of course, they're limited by our ability to gather data and by the biases of the data. But even more, they're limited to situations where you have large chunks of unchanging data that you can extrapolate from. Given that deeps are more or less perfect for the train-test-deploy cycle, it seems like the problem is with this cycle itself. And it's easy to see human beings somehow acting "intelligently" without using this cycle. So figuring out an alternative to this might be something to look at. [1] For example: Nonlinear Approximation and (Deep) ReLU Networks I. Daubechies, et al https://arxiv.org/abs/1905.02199 https://arxiv.org/abs/1905.02199 [2] For example: Gradient descent optimizes over-parameterized deep ReLU networks Difan Zou et al https://link.springer.com/article/10.1007/s10994-019-05839-6 https://link.springer.com/article/10.1007/s10994-019-05839-6