10 ms·
The Modern Mathematics of Deep Learning
- amelius 5y agoWhat are the prerequisites?
- thanksok 5y agoLooks like a little bit of everything except the likes of abstract algebra, logic, category theory. These include linear algebra, graph theory, probability, algorithms, mathematical analysis, topology, differential geometry. But the most important prereqs are math maturity and mental toughness/endurance.
- SilurianWenlock 5y agomental toughness/endurance haha!
- keithalewis 5y agoMind reading. They use terminology without defining it or giving a reference.
- beforeolives 5y agoSeriously, I'm struggling to understand things that I already know.
- sundarurfriend 5y agoAny examples? I haven't yet come across something like that yet, but I'm only a short way into the article.
- keithalewis 5y agoThe terms "measurable" and "tempered" for starters.
- ganzuul 5y agoFor the latter maybe this? https://en.wikipedia.org/wiki/Parallel_tempering https://en.wikipedia.org/wiki/Parallel_tempering
- ycreader 5y agoIs parallel tempering related to https://en.wikipedia.org/wiki/Bennett_acceptance_ratio https://en.wikipedia.org/wiki/Bennett_acceptance_ratio ?
- keithalewis 5y agoNyet, and nyet. This is why conscientious authors define the terms they use. A tempered distribution is a linear functional on a space of differentiable functions, for example, D_x(f) = f'(x), the derivative of f at x. This is why tempered distributions cannot be composed. In general, the dual of a space of functions is a space of set functions, aka measures. https://keithalewis.github.io/math/dual.html https://keithalewis.github.io/math/dual.html
- cpp_frog 5y agoThe term measurable is referring to "measurable functions" in measure theory, which correspond to functions verifying that the pre-image of any measurable set belonging to the sigma-algebra of the codomain belongs to the sigma-algebra of the domain (https://en.wikipedia.org/wiki/Measurable_function https://en.wikipedia.org/wiki/Measurable_function). I do not know how to state it in simpler terms, sorry. When the measure of the domain is 1 (as in a probability space), we call measurable functions random variables, hence their relevance to this topic. Now, tempered distributions are functions that assign a complex number to a very rapidly decaying function (a Schwarz space function), and it satisfies linearity properties. So this is a function that takes functions and maps them to complex numbers. https://secure.math.ubc.ca/~feldman/m321/distributions.pdf https://secure.math.ubc.ca/~feldman/m321/distributions.pdf
- keithalewis 5y ago
- godelski 5y agoI skimmed it. Looks like just some basic calc and linear algebra. Nothing that crazy.
- fspeech 5y agoMostly analysis. If you understand section 1 notations, you are obviously set. But even if you don't you should still be able to get the ideas with a bit of mental translation. In a word the notation seemed unnecessarily heavy for the level of discussion.
- 0-_-0 5y agoDeep learning papers often use math in a way that obscures rather than enlightens. And when you finally understand what they are saying, you realize it's not interesting at all, or they made a mistake in the math.
- cpp_frog 5y agoWhile I can't give the exact prerequisites, I know that all of the things that appear in the paper relate to: (1) Linear Algebra (2) Optimization Theory (Convex Analysis, non-convex optimization) [0], [2] (3) Probability Theory and Statistics (Measure Theory, Multivariate Statistics) [1], [3], [4], [5] (4) Analysis, to a lesser extent. (2) and (3) are the most important. I would give more references, but my background is too theoretical (and my field is Numerical Analysis of PDE). From the classes I took in college, three or four on each of (1-4), a person with a similar background can recognize the tools without much digging. Maybe some folks here can provide some insights into books that center on applications. So I'm trying not to diverge into too much theory (i.e. for measures, [4] instead of Folland). There also seems to be good use of Analysis techniques in the paper, see theorem 2.1. I love that the paper references the Moore-Penrose pseudo-inverse, an object of study in both statistics and optimization for which I had to give a lecture for a course. [0] https://web.stanford.edu/~boyd/cvxbook/bv_cvxbook.pdf https://web.stanford.edu/~boyd/cvxbook/bv_cvxbook.pdf Convex Optimization, Boyd and Vandenberghe [1] An Introduction to Multivariate Statistical Analysis, Anderson [2] Convex Analysis and Monotone Operator Theory in Hilbert Spaces, Bauschke-Combettes [3] Theory of Multivariate Statistics, Bilodeau-Brenner [4] The Elements of Integration and Lebesgue Measure, Bartle [5] Probability: Theory and Examples, Durrett
- julbern 5y agoI would recommend a solid background in linear algebra, probability theory, and analysis. Moreover, for some sections, it is helpful to have experience with functional analysis, optimization, and statistical learning theory. Some helpful resources are linked here: https://www.reddit.com/r/MachineLearning/comments/najnjg/r_the_modern_mathematics_of_deep_learning/gyhjuec/ https://www.reddit.com/r/MachineLearning/comments/najnjg/r_t...
- joe_the_user 5y agoIt seems like this can leave the reader with the wrong impression. Calculus really is "the mathematics of Newtonian physics". This is just "some mathematics that might help a bit in your intuitions of deep learning". IE, Deep learning is fundamentally just about getting the mathematically simple but complex and multi-layerd "neural networks" to do stuff. Training them, testing them and deploying them. There are many intuitions about these things but there's no complete theory - some intuitions involve mathematical analogies and simplifications while other involve "folk knowledge" or large scale experiments. And that's not saying folks giving math about deep learning aren't proving real things. It's just they characterizing the whole or even a substantial part of such systems. It's not surprising that a complex like a many-layered Relu network can't fully characterized or solved mathematically. You'd expect that of any arbitrarily complex algorithmic construct. Differential equations of many variables and arbitrary functions also can't have their solutions fully characterized.
- conformist 5y agoIt seems like it aims at giving somebody who would like to get started doing theoretical research in the field some pointers and basic insights. I don't think it does a particularly bad job at this, in particular given that it will be a book chapter? The target audience are probably people who have had some exposure to Functional Analysis and the likes before.
- jhrmnn 5y agoThere are a few works that try to put deep learning on some theoretical basis, I like this one, for example: https://arxiv.org/abs/1703.00810 https://arxiv.org/abs/1703.00810 This goes beyond mere intuition, but it is also still very far from a “complete theory”. I find it disappointing that so few people in deep learning work on the theoretical foundations.
- quibono 5y agoWhat are some subfields of mathematics that you would say are crucial for gaining a proper understanding of all the things related to deep learning (e.g. let's say the paper you linked)? Even though the theory isn't complete, I'm sure a grounding in certain fields of mathematics will be helpful.
- pcbro141 5y agoTangent, but has anyone taken Fast.ai or similar courses and transitioned into the Deep Learning/ML field without a MS/PhD? To be honest, I don't even know what 'doing ML/DL' looks like in practice, but I'm just curious if a lot of folks get in to the field without graduate degrees.
- mustafa_pasi 5y agoYou can learn all you need to know in 2 to 3 university level courses. So we are talking less than a year of university courses. Fast.ai is too high level. I don't like it. You would be better served taking actual university courses. A few days ago people linked to LeCun's university class[1]. This is a solid introduction. Does not cover everything but that is OK. Seems like it is missing Bayesian approaches. Then if you want to specialize in vision or speech or robotics or whatever, you take special classes on that topic and learn all the SOTA techniques. Then you are ready to do research already, or apply your knowledge to build stuff. Of course you still have to learn how to do real machine learning, which involves all the data manipulation stuff, but that is learned by doing. [1] https://cds.nyu.edu/deep-learning/ https://cds.nyu.edu/deep-learning/
- tracyhenry 5y agoAnother one I really liked is Berkeley CS182: https://cs182sp21.github.io/ https://cs182sp21.github.io/ The youtube playlist is here: https://www.youtube.com/playlist?list=PL_iWQOsE6TfVmKkQHucjPAoRtIJYt8a5A https://www.youtube.com/playlist?list=PL_iWQOsE6TfVmKkQHucjP... Prof. Sergey Levine is REALLY good at explaining the intuitions of DL algorithms. This class also includes lectures on ML basics and very approachable assignments. Many classes/blog posts start with describing what a neuron is - that IMHO is a super terrible way to teach a beginner. To understand DL, one should know why we need activations (because linear models are not enough), why we need back-propagation (because we are optimizing a loss using SGD). This class is very great at explaining those things in an intuitive way. Following through I felt I built a pretty solid ML/DL foundation for myself.
- akgoel 5y agoI am in a Fintech boot camp, and it’s clear that doing ML/DL requires very little math, as the math is all abstracted away.
- rohittidke 5y agoI believe that the curse of dimensionility doesn't apply here as we are optimizing the "universal apppriximator" of the "surface" of the possible real world function.
- antipaul 5y agoDoes “possible” in your statement refer to the inherent constraints of the architecture as specified by the researcher, or something else?
- ganzuul 5y ago> Kernel methods owe their name to the use of kernel functions, which enable them to operate in a high-dimensional, implicit feature space without ever computing the coordinates of the data in that space, but rather by simply computing the inner products between the images of all pairs of data in the feature space. - https://en.wikipedia.org/wiki/Kernel_method https://en.wikipedia.org/wiki/Kernel_method As it relates to this: https://en.wikipedia.org/wiki/Neural_tangent_kernel https://en.wikipedia.org/wiki/Neural_tangent_kernel To me, this is JFM. Not sure if I'm connecting the dots right either. I just don't know of anything else claiming to solve the curse.
- somewhereoutth 5y agoWake me up when 'deep learning' has independently created a language to communicate within a group of peers while under environmental pressure. (and that language is co-expressive with human languages)
- visarga 5y agoSelf play applied to language creation as opposed to go and chess? I like this idea.
- scaraffe 5y ago> a language to communicate within a group of peers while under environmental pressure what does this mean?
- scaraffe 5y ago> a language to communicate within a group of peers while under environmental pressure what does this mean?
- deleted 5y ago[deleted]
- mathgenius 5y agoAfter skimming through the paper it's clear that the title should be read as "The Modern (Mathematics of Deep Learning)" and not my original parse which was "The (Modern Mathematics) of Deep Learning." Very different things.