11 ms·
There isnt really any math to deep learning other than the concept of a derivative which is taught in high school calculus. The reason deep learning papers seem
by reader5000 9y ago
There isnt really any math to deep learning other than the concept of a derivative which is taught in high school calculus. The reason deep learning papers seem mathy is people take network architectures and various elementary operations on them and try to express them symbolically in latex using summations and indexing-hell. For example the easy concept of "updating all the neurons in one layer based on the neurons in the previous layer and connecting weights" is expressed as matrix-vector multiplication for not really any apparent reason other than it is technically correct and makes for slicker notation, and I guess makes it easier to use APIs that compute gradients for you. Deep learning however is broadly an experimental science, which in many ways is the opposite of math as traditionally envisioned, in which great insights follow deductively from prior great insights. If you ask a basic question like "why should use 4 layers instead of 3?" there is no answer other than "4 works better". Similarly with gradient descent versus random search in weight space. There are many problem domains where random search is as good as any known hill-climbing heuristic search (like gradient descent). Why is GD so effective when learning image classifiers expressed as stacked weight sums? Who knows.
- drewbuschhorn 9y agoAs someone who got as far as diffeq in college math, and is working his way through fast.ai right now, the impression I get is of a field that's at the start of formalization. It's like they've got the basic operations like addition and subtraction, but multiplication is still on the horizon. Or like the early days of calculus when some mathematicians called it black magic.
- lawless123 9y agoWell neural nets have been called a Dark Art in the past, though that seems to be changing now.
- davedx 9y agoUsing matrices to perform the calculations is an optimization over doing a bunch of for loops. This vectorization results in faster code within higher level languages and on certain hardware platforms (SIMD). It's nothing to do with "slicker notation", although having written gradient descent with for loops and matrix operations, the vectorized version is simpler and cleaner to read in my opinion.
- return0 9y agoThis has nothing to do with understanding backpropagation (which (correct me if i m wrong) is really the core of DL). In fact in the old days backpropagation was all about "propagating the deltas" and nothing about vectorizing.
- Houshalter 9y agoHe's not complaining about using vectorization in code. The problem is papers and even explanations targeted at non-experts, often use obfuscated math in place of clear explanations. I've complained about this before here: https://news.ycombinator.com/item?id=13953530 https://news.ycombinator.com/item?id=13953530 Mathematical notation is basically a programming language. A programming language with weird symbols you can't type to search for, single letter variable names for everything, and no comments. And it's written by programmers that are obsessed with fitting everything into a simple line and making it as small as possible, no matter how difficult it is to read. Any programmer understands this is incredibly bad practice. And even if parse every step and perfectly follow what the code is doing, without explanation, it's pretty difficult to figure out why.
- gnaritas 9y ago> Mathematical notation is basically a programming language. A very bad one that can only be executed by brains with the requisite existing historical knowledge; in fact it's more like bad pseudo-code that lacks the explicitness necessary to translate into actual instructions. It's basically condensed jargon intended for the already converted. It'd probably be vastly easier to teach math with an actual programming language than with traditional notation. Scheme would be ideal for this.
- davedx 9y agoOK, I see what you're saying. I think you have the same issue with "real" programming languages too. If you compare some very concise Clojure or Scala code with the equivalent in Java, it can be quite hard to understand if you're not very familiar with the language. But I wouldn't necessarily say it's "incredibly bad practice". A Scala programmer can write concise and elegant code that to another Scala programmer is actually faster to understand because of that conciseness. Whereas the same code written with for loops and class method calls and all the boilerplate in Java would take more studying to filter out the low level constructions. It's about the level of abstraction. And yeah if you don't understand the notation or syntax at the level of abstraction you're studying, it will be very hard. (FWIW I find Scala code quite hard to understand sometimes, but I also find the more I know about the language, the more comprehensible it gets).
- nahumfarchi 9y agoThat's interesting, I always thought that the graphical explanation was elaborate and confusing since a NN is just a bunch of matrices with non-linearities inbetween. To each his own I guess. I definitely agree though that it's more of an experimental science at the moment.
- return0 9y agoOTOH thinking in matrices only may be limiting, as there are potential designs one might want to try (brain-inspired for example) that cant be expressed in matrix operations.
- _nx010_ 9y agoThey use matrices for computational efficiency. That's why linear algebra (along with diff equations and probability theory) is one of the prerequisites for any non-mooc machine learning course.
- Houshalter 9y agoI've programmed neural networks without knowing any linear algebra. When I needed to figure out how to use vector operations for speed, it took like 5 minutes to search for matrix multiplication on wikipedia. You can't get by without even knowing that, as element wise operations can do everything just as fast.
- rs86 9y agoThere is a clear theoretical reason for using 4 layers vs 3. It allows for more degrees of freedom which translates to a higher VC dimension. This implies numerous trade-offs in model behavior. Besides this point, there is much more than simple derivatives in deep learning. For example regularization can yield quadratic programming problems. Different optimization algorithms can have tremendous impact on training time and model performance. Models can be quite sensitive to specific parameters that you can't just set at random. More ingenious architectures like GAN also require some fairly technical thinking to get right. There is much more than image classification and vanilla NN or CNNs.
- Houshalter 9y ago>There is a clear theoretical reason for using 4 layers vs 3. It allows for more degrees of freedom which translates to a higher VC dimension. But then why does using 5 layers work worse than 4? Your theory is no good at predicting what the hyperparameters should be. The only way to find the correct hyperparameters is through empirical search. >there is much more than simple derivatives in deep learning. For example regularization can yield quadratic programming problems. Different optimization algorithms can have tremendous impact on training time and model performance. All these concepts are fairly simple also and can be expressed with little math. Additionally, a casual user doesn't need to have a deep understanding of them and the library will usually take care of it. Any more than a programmer needs to have a deep understanding of how an optimizing compiler works. >More ingenious architectures like GAN also require some fairly technical thinking to get right. The idea of using NNs to trick each other, is also fairly simple. It doesn't even involve any math.
- return0 9y agoindeed the most unfamiliar/off-putting part might be the matrix formulation of the designs (and the matching of dimensions), which is not even useful if you are trying to implement a toy example in programming. But you can equally well understand backpropagation by following the updating of a single weight, which is much more intuitive. The other thing is the unfortunate/misleading/atrocious jargon that has been adopted.
- daddyo 9y agoIf we could calculate the perfect model architecture for all unseen data, without relying on evaluation/experimentation/heuristics, then we'd have effectively solved the halting problem. Mathematically, closest to that would be Hilbert's program. Though neural nets can paint like Van Gogh nowadays, asking them to come up with Hilbert's program may be a bit too much of an ask. Yet I would not deeply mind if researchers would revisit papers like http://www.ics.uci.edu/~rickl/publications/1996-icml.pdf http://www.ics.uci.edu/~rickl/publications/1996-icml.pdf "On the Learnability of the Uncomputable".