13 ms·
Each time I teach neural nets to an engineer, there's only a 50% chance they can write down the chain rule. Colah's blog on backprop used to be my favorite reso
by minihat 5y ago
Each time I teach neural nets to an engineer, there's only a 50% chance they can write down the chain rule. Colah's blog on backprop used to be my favorite resource to leave them with (https://colah.github.io/posts/2015-08-Backprop https://colah.github.io/posts/2015-08-Backprop).
The explanation of the calculus in this tool is equally fantastic. And the art is very cute.
There are many ways to skin a cat, of course, but this is as good a tutorial as I've seen for getting you through backprop as fast as possible.
- jhgb 5y ago> there's only a 50% chance they can write down the chain rule I blame the common mathematical notation for that.
- ThinkingAgain 5y agoAfter reading the tutorial I was not sure why it was called backpropagation. Thanks for the Colah's blog link. I think the two links together explains the things beautifully. Backpropagation seems just like an optimization for the gradient descent calculation as per Colah's blog.
- shaan7 5y agoAny recommendations for a 101 book for neural nets for someone who is "just a programmer"? OP's tutorial is quite nice, but I love to read books and find it easier to learn from them.
- jacobcmarshall 5y agoDeep Learning with Python by Chollet is an excellent beginner resource if you are a hands-on learner. It starts off with some tutorials using the Keras library, and then gets into the math later on. By the end of the book, you create multiple different types of neural networks for identifying images, text, and more! I highly recommend it.
- shaan7 5y agoThanks for the recommendations folks! :)
- wesleywt 5y agoFastai has a course: practical deep learning for programmers.
- carom 5y agoThere are also coursera specializations from Andrew Ng at https://deeplearning.ai https://deeplearning.ai.
- matsemann 5y agoNg's course is bottom up: Start with the basic math, expand upon it, until you arrive at ML and neural nets. Fastai is top down: learn to use practical ML with abstractions, and then dig deeper and explain as needed. I preferred fastai's approach, even though I enjoyed both. Ng's could be a bit too low level and fundamental for what I wanted to learn.
- carom 5y agoThis is a valuable take. Fastai was very frustrating for me because I wanted to understand the internals. I ended up not finishing it, so take my opinion with a grain of salt.
- windsignaling 5y agoI much prefer Andrew Ng's courses as well. I tried Fast AI, but it seems to be trying too hard to take out the math, which oddly for me (as a STEM grad) makes it much more difficult to understand. Had to stop when I saw him using Excel spreadsheets to explain convolution.
- knicholes 5y agoThe Excel spreadsheets is where it all clicked for me. It's not done with some magical library but with simple math being performed on some weights.
- carom 5y agoThere is a book called neural networks from scratch at https://nnfs.io https://nnfs.io.
- aeg42x 5y agoI highly recommend http://neuralnetworksanddeeplearning.com/ http://neuralnetworksanddeeplearning.com/ it’s an online book that has some great code examples built in.
- bernulli 5y agoI also found his ‘visual proof’ for neural networks as general function approximators super intuitive.
- baron_harkonnen 5y agoGiven the current state of automatic differentiation I'm not so sure it's even necessary or particularly useful to focus on backpropagation any more. While backprop has major historic significance, in the end it's essentially just a pure calculation which no longer needs to be done by hand. Don't get me wrong, I still believe that understanding the gradient is hugely important, and conceptually it will always be essential to understand that one is optimizing a neural network by taking the derivative of the loss function, but backprop is not necessary nor is it particularly useful for modern neural networks (nobody is computing gradients by hand for transformers). IMHO a better approach is to focus on a tool like JAX where taking a derivative is abstracted away cleanly enough, but at the same time you remain fully aware of all the calculus that is being done. Especially for programmers, it's better to look at Neural Networks as just a specific application of Differentiable Programing. This makes them both easier to understand and also enables the learner to open a much broader class of problems they can solve with the same tools.
- medo-bear 5y agoBackpropagation is a particular implementation of reverse mode auto-differentiation, and it is the basis for all implementaions of DL models. It is very strange for me to read this as though it is very obvious and commonly accepted fact, which I don't think it is.
- baron_harkonnen 5y ago> to read this as though it is very obvious and commonly accepted fact I'm not entirely sure what you're referring to by "this" but assuming you mean my comment, I think what I'm saying is very much up for debate and not an "obvious and commonly accepted fact". Karpathy has a very reasonably argument that directly disagrees with what I'm suggesting [0]. Of course he also agrees that in practice nobody will every use backprop directly. Whether it's JAX, TF, PyTorch, etc the chain rule will be applied for you. I'm arguing that I think it's helpful to not have to worry about the details of how your derivative is being computed, and rather build an intuition about using derivatives as an abstraction. To be fair I think Karpathy is correct for people who are going to be learning to explicitly be experts in Neural Networks. My point is more that given how powerful our tools today are for computing derivatives (I think JAX/Autograd have improved since Karpathy wrote that article), it's better to teach programmers to learn think of derivatives, gradients, hessians etc as high level abstractions. Worrying less about how to compute them and more about how to use them. In this way thinking about modeling doesn't need to be restricted to strictly NNs, but rather use NNs and example and then demonstrate to the student that they are free to build any model by defining how the model predicts, scoring the prediction and using the tools of calculus to answer other common questions you might have. edit: a good analogy is logic programming and backtracking/unification. The entire point of logic programming is to abstract away backtracking. Sure experts in Prolog do need to understand backtracking, but it's more helpful to get beginners understanding how Prolog behaves than understand the details of backtracking. [0] https://karpathy.medium.com/yes-you-should-understand-backprop-e2f06eab496b https://karpathy.medium.com/yes-you-should-understand-backpr...
- matsemann 5y ago> there's only a 50% chance they can write down the chain rule Why should I, though? I remember the concept from calculus. I know pytorch keeps track of the various stuff I do to a vector and calculates a gradient based on it. What more do I need to know when all I want to do is to play with applications, not implement backprop myself?
- medo-bear 5y agoif you don't understand chain rule then you dont understand backprop, which means you do not really understand how deep learning works. at most you can follow recipes cook book style. it is kind of how one can make a website without a deep understanding of networking
- baron_harkonnen 5y ago> at most you can follow recipes cook book style. Here I disagree with you pretty strongly. Once someone is comfortable with differentiable programming it's much more obvious how to build and optimize any type of model. People should be more concerned about when to use derivatives, gradients, hessians, Laplace approximation etc rather than worry about the implementation details of these tools. Abstraction can also aid depth of understanding. I know plenty of people who can implement backprop, but then don't understand how to estimate parameter uncertainty from the Hessian. The latter is much more important for general model building.
- medo-bear 5y agoi am not sure what you are disagreeing with. chain rule is basic calculus that precedes understanding hessians. my argument is, if you can not understand what the chain rule is, you will not understand more complicated mathematics in ML. do you think i am wrong ? EDIT: also uncertainty estimation is the stuff of probabalistic approach to ML. i would say that people who do probabalistic ML are quite mathematically capable (at least to my experience)
- 5y ago
- friebetill 5y agoI found this 13 min explanation very helpful in understanding backpropagation (https://youtu.be/c36lUUr864M?t=2520 https://youtu.be/c36lUUr864M?t=2520). First he explains the necessary concepts: 1) Chain Rule 2) Computational Graph Then he explains backpropagation in these three steps (first in general and then with examples): 1) Forward pass: Compute loss 2) Compute local gradients 3) Backward pass: Compute dLoss/dWeights using the Chain Rule