9 ms·
Backpropagation is a leaky abstraction (2016)
- deleted 11mo ago[deleted]
- joshdavham 11mo agoGiven that we're now in the year 2025 and AI has become ubiquitous, I'd be curious to estimate what percentage of developers now actually understand backprop. It's a bit snarky of me, but whenever I see some web developer or product person with a strong opinion about AI and its future, I like to ask "but can you at least tell me how gradient descent works?" I'd like to see a future where more developers have a basic understanding of ML even if they never go on to do much of it. I think we would all benefit from being a bit more ML-literate.
- kojoru 11mo agoI'm wondering: how can understanding gradient descent help in building AI systems on top of LLMs? To mee it feels like the skills of building "AI" are almost orthogonal to skills of building on top of "AI"
- joshdavham 11mo agoI take your point in that they are mostly orthogonal in practice, but with that being said, I think understanding how these AI's were created is still helpful. For example, I believe that if we were to ask the average developer about why LLM's behave randomly, they would not be able to answer. This to me exposes a fundamental hole in their knowledge of AI. Obviously one shouldn't feel bad about not knowing the answer, but I think we'd benefit from understanding the basic mathematical and statistical underpinnings on these things.
- Al-Khwarizmi 11mo agoYou can still understand that quite well without understanding backprop, though. All you need is: - Basic understanding of how a Markov chain can generate text (generating each word using corpus statistics on the previous few words). - Understanding that you can then replace the Markov chain with a neural model which gives you more context length and more flexibility (words are now in a continuous space so you don't need to find literally the same words, you can exploit synonyms, similarity, etc., plus massive training data also helps). - Finally, you add the instruction tuning (among all the plausible continuations the model could choose, teach it to prefer the ones human prefer - e.g. answering a question rather than continuing with a list of similar questions. You give the model cookies or slaps so it learns to prefer the answers humans prefer). - But the core is still like in the Markov chain (generating each word using corpus statistics on the previous words). I often give dissemination talks on LLMs to the general public and I have the feeling that with this mental model, you can basically know everything a lay user needs to know about how they work (you can explain things like hallucinations, stochastic nature, relevance of training data, relevance of instruction tuning, dispelling myths like "they always choose the most likely word", etc.) without any calculus at all; although of course this is subjective and maybe some people will think that explaining it in this way is heresy.
- HarHarVeryFunny 11mo agoSure, but it'd be similar to being a software developer and not understanding roughly what a compiler does. In a world full of neural network based technology, it'd be a bit lame for a technologist not to at least have a rudimentary understanding of how it works. Nowadays, fine tuning LLMs is becoming quite mainstream, so even if you are not training neural nets of any kind from scratch, if you don't understand how gradients are used in the training (& fine tuning) process, then that is going to limit your ability to fully work with the technology.
- lock1 11mo ago> I'd like to see a future where more developers have a basic understanding of ML even if they never go on to do much of it. I think we would all benefit from being a bit more ML-literate. Why "ML-literate" specifically? Also, there are some people against calculus and statistic in CS curriculum because it's not "useful" or "practical", why does ML get special treatment here? Plus, I don't think a "gotcha" question like "what is gradient descent" will give you a good signal about someone if it get popularized. It probably will lead to the present-day OOP cargo cult, where everyone just memorizes whatever their lecturer/bootcamp/etc and repeats it to you without actually understanding what it does, why it's the preferred method over other strategies, etc.
- joshdavham 11mo ago> Why "ML-literate" specifically? We could also say AI-literate too, I suppose. I guess I just like to focus on ML generally because 1) most modern AI is possible only due to ML and 2) it’s more narrow and emphasizes the low level of how AI works.
- confirmmesenpai 11mo agoso if you want to have a strong opinion on electric cars you need to be able to explain how an electric engine works right?
- oceanplexian 11mo agoI’d say so, the hallmark of being a car guy is understanding the basics, like the difference between a four cylinder and a six cylinder, a turbocharger from a supercharger, the different types of gearboxes (DCT vs AT or a CVT), and so on. They all affect the feel, capabilities, and limitations of the car. Electric cars have similar complexities and limitations, for example the Bolt I owned could only go ~92MPH due to limitations in the gearing as a result of having a 1 speed gearbox. I would expect someone with a strong opinion of a car to know something as simple as the top speed.
- chermi 11mo agoDepends on what exactly the opinion is, but generally I'd say yes. If it's about their looks, maybe not.. but even then understanding the basics that determine things like not needing air intake, exhaust, the placement of batteries, etc. can be helpful. If it's about the supply chain, understanding at least the requirements for magnets is helpful. On way to make sure you understand all of these things is to understand the electric motor. But you could learn the separate pieces of knowledge on the fly too. The more you understand the fundamentals of what you're talking about, the more likely you are to have genuine insight because you can connect more aspects of the problem and understand more of the "why". TL;DR it depends, but it almost always helps.
- esafak 11mo agoThe electric engine? I won't give someone the time of day if they don't understand the battery chemistry.
- augment_me 11mo agoImpossible requirement. The inherent quality of abstractions is to allow us to get more done without understanding everything. We dont write raw assembly for the same reason, you dont make fire by rubbing sticks, you dont go hunting for food in the woods, etc. There is no need for the knowledge that you propose in a world where this is solved, you will achieve more goals by utilizing higher-level tools.
- joshdavham 11mo agoI get your point and this certainly applies to most modern computing where each new layer of abstraction becomes so solid and reliable that devs can usually afford to just build on top of it without worrying about how it works. I don’t believe this applies to modern AI/ML however. Knowing the chain rule, gradient descent and basic statistics IMO is not the same level of solid as other abstractions in computing. We can’t afford to not know these things. (At least not yet!)
- esafak 11mo agoAs ML becomes 'democratized' the fraction of people with technical understanding will trend towards zero. They will operate at a higher level of abstraction, and worry about novel things like hallucinations, prompt injection, and misinformation. Such is the nature of advancement.
- gchadwick 11mo agoKarpathy's contribution to teaching around deep learning is just immense. He's got a mountain of fantastic material from short articles like this, longer writing like https://karpathy.github.io/2015/05/21/rnn-effectiveness/ https://karpathy.github.io/2015/05/21/rnn-effectiveness/ (on recurrent neural networks) and all of the stuff on YouTube. Plus his GitHub. The recently released nanochat https://github.com/karpathy/nanochat https://github.com/karpathy/nanochat is fantastic. Having minimal, understandable and complete examples like that is invaluable for anyone who really wants to understand this stuff.
- throwaway290 11mo agoAnd to all the LLM heads here, this is his work process: > Yesterday I was browsing for a Deep Q Learning implementation in TensorFlow (to see how others deal with computing the numpy equivalent of Q[:, a], where a is an integer vector — turns out this trivial operation is not supported in TF). Anyway, I searched “dqn tensorflow”, clicked the first link, and found the core code. Here is an excerpt: Notice how it's "browse" and "search" not just "I asked chatgpt". Notice how it made him notice a bug
- stingraycharles 11mo agoFirst of all, this is not a competition between “are LLMs better than search”. Secondly, the article is from 2016, ChatGPT didn’t exist back then
- code51 11mo agoI doubt he's letting LLM creep in to his decision-making in 2025, aside from fun side projects (vibes). We don't ever come across Karpathy going to an LLM or expressing that an LLM helped in any of his Youtube videos about building LLMs. He's just test driving LLMs, nothing more. Nobody's asking this core question in podcasts. "How much and how exactly are you using LLMs in your daily flow?" I'm guessing it's like actors not wanting to watch their own movies.
- 11mo ago
- drivebyhooting 11mo agoI have a naive question about backprop and optimizers. I understand how SGD is just taking a step proportional to the gradient and how backprop computes the partial derivative of the loss function with respect to each model weight. But with more advanced optimizers the gradient is not really used directly. It gets per weight normalization, fudged with momentum, clipped, etc. So really, how important is computing the exact gradient using calculus, vs just knowing the general direction to step? Would that be cheaper to calculate than full derivatives?
- mgh95 11mo ago> But with more advanced optimizers the gradient is not really used directly. It gets per weight normalization, fudged with momentum, clipped, etc. Why would these things be "fudging"? Vanishing gradients (see the initial batch norm paper) are a real thing, and ensuring that the relative magnitudes are in some sense "smooth" between layers allows for an easier optimization problem. > So really, how important is computing the exact gradient using calculus, vs just knowing the general direction to step? Would that be cheaper to calculate than full derivatives? Very. In high dimensional space, small steps can move you extremely far from a proper solution. See adversarial examples.
- ssivark 11mo ago> So really, how important is computing the exact gradient using calculus, vs just knowing the general direction to step? Would that be cheaper to calculate than full derivatives? Yes, absolutely -- a lot of ideas inspired by this have been explored in the field of optimization, and also in machine learning. The very idea of "stochastic" gradient descent using mini-batches basically a cheap (hardware compatible) approximation to the gradient for each step. For a relatively extreme example of how we might circumvent the computational effort of backprop, see Direct Feedback Alignment: https://towardsdatascience.com/feedback-alignment-methods-7e6c41446e36/ https://towardsdatascience.com/feedback-alignment-methods-7e... Ben Recht has an interesting survey of how various learning algorithms used in reinforcement learning relate with techniques in optimization (and how they each play with the gradient in different ways): https://people.eecs.berkeley.edu/~brecht/l2c-icml2018/ https://people.eecs.berkeley.edu/~brecht/l2c-icml2018/ (there's nothing special about RL... as far as optimization is concerned, the concepts work the same even when all the data is given up front rather than generated on-the-fly based on interactions with the environment)
- emil-lp 11mo ago... (2016) 9 years ago, 365 points, 101 comments https://news.ycombinator.com/item?id=13215590 https://news.ycombinator.com/item?id=13215590
- alyxya 11mo agoMore generally, it's often worth learning and understanding things one step deeper. Having a more fundamental understanding of things explains more of the "why" behind why some things are the way they are, or why we do some things a certain way. There's probably a cutoff point for balancing how much you actually need to know though. You could potentially take things a step further by writing the backwards pass without using matrix multiplication, or spend some time understanding what the numerical value of a gradient means.
- phplovesong 11mo agoSidenote why are people still using medium?
- evbogue 11mo agoarticle is from 2016
- joaquincabezas 11mo agoI took a course in my Master's (URV.cat) where we had to do exactly this, implementing backpropagation (fwd and backward passes) from a paper explaining it, using just basic math operations in a language of our choice. I told everyone this was the best single exercise of the whole year for me. It aligns with the kind of activity that I benefit immensely but won't do by myself, so this push was just perfect. If you are teaching, please consider this kind of assignments. P.S. Just checked now and it's still in the syllabus :)
- LPisGood 11mo agoI did this in highschool from some online textbook in plain Java. I recall implementing matrix multiplication myself being the hardest part. I made a UI that showed how the weights and biases changed throughout the training iterations.
- aDyslecticCrow 11mo agoI had a whole course just about how computers do maths. Matrix multiplication, linear fit, finding eigenvectors, multiplication and division, square root, solving linear systems, numerically calculating differential equations, spline interpolation, FEM analysis. "Computers are good at maths" is normally a pretty obvious statement... but many things we take for granted from analytical mathematics, is quite difficult to actually implement in a computer. So there is a mountain of clever algorithms hiding behind some of the seemingly most obvious library operations. One of the best courses I've ever had.
- e-master 11mo agoWould you mind sharing which course it was? Is it available online by any chance?
- aDyslecticCrow 11mo agoUnfortunately it was a course at my university, and in Swedish. But it wouldn't surprise me if there are similar courses online.
- littlestymaar 11mo agoI was happy to see Karpathy writing a new blog post instead of simply Twitter threads, but when I opened the link I just got dispointed to realize it's from 9 years ago… I really hate what Twitter did to blogging…
- Geee 11mo agoHe has a new blog at https://karpathy.bearblog.dev/blog/ https://karpathy.bearblog.dev/blog/
- littlestymaar 11mo agoOh, I wasn't aware of it, thank you very much!
- jamesblonde 11mo agoI have to be contrarian here. The students were right. You didn't need to learn to implement backprop in NumPy. Any leakiness in BackProp is addressed by researchers who introduce new optimizers. As a developer, you just pick the best one and find good hparams for it.
- _diyar 11mo agoFrom the perspective of the university, the students are being trained to become researchers, not engineers.
- PeterStuer 11mo agoThe problem with your reasoning is you never tackle your "unknown unknowns". You just assume they are "known unknowns". Diving through the abstraction reveals some of those.
- gchadwick 11mo agoIt's for a CS course at Stanford not a PyTorch boot camp. It seems reasonable to expect some level of academic rigour and need to learn and demonstrate understanding of the fundamentals. If researchers aren't learning the fundamentals in courses like these where are they learning them? You've also missed the point of the article, if you're building novel model architectures you can't magic away the leakiness. You need to understand the back prop behaviours of the building blocks you use to achieve a good training run. Ignore these and what could be a good model architecture with some tweaks will either entirely fail to train or produce disappointing results. Perhaps you're working at a level of bolting pre built models together or training existing architectures on new datasets but this course operates below that level to teach you how things actually work.
- froobius 11mo ago> Any leakiness in BackProp is addressed by researchers who introduce new optimizers > As a developer, you just pick the best one and find good hparams for it It would be more correct to say: "As a developer, (not researcher), whose main goal is to get a good model working — just pick a proven architecture, hyperparameters, and training loop for it." Because just picking the best optimizer isn't enough. Some of the issues in the article come from the model design, e.g. sigmoids, relu, RNNs. And some of the issues need to be addressed in the training loop, e.g. gradient clipping isn't enabled by default in most DL frameworks. And it should be noted that the article is addressing people on the academic / research side, who would benefit from a deeper understanding.
- brcmthrowaway 11mo agoDo LLMs still use backprop?
- ForceBru 11mo agoAre LLMs still trained by (variants of) stochastic GRADIENT descent? AFAIK what used to be called "backprop" is nowadays known as "automatic differentiation". It's widely used in PyTorch, JAX etc
- imtringued 11mo agoGradient descent doesn't matter here. Second order and higher methods still use lower order derivatives. Back propagation is reverse mode auto differentiation. They are the same thing. And for those who don't understand what back propagation is, it is just an efficient method to calculate the gradient for all parameters.
- deleted 11mo ago[deleted]
- samsartor 11mo agoYes. Pretraining and fine-tuning use standard Adam optimizers (usually with weight-decay). Reinforcement learning has been the odd-man out historically, but these days almost all RL algorithms also use backprop and gradient descent.
- stared 11mo agoThe original title is "Yes you should understand backprop" - which is good and descriptive.
- dpflan 11mo agoAgree, better title for this post; the fact that back prop is a leaky abstraction is a reason one should understand it and know how to do the mechanics by hand to truly experience it and develop understanding and intuition. Software / code abstracting away even more of the process leaves it open to magical thinking. I had to do hand calculations and convolutions in my Deep Learning graduate school course.
- xpe 11mo agoYep. Also, I don’t find the metaphorical connection to leaky abstractions useful at all. It feels strained.
- WithinReason 11mo agoKarpathy suggests the following error: def clipped_error(x): return tf.select(tf.abs(x) < 1.0, 0.5 * tf.square(x), tf.abs(x) - 0.5) # condition, true, false Following the same principles that he outlines in this post, the "- 0.5" part is unnecessary since the gradient of 0.5 is 0, therefore -0.5 doesn't change the backpropagated gradient. In addition, a nicer formula that achieves the same goal as the above is √(x²+1)
- macleginn 11mo agoIf we don't subtract from the second branch, there will be a discontinuity around x = 1, so the derivative will not be well-defined. Also the value of the loss will jump at this value, which will make it hard to inspect the errors, for one thing.
- WithinReason 11mo agoNo, that's not how backprop works. There will be no discontinuity in a backpropagated gradient.
- macleginn 11mo agoI did not say there will be a discontinuity in the gradient; I said that the modified loss function will not have a mathematically well-defined derivative because of the discontinuity in the function.
- WithinReason 11mo agoWhich is completely irrelevant to the point I was making
- kingstnap 11mo agoYou do that to make things smoother when plotted. You could in theory add some crazy stairstep that adds a hundred to the middle part. It would make your loss curves spike and increase towards convergence but then those spikes are just visual artifacts from doing weird discontinuous nonsense with yoru loss.
- joaquincabezas 11mo agooff-topic, anybody knows what's going on with EurekaLabs? It's been a while since the announcement
- meken 11mo agoHe gives an update in the Dwarkesh interview: https://youtu.be/lXUZvyajciY?si=vbqKDOOY7l-491Ka&t=7028 https://youtu.be/lXUZvyajciY?si=vbqKDOOY7l-491Ka&t=7028 Not too many details on timeline - just that he's working on it.
- leobg 11mo agoHe does have a history of abandoning projects. OpenAI. Tesla. OpenAI again... Then again, it might have been the corporate stuff that burned him out rather than the engineering.
- leobg 11mo agoAlso, he published nanochat 3 weeks ago [0], and he says in the readme: > nanochat will become the capstone project of the course LLM101n being developed by Eureka Labs. [0] https://github.com/karpathy/nanochat https://github.com/karpathy/nanochat
- away74etcie 11mo agoKarpathy's work on large datasets for deep neural flow is conceiving of the "backward pass" as the preparation for initializing the mechanics for weight ranges, either as derivatives in -10/+10 statistic deviations.
- sebastianconcpt 11mo agoThis comment: > “Why do we have to write the backward pass when frameworks in the real world, such as TensorFlow, compute them for you automatically?” worries me because is structured with the same reasoning of "why we have to demonstrate we understand addition if in the real world we have calculators"
- rossdavidh 11mo agoThe counter-argument would be that you can make excellent arguments for why we should understand what the compiler is doing, understand what the transistors are doing (e.g. to understand the limitations and risks of overclocking), understand how sort algorithms work (doing sort algorithms manually was at one time a typical interview question). It's not that these things aren't useful, but whether or not that learning would be useful is not the right question. Is this _more_ useful than the other things, which I could be learning but won't because I spent the time and effort to learn this instead? We have a finite capacity for learning, if for no other reason then at least because we have a finite amount of time in this life, and infinite topics to learn (and there are plenty of other constraints besides time). The reason given for learning this topic, is that it has hidden failure modes which you will not be on the lookout for if you didn't know how it worked "under the hood". Is this a good enough reason to spend time learning this rather than, say, how to model the physics of the system you're training the neural network to deal with? Tough question; maybe, maybe not. If you have time to learn both, do that, but if not, then you will have to choose which is most important. And in our education system, we do things like teach calculus but not intermediate statistics, and it would have been better to do the opposite for something like 90% of the people taking calculus. That said, I've implemented backpropagation multiple times, it's a good way to evaluate a new language (just complex enough to reveal problems, not so complex that it takes forever).
- sebastianconcpt 11mo agoBut you went all the way to the extreme. There is a gradient of know-how that is essential for mastery. You can ignore every detail of a transistor, but knowing it can be like a valve is enough to understand that it can model consistently a logic gate of a stream of electrons and that can be used to compute. If wouldn't be for that, then the laws that govern your whole computer become magic for you brain. I fear the consequences of minds that value more their laziness rather to understand, let's say with 5 degrees deep of "why", the things in reality work the way they do. The alternative is that they are hallucinating all the time on narratives that can be psychotic or real and they are not equipped to discern which is which.
- nirinor 11mo agoIts a nit pick, but backpropagation is getting a bad rep here. These examples are about gradients+gradient descent variants being a leaky abstraction for optimization [1]. Backpropagation is a specific algorithm for computing gradients of composite functions, but even the failures that do come from composition (multiple sequential sigmoids cause exponential gradient decay) are not backpropagation specific: that's just how the gradients behave for that function, whatever algorithm you use. The remedy, of having people calculate their own backwards pass, is useful because people are _calculating their own derivatives_ for the functions, and get a chance to notice the exponents creeping in. Ask me how I know ;) [1] Gradients being zero would not be a problem with a global optimization algorithm (which we don't use because they are impractical in high dimensions). Gradients getting very small might be dealt with by with tools like line search (if they are small in all directions) or approximate newton methods (if small in some directions but not others). Not saying those are better solutions in this context, just that optimization(+modeling) are the actually hard parts, not the way gradients are calculated.
- xpe 11mo agoYes. No need to be apologetic or timid about it — it’s not a nit to push back against a flawed conceptual framing. I respect Karpathy’s contributions to the field, but often I find his writing and speaking to be more than imprecise — it is sloppy in the sense that it overreaches and butchers key distinctions. This may sound harsh, but at his level, one is held to a higher standard.
- embedding-shape 11mo ago> often I find his writing and speaking to be more than imprecise I think that's more because he's trying to write to an audience who isn't hardcore deep into ML already, so he simplifies a lot, sometimes to the detriment of accuracy. At this point I see him more as a "ML educator" than "ML practitioner" or "ML researcher", and as far as I know, he's moving in that direction on purpose, and I have no qualms with it overall, he seems good at educating. But I think shifting the mindset of what the purpose of his writings are maybe help understand why sometimes it feels imprecise.
- 11mo ago
- xpe 11mo agoKarpathy is butchering the metaphor. There is no abstraction here. Backprop is an algorithm. Automatic differentiation is a technique. Neither promises to hide anything. I agree that understanding them is useful, but they are not abstractions much less leaky abstractions.
- t-vi 11mo agoIt seems to me that in 2016 people did (have to) play a lot more tricks with the backpropagation than today. Back then it was common to meddle with gradients in between the gradient propagation. For example, Alex Graves's (great! with attention) 2013 paper "Sequence Generation with Recurrent Neural Networks" has this line: One difficulty when training LSTM with the full gradient is that the derivatives sometimes become excessively large, leading to numerical problems. To prevent this, all the experiments in this paper clipped the derivative of the loss with respect to the network inputs to the LSTM layers (before the sigmoid and tanh functions are applied) to lie within a predefined range. with this footnote: In fact this technique was used in all my previous papers on LSTM, and in my publicly available LSTM code, but I forgot to mention it anywhere—mea culpa. That said, backpropagation seems important enough to me that I once did a specialized videocourse just about PyTorch (1.x) autograd.
- HarHarVeryFunny 11mo ago> It seems to me that in 2016 people did (have to) play a lot more tricks with the backpropagation than today Perhaps, but maybe because there was more experimentation with different neural net architectures and nodes/layers back then? Nowadays the training problems are better understood, clipping is supported by the frameworks, and it's easy to find training examples online with clipping enabled. The problem itself didn't actually go away. ReLU (or GELU) is still the default activation for most networks, and training an LLM is apparently something of a black art. Hugging Face just released their "Smol Training Playbook: a distillation of hard earned knowledge to share exactly what it takes to train SOTA LLMs", so evidentially even in 2025 training isn't exactly a turn-key affair.
- joe_the_user 11mo agoI wonder if its just about the important neural nets now being trained by large, secretive corporations that aren't interested in sharing their knowledge.
- HarHarVeryFunny 11mo ago
- raindear 11mo agoAre dead ReLUs still a pronlem today? Why not?
- intelkishan 11mo agoThere are alternative activation functions, which are also widely used.
- mirawelner 11mo agoI feel like my learning curve for AI is: 1) Learn backprop, etc, basic math 2) Learn more advanced things, CNNs, LMM, NMF, PCA, etc 3) Publish a paper or poster 4) Forget basics 5) Relearn that backprop is a thing repeat. Some day I need to get my education together.
- Huxley1 11mo agoWhen I first started learning deep learning, I only had a vague idea of how backprop worked. It wasn't until I forced myself to implement it from scratch that I realized it was not magic after all. The process was painful, but it gave me much more confidence when debugging models or trying to figure out where the loss was getting stuck. I would really recommend everyone in deep learning try writing it out by hand at least once.
- vrighter 11mo agoImplementing one also finally gave me a way to intuitively grasp (and remember) the chain-rule from calculus.
- rao-v 11mo agoI wonder if in the long term, with compute being cheap and parameter volumes being the constraint, if it will make sense to train models to be robust to different activation functions that look like ReLU (I.e. swish, gelu etc.) You might even be able to do a ugly version of this (akin to dropout) where you swap activation functions (with adjusted scaling factors so they mostly yield similar output shapes to ReLU for most input) randomly during training. The point is we mostly know what an ReLU like activation function is supposed to do, so why should we care about the edge cases of the analytical limits of any specific one. The advantage would be that you’d probably get useful gradients out of one of them(for training), and could swap to the computationally cheapest one during inferencing.