14 ms·
The maths you need to start understanding LLMs
- kingkongjaffa 1y agoThe steps in this article are also the same process for doing RAG as well. You computer an embedding vector for your documents or chunks of documents. And then you compute the vector for your users prompt, and then use the cosine distance to find the most semantically relevant documents to use. There are other tricks like reranking the documents once you find the top N documents relating to the query, but that’s basically it. Here’s a good explanation http://wordvec.colorado.edu/website_how_to.html http://wordvec.colorado.edu/website_how_to.html
- stared 1y agoWell, in short - basic linear algebra, basic probability, analysis (functions like exp), gradient. At some point I tried to create an introduction step-by-step, where people can interact with these concepts and see how to express it in PyTorch: https://github.com/stared/thinking-in-tensors-writing-in-pytorch https://github.com/stared/thinking-in-tensors-writing-in-pyt...
- MichaelRazum 1y agoAlthough is it really "understanding" or just able to write down the formulas...?
- misternintendo 1y agoIn some way it is true. Like understanding how a car works purely on physics laws.
- deleted 1y ago[deleted]
- stared 1y agoBeing able to use a formula is the first, and necessary, step for understanding. Then it is able to work at different levels of abstraction and being able to find analogies. But at this point, in my understanding, "understanding" is a never-ending well.
- MichaelRazum 1y agoHow about elliptic curve cryptography then? I just think coming with a formula is not really understanding. Actually most often the “real” formula is the end step of understanding through derivation. ML does it up side down in this regard
- deleted 1y ago[deleted]
- 11101010001100 1y agoApologies for the metacomment, but HN is a funny place. There is a certain type of learning that is deemed good ('math for AI') and a certain type of learning that is deemed bad ('leetcode for AI').
- sgt101 1y agocould you give an example of "HN would not like this AI leetcode"?
- enjeyw 1y agoI mean I kind of get it - overgeneralising (and projecting my own feelings), but I think HN favours introducing and discussing foundational concepts over things that are closer to memorising/wrote-learning. I think AI Math vs Leetcode broadly fits into that category.
- boppo1 1y agoWhat would leetcode for AI be?
- krackers 1y agoI suppose the closest thing might be the type of counting/probability questions asked at quant firms as a way to assess math skill
- raincole 1y agoWhat's leedcode for AI and which site is deemed bad by HN? Without a concrete example it's just a strawman. It could be the site is deemed bad for other reasons. It could be a few vocal negative comments. It could be just not happening.
- apwell23 1y agohonestly i would love 'leetcode for AI' . I am just so sick of all the videos and articles about it.
- pangeranslot 1y ago[dead]
- paradite 1y agoI recently did a livestream on trying to understand attention mechanism (K, Q, V) in LLM. I think it went pretty well (was able to understand most of the logic and maths), and I touched on some of these terms. https://youtube.com/live/vaJ5WRLZ0RE?feature=share https://youtube.com/live/vaJ5WRLZ0RE?feature=share
- kekebo 1y agoI keep having had the best time with Andrej Karpathy's Youtube intros into LLM math. But I haven't compared scope or quality to this submission
- ozgung 1y agoThis is not about _Large_ Language models though. This explains math for word vectors and token embeddings. I see this is the source of confusion for many people. They think LLMs just do this to statistically predict the next word. That was pre-2020s. They ignore the 1.8+ Trillion-parameter Transformer network. Embeddings are just the input of that giant machine. We don't know what is going on exactly in those trillions of parameters.
- ants_everywhere 1y agoBut surely you need this math to start understanding LLMs. It's just not the math you need to finish understanding them.
- HSO 1y ago"necessary but not sufficient"
- ants_everywhere 1y agoyes exactly :)
- HarHarVeryFunny 1y agoIt depends on what level of understanding, and who you are talking about. For the 99% of people outside of software development or machine learning, it is totally irrelevant, as is any details of the Transformer architecture, or the mechanism by which a trained Transformer operates. For the man in the street, inclined to view "AI" as some kind of artificial brain or sentient thing, the best explanation is that basically it's just matching inputs to training samples and regurgitating continuations. Not totally accurate of course, but for that audience at least it gives a good idea and is something they can understand, and perhaps gives them some insight into what it is, how it works/fails, and that it is NOT some scary sentient computer thingy. For anyone in the remaining 1% (or much less - people who actually understand ANNs and machine learning), then learning about the Transformer architecture and how a trained Transformer works (induction heads etc) is what they need to learn to understand what an (Transformer-based, vs LSTM-based) LLM is and how it works. Knowing about the "math" of Transformers/ANNs is only relevant to people who are actually implementing them from ground up, not even those who might just want to build one using PyTorch or some other framework/lbrary where the math has already been done for you. Finally, embeddings aren't about math - they are about representation, which is certainly important to understanding how Transformers and other ANNs work, but still a different topic. * US population of ~300M has ~1M software developers, of which a large fractions are going to be doing things like web development and only at a marginal advantage over someone smart outside of development in terms of learning how ANNs/etc work.
- apwell23 1y ago> Actually coming up with ideas like GPT-based LLMs and doing serious AI research requires serious maths. Does it ? I don't think so. All the math involved is pretty straightforward.
- ants_everywhere 1y agoIt depends on how you define the math involved. Locally it's all just linear algebra with an occasional nonlinear function. That is all straightforward. And by straightforward I mean you'd cover it in an undergrad engineering class -- you don't need to be a math major or anything. Similarly CPUs are composed of simple logic operations that are each easy to understand. I'm willing to believe that designing a CPU requires more math than understanding the operations. Similarly I'd believe that designing an LLM could require more math. Although in practice I haven't seen any difficult math in LLM research papers yet. It's mostly trial and error and the above linear algebra.
- apwell23 1y agoyea i would love to see what complicated math all this came out of. I thought rigorous math was actually an impediment to AI progress. Did any math actually predict or prove that scaling data would create current AI ?
- ants_everywhere 1y agoI was thinking more about the everyday use of more advanced math to solve "boring" engineering challenges. Like finite math to layout chips or kernels. Or improvement to Strassen's algorithm for matrix multiplication. Or improving the transformer KV cache etc. The math you would use to, for example, prove that search algorithm is optimal will generally be harder than the math needed to understand the search algorithm itself.
- matusp 1y agoIt is straight forward because you have been probably exposed to a ton of AI/ML content in your life.
- oulipo2 1y agoAdditions and multiplications. People are making it sound like it's complicated, but NNs have the most basic and simple maths behind The only thing is that nobody understand why they work so well. There are a few function approximation theorems that apply, but nobody really knows how to make them behave as we would like So basically AI research is 5% "maths", 20% data sourcing and engineering, 50% compute power, and 25% trial and error
- amelius 1y agoGradient descent is like pounding on a black box until it gives you the answers you were looking for. Ihere is little more we know about it. We're basically doing Alchemy 2.0. The hard technology that makes this all possible is in semiconductor fabrication. Outside of that, math has comparatively little to do with our recent successes.
- p1dda 1y ago> The only thing is that nobody understand why they work so well. This is exactly what I have ascertained from several different experts in this field. Interesting that a machine has been constructed that performs better than expected and/or is performing more advanced tasks than the inventors expected.
- skydhash 1y agoThe linear regression model "ax + b" is the most simplest one and is still quite useful. It can be interesting to discover some phenomenon that fits the model, but that's not something people have control over. But imagine spending years (expensively) training stuff with millions of weight to ultimately discover it was as simple as "e = mc^2" (and c^2 is basically a constant, so the equation is technically linear)
- rsanek 1y agoAnyone else read the book that the author mentions, Build a Large Language Model (from Scratch) [0]? After watching Karpathy's video [1] I've been looking for a good source to do a deeper dive. [0] https://www.manning.com/books/build-a-large-language-model-from-scratch https://www.manning.com/books/build-a-large-language-model-f... [1] https://www.youtube.com/watch?v=7xTGNNLPyMI https://www.youtube.com/watch?v=7xTGNNLPyMI
- kamranjon 1y agoIt’s good - I’m working through it right now
- ForceBru 1y agoYes, it's really good
- tra3 1y agoIs [1] worth a watch if I want to get a high level/basic understanding of how LLMs work?
- rsanek 1y agoYeah, it's very well done
- malshe 1y agoHere is the code used in the book - https://github.com/rasbt/LLMs-from-scratch https://github.com/rasbt/LLMs-from-scratch
- horizion2025 1y agoIs there a non-video equivalent. I always prefer reading/digesting at my own pace compared to following a video.
- gpjt 1y agoCheck the first link in the parent comment, it's a link to the book.
- InCom-0 1y agoThese are technical details of computations that are performed as part of LLMs. Completely pointless to anyone who is not writing the lowest level ML libraries (so basically everyone). This does now help anyone understand how LLMs actually work. This is as if you started explaining how an ICE car works by diving into chemical properties of petrol. Yeah that really is the basis of it all, but no it is not where you start explaining how a car works.
- ivape 1y agoAlso, people need to accept that they’ve been doing regular ass programming for many years and can’t just jump into whatever they want. The idea that developers were well rounded general engineers is a myth mostly propagated from within the bubble. Most people’s educations right here probably didn’t even involve Linear Algebra (this is a bold claim, because the assumption is that everyone here is highly educated, no cap).
- 49pctber 1y agoAnyone who would like to run an LLM would need to perform their computations on hardware. So picking hardware that is good at matrix multiplication is important for them, even if they didn't develop their LLM from scratch. Knowing the basic math also explains some of the rush to purchase GPUs and TPUs on recent years. All that is kind of missing the point though. I think people being curious and sharpening their mental models of technology is generally a good thing. If you didn't know an LLM was a bunch of linear algebra, you might have some distorted views of what it can or can't accomplish.
- InCom-0 1y agoBeing curious is good ... nothing wrong with that. What I took issue with above is (what I see as) attempt to derail people into low level math when that is not the crux of the question at all. Also: nobody who wants to run LLMs will write their own matrix multiplications. Nobody doing ML / AI comes close to that stuff ... its all abstracted and not something anyone actually thinks about (except the few people who actually write the underlying libraries ie. at Nvidia).
- d_sem 1y agoI think the author did a sufficient job caveating his post without being verbose. While reading through past posts I stumbled on a multi part "Writing an LLM from scratch" series that was an enjoyable read. I hope they keep up writing more fun content.
- petesergeant 1y agoYou need virtually no maths to deeply and intuitively understand embeddings: https://sgnt.ai/p/embeddings-explainer/ https://sgnt.ai/p/embeddings-explainer/
- Mallowram 1y ago[dead]
- armcat 1y agoOne of the most interesting mathematical aspects to me are the fact that LLMs are logit emitters. And associated with this output is uncertainty. Lot of ppl talk about networks of agents. But what you are doing is accumulating uncertainty - every model in the chain introduces its own uncertainty on top of what it inherits. In some situations I've seen a complete collapse after 3 LLM calls chained together. Hence why lot of people recommend "human in the loop" as much as possible to try and reduce that uncertainty (shift the posterior if you will); or they recommend more of a workflow approach - where you have a single orchestrator that decides which function to call, and most of the emphasis (and context engineering) is placed on that orchestrator. But it all ties together in the maths of LLMs.
- deleted 1y ago[deleted]
- Mallowram 1y ago[dead]
- ryanchants 1y agoI'm currently working through Mathematics for Machine Learning and Data Science Specialization from Deeplearning.AI. It's been the best into to Linear Algebra I've found. It's worth the $50 a month just for the quizzes, labs, etc. I'm simultaneously working through the book Math and Architectures of Deep Learning, which is helping re-inforce and flesh out the ideas from the course. [0] https://www.coursera.org/specializations/mathematics-for-machine-learning-and-data-science https://www.coursera.org/specializations/mathematics-for-mac... [1] https://www.manning.com/books/math-and-architectures-of-deep-learning https://www.manning.com/books/math-and-architectures-of-deep...
- zahlman 1y agoIt appears that the "softmax" is found (as I hypothesized by looking at the results, before clicking the link) by exponentiating each value and normalizing to a sum of 1. It would be worthwhile to be explicit. The exponential function is also "high-school maths", and an explanation like that is much easier to follow than the Wikipedia article (since not a lot of rigour is required here).
- jokoon 1y agoML is interesting, but honestly I have trouble knowing the future of it, to see if I should learn the techniques to land a job or not be too obsolete. There is certainly some hype, a lot of what is the market is just not viable.
- fnord77 1y agonothing about vector calculus to minimize loss functions or needing to find Hessians to do Newton's method.
- spinlock_ 1y agoFor me, working through Karpathy's video series (instead of just "watching" them) helped me tremendously to understand how LLMs work and gave me the confidence to read through more advanced material, if I feel like it. But to be honest, the knowledge I gained through his videos are already enough for me. It's kind of like learning how a CPU works in general and ignoring all the fancy optimization steps that I'm not interested in. Thanks Andrej for the time and effort you put into your videos.
- romanoonhn 1y agoCan you share what you mean by "working through" the videos? This playlist has been on my todo for a few weeks so I'm interested.
- spinlock_ 1y agoSure, I was talking about Andrej's "Zero to Hero" playlist: https://youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ&si=yQ87nnq_41uexOGF https://youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9Gv...
- meken 1y ago+1. His cs231n class he taught at Stanford gave me a great foundation.
- karpathy 1y ago<3
- tsunamifury 1y agoI’m sure no one will read this but I was on the team that invented a lot of this early pre-LLM math at Google. It was a really exciting time for me as I had pushed the team to begin looking at vectors beyond language (actions and other predictable perimeters we could extract from linguistic vectors.) We had originally invented a lot of this because we were trying to make chat and email easier and faster, and ultimately I had morphed it into predicting UI decisions based on conversations vectors. Back then we could only do pretty simple predictions (continue vector strictly , reverse vector strictly or N vector options on an axis) but we shipped it and you saw it when we made hangouts, gmail and allo predict your next sentence. Our first incarnation was interesting enough that eric Schmidt recognized it and took my work to the board as part of his big investment in ML. From there the work in hangouts became all/gmail etc. Bizarrely enough though under sundar, this became the Google assistant but we couldn’t get much further without attention layers so the entire project regressed back to fixed bot pathing. I argued pretty hard with the executives that this was a tragedy but sundar would hear none of it, completely obsessed with Alexa and having a competitor there. I found some sympathy with the now head of search who gave me some budget to invest in a messaging program that would advance prediction to get to full action prediction across the search surface and UI. We launched and made it a business messaging product but lost the support of executives during the LLM panic. Sundar cut us and fired the whole team, ironically right when he needed it the most. But he never listened to anyone who worked on the tech and seemed to hold their thoughts in great disdain. What happened after that is of course well known now as sundar ignored some of the most important tech in history due to this attitude. I don’t think I’ll ever fully understand it.
- throwaway-49203 1y ago[dead]
- Rubio78 1y agoWorking through Karpathy's series builds a foundational understanding of LLMs, providing enough confidence to explore further. A key insight is that LLMs are logit emitters, and their inherent uncertainty compounds dangerously in multi-agent chains, often requiring a human-in-the-loop or a single orchestrator to manage it. Crucially, people confuse word embeddings with the full LLM; embeddings are just the input to a vast, incomprehensible trillion-parameter transformer. The underlying math of these networks is surprisingly simple, built on basic additions and multiplications. The real mystery isn't the math but why they work so well. Ultimately, AI research is a mix of minimal math, extensive data engineering, massive compute power, and significant trial and error.
- libraryofbabel 1y agoWay back when, I did a masters in physics. I learned a lot of math: vectors, a ton of linear algebra, thermodynamics (aka entropy), multi-variable and then tensor calculus. This all turned out to be mostly irrelevant in my subsequent programming career. Then LLMs came along and I wanted to learn how they work. Suddenly the physics training is directly useful again! Backprop is one big tensor calculus calculation, minimizing… entropy! Everything is matrix multiplications. Things are actually differentiable, unlike most of the rest of computer science. It’s fun using this stuff again. All but the tensor calculus on curved spacetime, I haven’t had to reach for that yet.
- psb217 1y agoThat past work will pay off even more when you start looking into diffusion and flow-based models for generating images, videos, and sometimes text.
- pornel 1y agoBreakthrough in image generation speed literally came from applying better differential equations for diffusion taken from statistical mechanics physics papers: https://youtu.be/iv-5mZ_9CPY https://youtu.be/iv-5mZ_9CPY
- JBits 1y agoFor me, it's the very basics of general relativity which made the distinction between the cotangent and tangents space click. Optimisation on Riemannian manifolds might give an opportunity to apply more interesting tensor calculus with a non-trivial metric.
- alguerythme 1y agoWell, calculus on curved space, please let me introduce you to: https://arxiv.org/abs/2505.18230 https://arxiv.org/abs/2505.18230 (This is self advertising) If you know how to incorporate time into that, I am interested.
- CrossVR 1y agoAny reason you didn't pick up computer graphics before? Everything is linear algebra and there's even actual physics involved.
- lazarus01 1y agoHere are the building blocks for any deep learning system and a little bit about llm towards the end. Graphs - It all starts with computational graphs. These are data structures that include element wise operations, usually matrix multiplication, addition, activation functions and loss function. The computations are differential, resulting in a smooth continuous space, appropriate for continuous optimization (gradient descent), which is covered later. Layers - Layers are modules comprised of graphs that apply some computation and store the results in a state, referred to as the learned weights. Each Layer learns a deeper, more meaningful representation from the dataset, ultimately learning a latent manifold, which is a highly structured, lower dimensional space, that interpolates between samples, achieving generalization for predictions. Different machine learning problems and data types use different layers, e.g. Transformers for sequence to sequence learning and convolutions for computer vision models, etc. Models - Organize stacks of layers for training. Includes a loss function that sends a feedback signal to an optimizer to adjust learned weights during training. Models also include an evaluation metric for accuracy, independent of the loss function. Forward pass - For training or inference, when an input sequence passes through all the network layers and a geometric transformation is applied producing an output. Backpropagation - Durring training, after the forward pass, gradients are calculated for each weight with respect to the loss, gradients are just another word for derivatives. The process for calculating the derivatives is called automatic differentiation, which is based on the chain rule of derivation. Once the derivatives are calculated the optimizers intelligently updates the weights, with respect to the loss. This is the process called “Learning” often referred to as gradient descent. Now for Large Language Models. Before models are trained for sequence to sequence learning, the corpus of knowledge must be transformed into embeddings. Embeddings are dense representations of language that includes a multidimensional space that can capture meaning and context for different combinations of words that are part of sequences. LLMs use a specific network layer called transformers, that includes something called an attention mechanism. The attention mechanism uses the embeddings to dynamically update the meaning of words when they are brought together in a sequence. The model uses three different representations of the input sequence, called the key, query and value matrices. Using dot product, an attention score is created to identify the meaning of the reference sequence, then a target sequence is generated The output sequence is predicted one word at a time, based on a sampling distribution of the target sequence, using a softmax function.
- nativeit 1y agoAh, I was hoping this would teach me the maths to start understanding the economics surrounding LLMs. That’s the really impossible stuff.
- cultofmetatron 1y agojust wanna plug https://mathacademy.com/courses/mathematics-for-machine-learning https://mathacademy.com/courses/mathematics-for-machine-lear.... happy customer and have found it to be one of the best paid resources for learning mathematics in general. wish I had this when I was a student.
- erdehaas 1y agoThe title is misleading. The maths explained in the blog is the math that is used to build an LLM (how it internally does calculations to do inference etc.). The math to understand LLMs, i.e. that explains in mathematical rigor why LLMs work, is not fully developed yet. That is what the LLM Explainability is about, the effort to understand and clarify the complex, "black-box" decision-making processes of Large Language Models (LLMs) in human-interpretable terms.
- odofog 1y ago[dead]
- orionuni 1y agoThanks for sharing!
- lazarus01 1y agoHere is the bible on deep learning, ”Deep learning with Python” written by Francois Chollet, the creator of Keras. https://www.manning.com/books/deep-learning-with-python https://www.manning.com/books/deep-learning-with-python