12 ms·
How LLMs work
- singpolyma3 4mo agoNext do "why LLMs work"
- sheeshkebab 4mo agoconsidering they work with any architecture/configuration given enough compute, just more or less efficiently - then maybe it's fundamental, in the same sense as why electricity works...
- soupspaces 4mo agoUniversal approximation theorem, embeddings, self-attention, gradient descent. And empirically, scaling laws.
- skydhash 4mo agoWhy does linear regression works? Why does computer works? Because it's about math and the encoding information. If we can encode words as numbers, then why can't we encode their order as a relation? It's just that neural networks are very apt at finding that relation even if it's noisy.
- krackers 4mo agoSee Tegmark's "why does deep cheap learning work so well" (well not so cheap anymore...) https://www.youtube.com/watch?v=5MdSE-N0bxs https://www.youtube.com/watch?v=5MdSE-N0bxs is remarkably prescient given that it was written before LLMs
- inkysigma 4mo agoThis is essentially an open research question. ML theory is unfortunately very weak relative to where the empirics are. I think there's a relatively optimistic paper that was posted a while back here but I would also take it with a grain of salt. https://arxiv.org/abs/2604.21691 https://arxiv.org/abs/2604.21691 There's of course empirical results and relatively weak theoretical results like the UAT but I also don't think that answers your question fully, especially since it seems impossible to definitively answer questions that the industry seems to betting on like whether or not there is a lower bound to their error rate or whether hallucination as a problem can be solved. We have much stronger ideas of what linear regression is doing relative to what LLMs are doing.
- qsera 4mo agoBecause there are patterns everywhere!
- andai 4mo agoI couldn't load the article directly due to an SSL issue, so here's the archive link: https://archive.ph/aWtFG https://archive.ph/aWtFG
- 10GBps 4mo agoI learned TCP/IP by watching and reading raw packets over packet radio at 1200 baud. I've noticed the same thing is possible if you watch the output of a slow LLM. Eventually you start to see the machinery. input tokens = output tokens, it's math. I can't exactly predict the tokens generated but I can see how they are formed. It's a lot like chess. You can't see every possible move but the mechanism is understandable.
- trollbridge 4mo agoComment <-> username synergy.
- fragmede 4mo agohttps://distill.pub/2019/activation-atlas/ https://distill.pub/2019/activation-atlas/ I can only imagine what sort of visualizations are going on today inside of the AI labs.
- Maledictus 4mo agoHow would I set this up?
- barrenko 4mo agoI'd recommend to maybe also specifically watching Karpathy's videos and focusing on the early parts where he specifically deals with tokenization / embeddings generation (which gets really overlooked), and he does this in most of his videos.
- helloplanets 4mo agoIt's basically possible build an LLM using just routers+packets, and then hook them up to Wireshark to see it compute!
- deleted 4mo ago[deleted]
- lhd1 4mo agofind it difficult to engage with AI generated text. What am I getting here that I couldn't get from a chatbot.
- dialsMavis 4mo agoIs this text generated by AI? I couldn't tell but I'd believe it if it was. I imagine if resources were spent writing this text then one benefit of using it is not using more resources or the pollution caused from a chatbot.
- rippeltippel 4mo agoThe voice of several passages resembles ChatGPT very closely.
- zemo 4mo agonormal people talk and write with some notion of meter, the cadence of communicating where pauses are inserted at places that naturally suit the speaker (and listener) to pause for thought. LLM's don't really do that, they just write a bunch of sentences. > Researchers have found that some neurons inside the FFN are strongly associated with specific concepts or facts. One neuron might activate strongly on Eiffel-Tower-related text. Another on programming languages. Another on past-tense verbs. People don't really write like this and they don't really talk like this (and no, people don't necessarily write exactly how they talk because they don't read exactly how they listen; the written word can be backtracked while the heard cannot, and speakers/writers know this, either consciously or unconsciously). A person would probably structure this more like: > Researchers have found that some neurons inside the FFN are strongly associated with specific concepts or facts. For example, there could be one neuron that activates strongly on Eiffel-Tower-related text, another that activates strongly on programming languages, a third neuron activating on past-tense verbs, and so on. Usually people wouldn't write "Another on programming languages." as a standalone sentence like that because the periods introduce an unnatural pause like they're giving a TED talk, unless of course they were punctuating that way for effect, but you'd essentially never communicate with that effect full time.
- malwrar 4mo agoBack when ChatGPT came out, I was so shocked by how _good_ it was for an “AI” product that I simply had to know how it worked. Over the next month I ended up drawing out a block diagram on a whiteboard I have in my office, with the math involved next to each step in the blackboard. I’d puzzle about each step along the way, and the triumph of completing the drawing was also that of this sense of deep understanding. I kept that drawing up for many months after, and would gaze at it often during meetings and idle moments in wonder. This is to say: the autoregressive decoder-only transformer llm architecture as pioneered by openai is wildly simple for how revolutionary its results are. I was reading about non-learned classical SLAM systems (uses video + handcrafted math to produce 3d mappings of physical spaces while also locating the camera in those spaces) at the time, and comparatively speaking I’d say the math is about as complicated as ONE of the components in those complex formulations. The only reason frontier LLMs need 6-figure computers to run is because the model designers made the middle bit in those models REALLY BIG, dimensionally speaking. They just took the steam engine, made a few gargantuan versions of it, and are selling them as the ultimate source of power. This was openai’s entire breakthrough. Making this particular model architecture larger leads to emergent capabilities like being able to pick the best ending to a story/set of instructions or answer questions about broad factual knowledge. I’ve been meanwhile watching these AI companies attempt, successfully, to sell this capability as some sort of robot consciousness hand-crafted by supergeniuses. The fact that they are getting away with it is almost as shocking to me as the discovery itself.
- 10GBps 4mo agoYep. It's nearly identical to the neural nets we were using in the 90s. Back then even a supercomputer wasn't big enough or fast enough to do what we do today. I have to wonder though. Is this all a human brain is? A similar thing to an LLM just scaled exponentially larger. I mean a brain is not just neurons with simple connections to each other. The neurons, axons, dendrites, <insert_unexplained_thing>, etc in a brain are all holding and processing information in different ways and doing it nearly 100% in parallel. That's a really big model. The biological discoveries show how complex a biological brain actually is. Even the tiny brains in a bee or spider are able to solve puzzles and use tools. That's crazy.
- codeakki 4mo agoWhat's the point of this? Im not here to engage with AI bots
- whateveracct 4mo agoaccidentally quadratic
- aabdi 4mo agothis is hard to read... it goes all over the place. i'm not actually sure who your target audience is. there's too many side tangents. just like, structure it plz. 1. customer feels bad cuz they don't understand how llms work 2. provide high level abstracted explanation (don't dive into concepts yet) 3. provide breakdown guide of overall set of components. 4. walk through each component. don't side track. no need to explain, ROPE,GQA etc... it just distracts. i.e. customers don't know how llms work, leading them to feel bad about their own intelligence. at a high level llms take in words, do some math on them, and then produce words, one by one. inside llms have these different components. we walk through them step by step. 1. tokenizer 2. embedding 3. attention 4. heads 5. ffn 6. sampling ## tokenizer
- barrenko 4mo agoIt's just slop.
- LearnYouALisp 4mo agoWithout seeing it, most likely answer based on any apparent confusion and the current broader situation
- lateral_cloud 4mo agoI don't understand how these AI written articles get so many votes.
- alansaber 4mo agoThere is a very high volume of them being posted every day, and they are a significant % of the total. Also, writing is hard, LLM articles can be slop whilst also being better written than average.
- transkey 4mo ago[dead]
- melvinroest 4mo agoI thought Karpathy’s microgpt explain how LLMs work
- disgruntledphd2 4mo agoMicrogpt is really good, if you want to understand exactly what happens. I still thought that this article was a good, higher-level complement to that article though.
- rbrown46 4mo ago[dead]
- spacebacon 4mo agoBut how do they “think”? This is the only repo that can tell you that. https://github.com/space-bacon/SRT https://github.com/space-bacon/SRT
- cubefox 4mo agoWe are living in a crazy science fiction world where on the top of the HN frontpage there is an article on how LLMs work which is likely itself LLM generated, and the only way to tell is its writing style rather than its factual accuracy.
- lionkor 4mo agoIt sucks that this article is clearly LLM edited, with common phrases like "same shape as", "the intuition: ", and the "tiny explainer" which clearly generalized from a prompt accidentally. Good article, but when sharing it I will have to preface "yes it's slop, but it's a good explanation". Absolutely embarrassing that the author didn't catch that these LLM-isms are a (and here I'll use one) bad signal. In fact, I would go so far as to say that publishing in this style stems from a lack of reading experience and writing experience, which does not bode well for someone pretending to be an expert. I gave this article to someone highly intelligent who doesn't know the first thing about how LLMs work internally, and she immediately called out that it reads like AI text.
- janalsncm 4mo agoI don’t think it’s absolutely embarrassing. First of all, the point of the author writing at all is to aid understanding, not produce prose. So from that standpoint, what would be embarrassing would be to include incorrect facts that suggest a fundamental misunderstanding of the topic. From my read, it is fine. The brief history of LLMs is complicated since every single component has papers introducing enhancements. So it’s easy to ignore them or get bogged down with details. The author appears to be a security researcher learning about LLMs for the purpose of defending against common attacks. So this piece is that person giving themselves a crash course on the topic. The fact that they cleaned up their notes with an LLM is frankly completely irrelevant.
- Ampersander 4mo agoYou're not supposed to read it, just like you're not supposed to write anything anymore. Claude can read and write more than any human. We just lean back and relax now.
- vocram 4mo agoSaying an article is of inferior quality just because editing was AI-assisted is like saying a book is lower quality just because it was printed rather than written by hand
- bspammer 4mo agoNo? One affects the actual text and the other doesn’t.
- lateral_cloud 4mo agoAI assisted is a stretch. And that analogy isn't even close to being relevant
- janalsncm 4mo agoNot just that, I think a lot of people are going to waste their time losing the battle (and make no mistake, they will lose) fighting against AI writing without ever asking themselves what makes writing good in the first place. There’s good AI writing and bad organic writing. But it’s easier to point out a few LLM-isms than to actually identify the problems with text.
- blharr 4mo ago> There's good AI writing Sure, but the LLM-isms in AI writing are mentally exhausting to see in every way at this point. The whole point of reading, frankly, is to understand the voice of other people. When you pass that through a distorted filter that makes everyone sound the same... its bad, lossy, frustrating communication It's also dishonest. When you publish something that is direct output without your wording. Digital catfishing at best. The only good AI writing is providing the prompt, because the question is way more interesting, and way more constructive to learning than the answer
- janalsncm 4mo agoThe point of writing is to convey an idea to another person or yourself at a future date. Authenticity has nothing to do with it. I frankly do not care about the “authentic voice” of the author of a random blog. I want to know if they have any interesting ideas.
- stalfie 4mo agoThis article describes how Transformers work, but not really how LLMs work. Explaining the underlying architecture gives you about as much insight into how a modern LLM behaves as an breakdown of neuronal biochemistry and a few pathways does for the brain. Meaning, almost no insight at all.
- helloplanets 4mo agoThe part about positional encoding is not correct. > The intuition: instead of adding position info to each token’s vector, RoPE rotates the vector by an angle that depends on its position You can't rotate the token's entire vector (or all three vectors, whatever is being implied is unclear). You rotate each token's Query and Key vectors only, so dot product can be used to tell how far apart the tokens are when comparing token 1's Query vector to token 2's Key vector. Positional embedding should just be explained after explaining the Query, Key and Value vectors. When the article explains those only after that, the reader is building up on a wrong intuition and it gets confusing.
- giardini 4mo agoCould you restate this another way: I don't follow.
- deleted 4mo ago[deleted]
- rhubarbtree 4mo agoYep, you’re correct. I got to that bit and thought that can’t be right. It’s obviously wrong. If you rotate a semantic vector, you change the semantics of it. You don’t want that. Makes me wonder if the whole thing is just slop. Is the rest of the article correct? Anyone suggest an alternative article?
- aaroninsf 4mo agoThe repetition of some phrases/statements means it's either poorly edited, machine generated, or both.
- AltruisticGapHN 4mo agoI don't like how most LLM explainer articles and videos say that essentially a LLM " predicts the next word". I'm a developer but not very good at maths and I still don't understand any of it. A LLM clearly has some "visual" capacity. You ask Gemini to build something with Canvas and it's able to reason about the shape of things. Like recently I waanted a checkbox that has like a gradient flowing around the edge. It figured out it could use a radial gradient from the center of the checkbox, and overlay that with a small inner div so you only see the edge that looks like the gradient is circling around the checkbox. How is that "predicting the next word"? Not saying AI is intelligent or conscious or anything like that, but the algorithm clearly is far more complex than "predicting words". What I mean, is the LLM is able to represent things in space . That part I don't understand. I also still dont understand the relationship between the chat based LLM and the multi modal stuff. I think I read somewhere when image is generated it is also tokens?
- Borealid 4mo agoYour casual understanding is imprecise. At all times the LLM is, indeed, predicting the next token. Anything it does emerges from that. It did not "figure anything out". It predicted that text describing the use of a radial gradient was likely to follow text describing your problem.
- layla5alive 4mo agoLol, the bird did not 'fly' - it just flapped its wings and generated lift!
- qsera 4mo agoMore like being suspended by a thread...
- Borealid 4mo agoNo. The how is relevant here because it leads to understanding of the resulting behavior. If you train the LLM on a corpus that shows people saying the sky is red, you get an LLM that is predisposed to say the sky is red. This is true even if it's also trained on all of the science that explains how and why the sky is blue. If it were to "figure out" or "reason", it would not have such a predisposition to emit "red" after "the sky is" just because that matches the reward during training. In other words, the token prediction is important because it both explains the successes AND the failures of the LLM. If there were situations in which a bird could fail to fly, then how it tried to fly would also be crucial knowledge.
- mgc_blackbox 4mo ago[dead]
- yukIttEft 4mo ago> so the model figures out during training what each token should look for and what it should offer But how does it learn this token-relationship? All it has is many text samples, but still, nowhere it says how the tokens relate to each other, so where does this information come from?
- inkysigma 4mo agoAt a high level, the text samples are how the relationships are derived. If we treat text samples as sequences of tokens, then the sequences of tokens describe the joint distributions they occur together which confers the relationship between them. Iirc, this is related to the idea of the distributional hypothesis in NLP: the idea the semantics of words should be similar if they occur in similar situations.
- dist-epoch 4mo agoHow does evolution learn the form-fitness relationship? It's the same thing here, you randomly try various token-relationship values and the ones which are slightly better will be favoured.
- MagicMoonlight 4mo agoIf I handed you thousands of documents which said “Jan-Michael Vincent” all over them, would you need to understand who that is in order to notice the relationship there?
- HarHarVeryFunny 4mo agoThe model is just trying to map from sequence to next token. You could say that it doesn't really care about the relationships between words/tokens - it is just being trained to learn the best attention/etc weights to make this mapping as accurate as possible. The model could just as well learn to predict next token from gibberish text as long as there were some statistical gibberish regularities to learn. However, if you train it on real meaningful text then the statistical regularities it needs to learn (and will, thanks to gradient descent, and the capable architecture) will be those reflecting "token relationships" - grammar, semantics, etc. So, you can say the "token relationships" (incl word meanings) are reflected in the statistical regularities of the training data, and the model architecture and training algorithm are just very capable of learning those regularities whatever they may be. You can consider it related to Word2Vec word embeddings, which are based on the idea that the meaning of words comes from how they are used, which to a first approximation can be implemented by considering the meaning of words to be defined by the words they appear next to(!), which is what the Word2Vec embedding training algorithm does, and famous examples such as "(king - man) + woman = queen" prove that this is in fact learning the meanings of words.
- rishbz 4mo agoGreat insights. RL training is the key
- alecco 4mo agoA better blog on Transformers: https://www.aleksagordic.com/blog/transformer https://www.aleksagordic.com/blog/transformer
- whyage 4mo agoStyle nit: the transitions between dark-mode text and large diagrams with a snow white background are jarring.
- mathisdev7 4mo agovery interesting and useful!
- miki123211 4mo agoThere's one thing I wish people understood about LLMs, and it doesn't really have anything to do with what's inside the neural network part. It's the fact that LLMs can only write in one direction — forward. When you are writing an essay and realize midway through a sentence that what you've written doesn't make sense, you go back and edit. An LLM can't do that, the only thing it can do is keep on generating. Because training data typically contains full essays and not half-finished sentences which were then edited, LLMs have a strong preference for "saving face" and producing grammatically correct, internally coherent outputs. They will often do so even if the only way to write themselves out of the corner they wrote themselves into is to lie. To maintain internal coherence, they'll then repeat that lie for the rest of the response. This is also why changing response structure used to affect LLM performance so dramatically. If you asked an LLM to solve a math problem and all-but-forced it to start with the answer, it would have had to calculate that answer before emitting any tokens, something which it very often wasn't able to do. If it was told to follow up the answer with an explanation, it would produce a plausible-sounding explanation to maintain coherence. If, on the other hand, it was told to start by "thinking step by step", it would often be able to solve the first step, and then the next one given the results of the first, and so on, until it was able to reach the answer. Because the answer came last, it wasn't committing to anything, so had no reason to "save face" and lie. This part of the problem is basically solved now with reasoning; reasoning is where all the step-by-step stuff happens, even if users aren't always able to see it. In the process of RLVR, models even train themselves into outputting phrases like "let me check my answer once again" in the chain-of-thought; those serve as their "life rafts" which they can use to both save face and change their answer.
- chris_money202 4mo agoIn terms of our brains though we can only think forward as well (if forward is time). Our brain in the future says something we did in the past was wrong (part of the sentence we wrote) and that informs our body (the agent) to go back and fix it
- kzrdude 4mo agoI think that's why pen and paper is such a good tool for thinking. :)
- runfuyngunasdlj 4mo ago[flagged]
- eddysir 4mo ago[flagged]
- zenfoxai 4mo agoNice article but chain of thought is what makes frontier LLMs smart, not really the token loop
- brcmthrowaway 4mo agoIs chain of thought same as test time compute?
- agumonkey 4mo agoNice intro, gonna help me dig further a lot now. Thanks a ton.
- oceansky 4mo agoOut of curiosity, I wondered if you could break a tokenizer by introducing weird characters not mapped to an id. But apparently, they either just emit a [UNK] token or translate the unrecognized character into raw UTF-8 bytes.
- dotdev_prem 4mo agoWhat is the minimum Nvidia graphics card required to train a 1B model from scratch??
- spaceisballer 4mo agoSolid read especially for someone not in this field. While everything I’ve learned about LLMs has been pretty interesting, all I can say for sure is that I’m more and more skeptical about wide scale adoption. Consumers are being pushed almost to the level of coercion to utilize LLMs. Especially in the case of the government, who in most cases will get a free pass for a year to help build that addiction before the real bill comes due. I’m sure people will object to LLMs being considered statistical bullshit generators, but I cant stop seeing that they just generate bullshit. Bullshit can be believable, sometimes bullshit is truthful. But the public is generally accepting that LLMs churn out truth, and they trust them blindly. I truly don’t see the upside, and the argument being pushed by the dealers is that we need to keep using. The real breakthrough is around the corner. It’s my believe that the tech bros are the new robber barons, they’ve moved on from crypto and NFTs and found something that is pretty impressive but a far cry from the GenAI that wish for. I just feel that technology was supposed to free us and also connect us, but instead it’s all things addictive and consuming.
- Gareth321 4mo agoI’m kind of amazed when I read comments like this, but I have to remind myself that I work in an industry which use these tools at the cutting edge and see what they can really do. In the space of 18 months I changed from skeptic to the belief that our world is going to RAPIDLY change, and soon. I sense that statistics and benchmarks and research and statements from the world’s greatest academics won’t sway you, so maybe I’ll give you a personal anecdote. I have suffered from a condition my whole life called bile acid malabsorption. It caused chronic diarrhoea, pain, arthritis, dehydration, insomnia, and more. I spent decades searching for an answer. Dozens of different tests. Eventually doctors just said I was depressed and prescribed me antidepressants. They didn’t help. On the bad days I considered ending my life. In desperation I turned to ChatGPT. Over months I described my symptoms, triggers, diet, timing, etc. We “sparred” with each other over assumptions and ideas. I gave it all my medical history. All the tests. Eventually it concluded that BAM was likely (plus another few options). So I pushed my doctor for a specialist referral. The specialist agreed to a scan based on the symptoms. It was confirmed. I’ve been taking some cheap medication each day now and it has changed my life. I know others for whom ChatGPT has changed their lives in similar ways. Research shows LLMs are better than doctors already in many cases at diagnosis. They are improving at an exponential rate.
- metaquestions 4mo agoThanks for putting this together. I found it very helpful
- youareartree 4mo ago[dead]
- rramadass 4mo agoAny article titled "How LLMs work" (or similar) and which does not start with conditional and joint probability (high-level and not necessarily detailed) and then show by hand how a trivial language with tokens (eg: "the", "mat", "cat", "sat", "on") can produce the semantically coherent and most likely sentence (eg. "The Cat sat on the Mat") is no good. The intuition for the whole should be built using the above before diving into details of transformers etc.
- enrich717 4mo agoGreat write-up, thanks! I went from zero to a partial understanding after the first read. Will definitely come back to this to fully wrap my head around it.
- sspoisk 4mo ago[flagged]