11 ms·
LLM Visualization
- holtkam2 3y agoThe visualization I've been looking for for months. I would have happily paid serious money for this... the fact that it's free is such a gift and I don't take it for granted.
- terminous 3y agoSame... this is like a textbook, but worth it
- athulsuresh123 3y agoThis should be in college textbooks
- gryfft 3y agoDamn, this looks phenomenal. I've been wanting to do a deep dive like this for a while-- the 3D model is a spectacular pedagogic device.
- quickthrower2 3y agoAndrej Karpathy twisting his hands as he explains it is also a great device. Not being sarcastic, when he explains it I understand it for a good minute it two. Then need to rewatch as I forget (but that is just me)!
- hodanli 3y agowhich video specifically?
- airesQ 3y agoIncredible work. So much depth; initially I thought it's "just" a 3d model. The animations are amazing.
- reqo 3y agoI feel like visualizations like this are what is missing from univeristy curricula. Now imagine a professor going trough each animation describing exactly what is happening, I am pretty sure students would get a much more in-depth understanding!
- deleted 3y ago[deleted]
- Arson9416 3y agoIsn't it amazing that a random person on the internet can produce free educational content that trumps university courses? With all the resources and expertise that universities have, why do they get shown up all the time? Do they just not know how to educate?
- cmiller1 3y agoHave you considered the considerably greater breadth of content required for a full course, as well as the other responsibilities of the people teaching them such as testing, public speaking, etc.
- dannyw 3y agoIt probably would be massively beneficial to society and progress if teaching professors could spend more time and attention on teaching.
- syntaxers 3y agoIt's an incentives problem. At research universities, promotion is contingent on research output, and teaching is often seen as a distraction. At so-called teaching universities, promotion (or even survival) is mainly contingent on throughput, and not measures of pedagogical outcome. If you are a teaching faculty at a university, it is against your own interests to invest time to develop novel teaching materials. The exception might be writing textbooks, which can be monetized, but typically are a net-negative endeavor.
- smy20011 3y agoReally cool!
- thefourthchime 3y agoFirst off, this is fabulous work. I went through it for the Nano, but is there a way to do the step-by-step for the other LLMs?
- Heidaradar 3y agoBelow the title, there's a few others you can choose from (GPT-2 small and XL and GPT-3)
- valdect 3y agoSuper cool! It's always nice to look something concrete
- deleted 3y ago[deleted]
- warkanlock 3y agoThis is an excellent tool to realize how an LLM actually works from the ground up! For those reading it and going through each step, if by chance you get stuck on why 48 elements are in the first array, please refer to the model.py on minGPT [1] It's an architectural decision that it will be great to mention in the article since people without too much context might lose it [1] https://github.com/karpathy/minGPT/blob/master/mingpt/model.py https://github.com/karpathy/minGPT/blob/master/mingpt/model....
- namocat 3y agoYes, thank you - It was unexplained, so I got stuck on "Why 48?", thinking I'd missing something right out of the gate.
- zombiwoof 3y agoI was thinking 42 ;-)
- riemannzeta 3y agoAre you referring specifically to line 141, which sets the number of embedding elements for gpt-nano to 48? That also seems to correspond to the Channel size C referenced in the explanation text? https://github.com/karpathy/minGPT/blob/master/mingpt/model.py#L141C70-L141C70 https://github.com/karpathy/minGPT/blob/master/mingpt/model....
- tomnipotent 3y agoThat matches the name of default model selected in the right pane, "nano-gpt". I missed the "bigger picture" at first before I noticed the other models in the right pane header.
- taliesinb 3y agoWow, I love the interactive wizzing around and the animation, very neat! Way more explanations should work like this. I've recently finished an unorthodox kind of visualization / explanation of transformers. It's sadly not interactive, but it does have some maybe unique strengths. First, it gives array axis semantic names, represented in the diagrams as colors (which this post also uses). So sequence axis is red, key feature dimension is green, multihead axis is orange, etc. This helps you show quite complicated array circuits and get an immediate feeling for what is going on and how different arrays are being combined with each-other. Here's a pic of the the full multihead self-attention step for example: https://math.tali.link/raster/052n01bav6yvz_1smxhkus2qrik_0736_0884_02kdqvrzq963t.jpg https://math.tali.link/raster/052n01bav6yvz_1smxhkus2qrik_07... It also uses a kind of generalization tensor network diagrammatic notation -- if anyone remembers Penrose's tensor notation, it's like that but enriched with colors and some other ideas. Underneath these diagrams are string diagrams in a particular category, though you don't need to know (nor do I even explain that!). Here's the main blog post introducing the formalism: https://math.tali.link/rainbow-array-algebra https://math.tali.link/rainbow-array-algebra Here's the section on perceptrons: https://math.tali.link/rainbow-array-algebra/#neural-networks https://math.tali.link/rainbow-array-algebra/#neural-network... Here's the section on transformers: https://math.tali.link/rainbow-array-algebra/#transformers https://math.tali.link/rainbow-array-algebra/#transformers
- rikafurude21 3y ago[flagged]
- beckingz 3y agoJust lots and lots and lots and lots of vectors. So Black Magic!
- tsunamifury 3y agoThis shows how the individual weights and vectors work but unless I’m missing something doesn’t quite illustrate yet how higher order vectors are created at the sentence and paragraph level. This might be an emergent property within this system though so it’s hard to “illustrate”. how all of this ends up with a world simulation needs to be understood better and I hope this advances further.
- visarga 3y agoThe deeper you go, the higher the order. It's what attention does at each layer, makes information circulate.
- tsunamifury 3y agoThanks, I assumed that was the case, but they didn't make that explicit. Then the question is, is the world simulation running at the highest order attention layer or is it an emergent property of the interaction cycle between the attention layers.
- deleted 3y ago[deleted]
- 29athrowaway 3y agoI big kudos to the author of this. Not only has the visualization, but it's interactive, has explanations for each item, has excellent performance and is open source: https://github.com/bbycroft/llm-viz/blob/main/src/llm https://github.com/bbycroft/llm-viz/blob/main/src/llm Another interesting visualization related thing: https://github.com/shap/shap https://github.com/shap/shap
- gbertasius 3y agoThis is AMAZING! I'm about to go into Uni and this will be useful for my ML classes.
- baq 3y agoCould as well be titled 'dissecting magic into matmuls and dot products for dummies'. Great stuff. Went away even more amazed that LLMs work as well as they do.
- Simon_ORourke 3y agoThis is brilliant work, thanks for sharing.
- reexpressionist 3y agoDitto. This is the most sophisticated viz of parameters I've seen...and it's also an interactive, step-through tutorial!
- hmate9 3y agoThis is a phenomenal visualisation. I wish I saw this when I was trying to wrap my head around transformers a while ago. This would have made it so much easier.
- wills_forward 3y agoMy jaw drop to see algorhythmic complexity laid out so clearly in a 3d space like that. I wish I was smart enough to know if it's accurate or not.
- block_dagger 3y agoTo know, you must perform intellectual work, not merely be smart. I bet you are smart enough.
- nocoder 3y agoWhat a nice comment!! This has been a big failing of my mental model. I always believed if I was smart enough I should understand things without effort. Still trying to unlearn this....
- blackbear_ 3y agoThat is a surprisingly common fallacy actually; I think you will find this book quite helpful to overcome it: https://www.penguinrandomhouse.com/books/44330/mindset-by-carol-s-dweck-phd/ https://www.penguinrandomhouse.com/books/44330/mindset-by-ca...
- wills_forward 3y agoAw thanks for such encouragement all
- jampekka 3y agoIt's also important to learn how to "teach yourself". Understanding transformers will be really hard if you don't understand basic fully connected feedforward networks (multilayer perceptrons). And learning those is a bit challenging if you don't understand a single unit perceptron. Transformers have the additional challenge of having a bit weird terminology. Keys, queries and values kinda make sense from a traditional information retrieval literature but they're more a metaphor in the attention system. "Attention" and other mentalistic/antrophomorphic terminology can also easily mislead intuitions. Getting a good "learning path" is usually a teacher's main task, but you can learn to figure those by yourself by trying to find some part of the thing you can get a grasp of. Most complicated seeming things (especially in tech) aren't really that complicated "to get". You just have to know a lot of stuff that the thing builds on.
- drdg 3y agoVery cool. The explanations of what each part is doing is really insightful. And I especially like how the scale jumps when you move from e.g. Nano all the way to GPT-3 ....
- atgctg 3y agoA lot of transformer explanations fail to mention what makes self attention so powerful. Unlike traditional neural networks with fixed weights, self-attention layers adaptively weight connections between inputs based on context. This allows transformers to accomplish in a single layer what would take traditional networks multiple layers.
- kmeisthax 3y agoNone of this seems obvious just reading the original Attention is all you need paper. Is there a more in-depth explanation of how this adaptive weighting works?
- WhitneyLand 3y agoIt’s definitely not obvious no matter how smart you are! The common metaphor used is it’s like a conversation. Imagine you read one comment in some forum, posted in a long conversation thread. It wouldn’t be obvious what’s going on unless you read more of the thread right? A single paper is like a single comment, in a thread that goes on for years and years. For example, why don’t papers explain what tokens/vectors/embedding layers are? Well, they did already, except that comment in the thread came 2013 with the word2vec paper! You might think wth? To keep up with this some one would have to spend a huge part of their time just reading papers. So yeah that’s kind of what researchers do. The alternative is to try to find where people have distilled down the important information or summarized it. That’s where books/blogs/youtube etc come in.
- andai 3y agoIs there a way of finding interesting "chains" of such papers, short of scanning the references / "cited by" page? (For example, Google Scholar lists 98797 citations for Attention is all you need!)
- WhitneyLand 3y agoAs a prerequisite to the attention paper? One to check out is: A Survey on Contextual Embeddings https://arxiv.org/abs/2003.07278 https://arxiv.org/abs/2003.07278 Embeddings are sort of what all this stuff is built on so it should help demystify the newer papers (it’s actually newer than the attention paper but a better overview than starting with the older word2vec paper). Then after the attention paper an important one is: Language Models are Few-Shot Learners https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165 I’m intentionally trying to not give a big list because they’re so time-consuming. I’m sure you’ll quickly branch out based on your interests.
- skadamat 3y agoIf folks want a lower dimensional version of this for their own models, I'm a big fan of the Netron library for model architecture visualization. Wrote about it here: https://about.xethub.com/blog/visualizing-ml-models-github-netron https://about.xethub.com/blog/visualizing-ml-models-github-n...
- thierrydamiba 3y agoThanks for sharing. What an exciting time to be learning about LLMs. Everyday I come across a new resource, and everything is free!
- arikrak 3y agoThis looks pretty cool! Anyone know of visualizations for simpler neural networks? I'm aware of tensorflow playground but that's just for a toy example, is there anything for visualizing a real example (e.g handwriting recognition)?
- Logge 3y agohttps://okdalto.github.io/VisualizeMNIST_web/ https://okdalto.github.io/VisualizeMNIST_web/
- atonalfreerider 3y agoWe made a VR visualization back in 2017 https://youtu.be/x6y14yAJ9rY https://youtu.be/x6y14yAJ9rY
- crimsoneer 3y agoI like this one: https://aegeorge42.github.io/ https://aegeorge42.github.io/
- rvz 3y agoRather than looking at the visuals of this network, it is more better to focus on the actual problem with these LLMs which the author already has shown: With in the transformer section: > As is common in deep learning, it's hard to say exactly what each of these layers is doing, but we have some general ideas: the earlier layers tend to focus on learning lower-level features and patterns, while the later layers learn to recognize and understand higher-level abstractions and relationships. That is the problem and yet these black boxes are just as explainable as a magic scroll.
- nlh 3y agoI find this problem fascinating. For decades we’ve puzzled at how the inner workings of the brain works, and thought we’ve learned a lot we still don’t fully understand it. So, we figure, we’ll just make an artificial brain and THEN we’ll be able to figure it out. And here we are, finally a big step closer to an artificial brain and once again, we don’t know how it works :) (Although to be fair we’re spending all of our efforts making the models better and better and not on learning their low level behaviors. Thankfully when we decide to study them it’ll be a wee less invasive and actually doable, in theory.)
- garte 3y agoIs it a brain, though? As far as I understand it it's mostly stochastic calculations based on language or image patterns whose rule sets are static and immutable. Every conceivable idea of plasticity (and with that: a form of fake consciousness) is only present during training. Add to that the fact that a model is being trained actively and the weights are given by humans and the possible realm of outputs is being heavily moderated by an army of faceless low paid workers I don't see any semblance of a brain but a very high maintenance indexing engine sold to the marketing departments of the world as a "brain".
- MRtecno98 3y agoIt's still a neural network, like your brain. It lacks plasticity and can't "learn" autonoumously but it's still one step closer in creating an artifical brain
- shaburn 3y agoVisualization never seems to get the credit due in software development. This is amazing.
- flockonus 3y agoTwitter thread by the author sharing some extra context on this work: https://twitter.com/BrendanBycroft/status/1731042957149827140 https://twitter.com/BrendanBycroft/status/173104295714982714...
- itslennysfault 3y agoThanks for sharing. This is a great thread. Since X now hides replies for non-logged in user here is a nitter link for those without an account (like me) that might want to see the full thread. https://nitter.net/BrendanBycroft/status/1731042957149827140 https://nitter.net/BrendanBycroft/status/1731042957149827140
- 3abiton 3y agoI wish it could integrate other open source LLMs in the backend, but this is already an amazing viz.
- russellbeattie 3y agoI've wondered for a while if as LLM usage matures, there will be an effort to optimize hotspots like what happened with VMs, or auto indexed like in relational DBs. I'm sure there are common data paths which get more usage, which could somehow be prioritized, either through pre-processing or dynamically, helping speed up inference.
- Solvency 3y agoWish it were mobile friendly.
- physPop 3y agoHonestly reading the pytorch implementation of minGTP is a lot more informative than an inscrutable 3d rendering. It's a well commended and pedagogical implementation. I applaud the intention, and it looks slick, but I'm not sure it really conveys information in an efficient way.
- thistoowontpass 3y agoThank you. I'd just completed doing this manually (much uglier and less accurate) and so can really appreciate the effort behind this.
- sva_ 3y agoThe score on this post just went down by a factor of 10 and the time went to "1 hour" ago?!
- myself248 3y agoAnother post was merged with it: https://news.ycombinator.com/item?id=38507672 https://news.ycombinator.com/item?id=38507672
- nikhil896 3y agoThis is by far the best resource I've seen to understand LLMs. Incredibly well done! Thanks for this awesome tool
- tikkun 3y agoWhat happened to this thread? When I saw it before it had 700+ upvotes.
- tikkun 3y agoOh: https://news.ycombinator.com/item?id=38511659 https://news.ycombinator.com/item?id=38511659
- RecycledEle 3y agoThis is excellent! This is why I love Hacker News!
- wdiamond 3y agoamazing router
- johnklos 3y agoAm I the only one getting "Application error: a client-side exception has occurred (see the browser console for more information)." messages?
- lopkeny12ko 3y agoSame here. I blame the popularity of Next.js. More and more of the web is slowly becoming more broken on Firefox on Linux, all with the same tired error: "Application error: a client-side exception has occurred"
- HellsMaddy 3y agoWorks fine on Firefox on Linux for me.
- dathinab 3y agoFor me too. Next.js was never really broken for Firefox for Linux in my experience. Through some "hidden" settings, disabling JS, and proably quite many plugins can brake things. The only thing which tends to be often "broken" for FF in my experience is often CORS and Mic/Camera APIs, ironically in case of CORS 100% because of bugs in non standard compliant websites and for Mic/Camera it's more complicated but most common websites simply refusing to work with FF without even trying (and if you trick them into believing it's no FF often working just fine...).
- altilunium 3y agoIt is possible that your machine does not yet support WebGL2. Check here : https://get.webgl.org/webgl2/ https://get.webgl.org/webgl2/
- fulafel 3y agoOr that you've blocked it (some sources recommend this to avoid fingerprinting so various extensions and privacy configuration receipes do it).
- 3y ago
- abrookewood 3y agoThis does an amazing job of showing the difference in complexity between the different models. Click on GPT-3 and you should be able to see all 4 models side-by-side. GPT-3 is a monster compared to nano-gpt.
- mark_l_watson 3y agoI am looking at Brenden’s GitHub repo https://github.com/bbycroft/llm-viz https://github.com/bbycroft/llm-viz Really nice stuff.
- haltist 3y agoVery cool.
- BSTRhino 3y agobbycroft is the GOAT!
- 8f2ab37a-ed6c 3y agoExpecting someone to implement an LLM in Factorio any day now, we're half-way there already with this blueprint.
- crotchfire 3y agoApplication error: a client-side exception has occurred (see the browser console for more information).
- nandhinianand 3y agoSeems brave blocks some js scripts .. this works in Firefox
- crotchfire 3y agoNot using brave. I get the same fail in both firefox and (chromium-based) qutebrowser. The web is where useful error messages go to die.
- meeb 3y agoThis is easily the best visualization I've seen for a long time. Fantastic work!
- cod1r 3y agoReally cool stuff. Looks like an entire computer but with software. Definitely need to dig into more AI/ML things.
- tysam_and 3y agoAnother visualization I would really love would be a clickable circular set of possible prediction branches, projected onto a Poincare disk (to handle the exponential branching component of it all). Would take forever to calculate except on smaller models, but being able to visualize branch probabilities angularly for the top n values or whatever, and to go forwards and backwards up and down different branches would likely yield some important insights into how they work. Good visualization precludes good discoveries in many branches of science, I think. (see my profile for a longer, potentially more silly description ;) )
- Arctic_fly 3y agoCurse you for being interesting enough to make me get on my desktop.
- SiempreViernes 3y agoAnyone know if there is a name for this 3D control schema? This feels like one of the most intuitive setups I've ever used.
- stareatgoats 3y agoNot sure if it has a name, but you might find out in the github repository: https://github.com/bbycroft/llm-viz/tree/main https://github.com/bbycroft/llm-viz/tree/main
- bbycroft 3y agoAuthor here! Thanks, you'll find the code in https://github.com/bbycroft/llm-viz/blob/main/src/llm/CanvasEventSurface.tsx https://github.com/bbycroft/llm-viz/blob/main/src/llm/Canvas... I don't know a name for it, I just made it up. But for me it's really broken haha 1) When you zoom, the cursor doesn't stay in the same position relative to some projected point 2) Panning also doesn't pin the cursor to a projected point, there's just a hacky multiplier there based on zoom The main issue is that I'm storing the view state as target (on 2D plane) + euler angles + distance. Which is easy to think about, but issues 1 & 2 are better solved by manipulating a 4x4 view matrix. So would just need a matrix -> target-vector-pair conversion to get that working.
- SiempreViernes 3y agoYou know what, I think the panning being pinned to a specific plane is actually great, it means you actually pan across the surface of the object instead of mostly moving it out of the viewport like in this example: https://connectivity.brain-map.org/3d-viewer?v=1&types=PLY&PLY=500%2C184%2C453%2C688 https://connectivity.brain-map.org/3d-viewer?v=1&types=PLY&P... This pinning to one single plane works really well in this particular case because what you are showing is mostly a flat thing anyway, so you don't have much reason to put the view direction close to the plane. A straightforward extension of this behaviour would be to add a few more planes, like one for each of the the cardinal directions and just switch which plane is the one the panning happens in, might be interesting to try for a more rounded object. To me the zoom seems to do what is expected, zooming around the cursor position tends to be disorientating in 3D anyway, though maybe I didn't understand what problem you complained about.
- nbzso 3y agoBeautiful. This should be the new educational standard for visualization of complex topics and systemic thinking.
- codedokode 3y agoThis is a great visualization because original paper on transformers is not very clear and understandable; I tried to read it first and didn't understand so I had to look for other explanations (for example it was unclear for me how multiple tokens are handled). Also, speaking about transformers: they usually append their output tokens to input and process them again. Can we optimize it, so that we don't need to do the same calculations with same input tokens?
- deleted 3y ago[deleted]
- Exuma 3y agoThis is really awesome but I at least wish there were a few added sentences around how I'm supposed to intuitively think about the purpose of why it's like that. For example, I see a T x C matrix of 6 x 48... but at this step, before it's fed into the net, what is this supposed to represent?
- singularity2001 3y agoAlso later why 8 and why is "A" expected in the sixth position
- Workaccount2 3y agoSuch an amazing tool