4 ms·
The funny thing about LLM is, you can build one from scratch and yet you still won't understand how it works. You would understand what kind of matrix multiplic
by nialv7 1mo ago
The funny thing about LLM is, you can build one from scratch and yet you still won't understand how it works. You would understand what kind of matrix multiplications the neural network performs (in fact that's not that hard. An OS is orders of magnitudes more complex), but you would still have no idea why it does what it does.
- michael0church 1mo ago[dead]
- seanmcdirmid 1mo agoLearning how to build emergent systems is also a skill kids should learn these days. The closest I got was coding up game of life for CSE 142 (intro programming). If stochastic gradient descent isn't taught in whatever CS Theory 101 is now, it really should be these days.
- libraryofbabel 1mo ago> there’s no way to “crack open” an LLM and see precisely where each skill or tendency lives Mechanistic Interpretability has entered the chat. For a classic example, see https://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language... The spirit of your point stands, though. This kind of research is interesting to read about, but it's very hard, more like neuroscience or biology than computer science ("LLMs are grown, not made"). You're dealing with a lot of extremely _messy_ complexity, for which organic life is really the only good point of comparison. Most of us here are't really equipped for that kind of work; it's not at all like, say, reverse-engineering a piece of software written by humans. And of course the only people who can do it on frontier models from Anthropic and OpenAI are people within the labs themselves. (But I'm optimistic we'll see more of this work on open weights models...)
- aschobel 1mo agoYah, "building" it is not sufficient. But a lot of times when I build I want to know the why. "Why does gradient Descent have some clever tricks that easily translate to matrix math"? Lot's of neat stuff to learn.
- danielmarkbruce 1mo agoThis is a silly point. Just because the size of these things is too large for our tiny brains, it doesn't mean we have no idea why it does what it does. If you run a physics simulation of a weather system, you have a situation that is unpredictable for a human - but it's not fair to say "we have no idea why the outputs are what they are!!"
- auggierose 1mo agoThat is not a silly point at all. You confuse understanding the mechanics of it with having a theory of why it works. For the physics simulation, physics provides us with the theories which give us the equations underlying the physics simulation. For neural networks, why is next word prediction giving us AI that can do math? We don't really know!
- danielmarkbruce 1mo agoThere is no "theory" of tomorrow's weather. We understand the math of every single individual equation of an LLM (for example), just like we understand every single equation in the physics simulation. It's the entirety of the system that we don't understand (and hence can't predict) in our heads. From a biology perspective - we have a good understanding of the carbon atom and how forces influence it. We really don't know much about a cell.