6 ms·
Mechanistic interpretability researchers applying causality theory to LLMs
https://arxiv.org/abs/2301.04709 https://arxiv.org/abs/2301.04709
- gfody 3mo agothere's a 2MP about the related paper: https://www.youtube.com/watch?v=l72ufA-4SzE https://www.youtube.com/watch?v=l72ufA-4SzE
- calf 3mo agoOne plausible reason I thought of that we may not understand neural nets is that by their nature their power grows with ever-more complex connections and weights. So it is like the opposite of logical systems, in that the very design of neural net architecture is a mess of parameter "spaghetti code" which renders the entire thing a metaphorical encrypted black box. The more powerful an AI/AGI the more this would be the case, and this is analogous a complexity curve. And so any effort to make sense of such black box computation would be like trying to reverse entropy, analogous to trying to recover information lost in waste heat. And that could be one fundamental barrier to understanding both human and artificial brains alike, relative to their internal complexity. (Just thinking aloud my handwavy pet theory recently, I am not an expert and could be totally mistaken on this)
- hsb3 3mo agoYou dont have to understand chemistry to be a good cook tho.
- RobRivera 3mo ago[flagged]
- antleys 3mo agoThis article is not about "reasoning" in the abstract, philosophical sense but is talking about "mechanistic interpretability" research. The title is more like, "can we understand if the 'knowledge' encoded into a neural networks actually corresponds to reasoning-like concepts" and doing that with actual experiments like tweaking weights and activations. There's an interesting example where researchers saw a model approached clock time calculations and calendar month-day calculations using the same methodology. So then is this because an underlying concept of "cyclical measures" has emerged in the network?
- dang 3mo agoThanks - I've attempted to put that in the title above, in the hope of representing the article accurately. (The trouble with a baity title like "Can We Understand How Large Language Models Reason?" is that it generates a barrage of shallow, reflexive responses having little to do with the article. What we want on HN are curious, reflexive responses instead - https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sort=byDate&type=comment&query=reflective%20reflex%20by:dang https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor....)
- danbruc 3mo agoI personally would not look for the way they reason in the weights, at least not directly. In principle I could replace a large language model with a map from all possible input strings to output token or output token distribution without any weights. I have a hard time imagining how you would even tell, at the level of weights and activations, if the next token being the is the result of some proper reasoning or a hallucination. But those weights do not exist for the sake of it, they encode a lot of text the model has seen during training, and I would imagine this is what drives the reasoning. Can you evaluate the following polynomial ... will be related to To evaluate a polynomial ... seen in the training data. This is the level at which I would look for the reasoning, memorized patterns how to do specific things, maybe with some kind of placeholder variables for generalization. Ultimately such a structure would of course also be represented in the weights but I could imagine that this makes it unnecessary hard to understand. Or maybe not, maybe the learned patterns are so complex that they do not have a simple representation.
- BobbyTables2 3mo agoI also suspect the learned patterns are not necessarily efficient, though might be by accident. One could “learn” addition by memorizing a truth table instead of understanding the concept… The truth table itself wouldn’t have much meaning.
- dpark 3mo ago> In principle I could replace a large language model with a map from all possible input strings to output token or output token distribution without any weights. What is an output token distribution except a set of weights?
- dominotw 3mo ago>“Mechanistic interpretability will probably never reduce large language models to a few simple equations,” Icard concluded, “but it may gradually turn deep neural networks into systems whose hidden algorithms can at least partly be understood.” what is the basis for this optimism ?
- azakai 3mo agoThe optimism is based on the successes so far, some of which are described in this article. Scientists have made progress here.
- dominotw 3mo agono they havent . success so far is totally meaningless and doesn't imply any sort of upward slope .
- azakai 3mo agoThe researchers in the field disagree with you. Look at conferences like NeurIPS and ICLR to see a steady stream of incremental progress in this area.
- dominotw 3mo agotrust the researchers bro
- hi-im-buggy 3mo ago[dead]
- Bjartr 3mo agoWhat would real progress look like?
- dang 3mo ago[stub for offtopicness] [[All: please don't post shallow-generic reactions to baity titles. Those are basically the same thing, a la https://en.wikipedia.org/wiki/Rubin_vase https://en.wikipedia.org/wiki/Rubin_vase, and we're trying for something more substantive here.]]
- JackSlateur 3mo agoDo they ?
- throw310822 3mo agoOf course they do, how else do you think they manage to implement new features in large codebases, or to prove new theorems? But you don't even have to assume they do because of the results- you can read their chain of thought.
- 3848499449 3mo ago[flagged]
- ToValueFunfetti 3mo agoFor the love of all that is sacred, please stop doing this. I'm begging you. The whole social media landscape is dying and you are creating a throwaway to participate in ruining this small corner. I assume this is not your first. And no one is convinced by this! The guidelines are there for your benefit as well. You achieve nothing but hastening the destruction of one of the last half-decent communities. Sorry for the melodrama.
- 3848499449 3mo ago[flagged]
- ToValueFunfetti 3mo ago
- 1vuio0pswjnm7 3mo agoOriginal HN title: "Can We Understand How Large Language Models Reason?"
- michaelhoney 3mo agoI also calculate six months from August as 8+6=14mod12. I wonder if anyone does it differently, this seems like the most plausible technique.
- ivandenysov 3mo agoTo calculate +/- 3/6/9 months I shift by seasons. 3rd month of summer becomes 3rd month of winter. That works well cause all months live in a primitive memory palace in my head: an analogue clock face with July at 12 and January at 6. So shifting by 6 means rotating the clock hand from 11 to 5 and immediately visualising what month it falls on. This might sound inefficient to an LLM but human brains had image processing before language.
- windenntw 3mo agoFor me: January = up, April = left, July = down, October = right. Never sure why I did this association, maybe it comes from a drawing in a book I read when I was six or somthn?
- TxFkdZ 3mo ago[dead]