12 ms·
For anyone who hasn’t seen this before, mechanistic interpretability solves a very common problem with LLMs: when you ask a model to explain itself, you’re play
by foundry27 2y ago
For anyone who hasn’t seen this before, mechanistic interpretability solves a very common problem with LLMs: when you ask a model to explain itself, you’re playing a game of rhetoric where the model tries to “convince” you of a reason for what it did by generating a plausible-sounding answer based on patterns in its training data. But unlike most trends of benchmark numbers getting better as models improve, more powerful models often score worse on tests designed to self-detect “untruthfulness” because they have stronger rhetoric, and are therefore more compelling at justifying lies after the fact. The objective is coherence, not truth.
Rhetoric isn’t reasoning. True explainability, like what overfitted Sparse Autoencoders claim they offer, basically results in the causal sequence of “thoughts” the model went through as it produces an answer. It’s the same way you may have a bunch of ephemeral thoughts in different directions while you think about anything.
- stavros 2y agoI want to point out here that people do the same: a lot of the time we don't know why we thought or did something, but we'll confabulate plausible-sounding rhetoric after the fact.
- sinuhe69 2y agoNot in math.
- TeMPOraL 2y agoYes in math. Formalisms come after casual thoughts, at every step.
- sinuhe69 2y agoWhat is a casual thought that you cannot explain in math?
- TeMPOraL 2y agoThat question makes no sense. You can explain anything in math, because math is a language and lets you define whatever terms and axioms you need at a given moment. (Whether or not such explanation is useful for anything is another issue entirely.)
- worldsayshi 2y agoCan you explain how intuition led you to try a certain approach?
- TeMPOraL 2y agoIs it enough if I hand-wave it with probability distributions, or do you want me to write out adjacency search in a high-dimensional space?
- mdp2021 2y agoIt's totally different: those formalisms are in a workbench, following a set of rules that either work or not. So, yes, that (math) is representative of the actual process: pattern recognition gives you spontaneous ideas, that you assess for truthfulness in conscious acts of verification.
- legel 2y agoMath comes from brains.
- HeavyStorm 2y agoThat's some misunderstanding of the human brain and thought process...
- LoganDark 2y agoThe split-brain experiment is one of my favorites! https://www.youtube.com/watch?v=wfYbgdo8e-8 https://www.youtube.com/watch?v=wfYbgdo8e-8
- btbuildem 2y agohttps://en.wikipedia.org/wiki/Peace_on_Earth_(novel) https://en.wikipedia.org/wiki/Peace_on_Earth_(novel)
- mdp2021 2y ago/Some/ people bullshit themselves stating the plausible; others check their hypotheses. The difference is total in both humans and automated processes.
- stavros 2y agoHow are you going to check your hypotheses for why you preferred that jacket to that other jacket?
- deleted 2y ago[deleted]
- DSingularity 2y agoIs that example representative for the LLM tasks for which we seek explainability ?
- stavros 2y agoAre we holding LLMs to a higher standard than people?
- f_devd 2y agoIdeally yes, LLMs are tools that we expect to work, people are inherently fallible and (even unintentionally) deceptive. LLMs being human-like in this specific way is not desirable.
- stavros 2y agoThen I think you'll be very disappointed. LLMs aren't in the same category as calculators, for example.
- f_devd 2y agoI have no illusions on LLMs, I have been working with them since og BERT, always with these same issues and more. I'm just stating what would be needed in the future to make them reliably useful outside of creative writing & (human-guided & checked) search. If an LLM provides an incorrect/orthogonal rhetoric without a way to reliably fix/debug it it's just not as useful as it theoretically could be given the data contained in the parameters.
- fsndz 2y agoI stopped at: "causal sequence of “thoughts” "
- benchmarkist 2y agoInterpretability research is basically a projection of the original function implemented by the neural network onto a sub-space of "explanatory" functions that people consider to be more understandable. You're right that the words they use to sell the research is completely nonsensical because the abstract process has nothing to do with anything causal.
- HeatrayEnjoyer 2y agoAll code is causal.
- benchmarkist 2y agoWhich makes it entirely irrelevant as a descriptive term.
- mdp2021 2y ago"Servers shall be strict in formulation and flexible in interpretation."
- Onavo 2y agoHow does the causality part work? Can it spit out a graphical model?
- benreesman 2y agoA lot of the mech interp stuff has seemed to me like a different kind of voodoo: the Integer Quantum Hall Effect? Overloading the term “Superposition” in a weird analogy not governed by serious group representation theory and some clear symmetry? You guys are reaching. And I’ve read all the papers. Spot the postdoc who decided to get paid. But there is one thing in particular that I’ll acknowledge as a great insight and the beginnings of a very plausible research agenda: bounded near orthogonal vector spaces are wildly counterintuitive in high dimensions and there are existing results around it that create scope for rigor [1]. [1] https://en.m.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_lemma https://en.m.wikipedia.org/wiki/Johnson%E2%80%93Lindenstraus...
- txnf 2y agoSuperposition code is a well known concept in information theory - I think there is certainly more to the story then described in the current works, but it does feel like they are going in the right direction
- drdeca 2y agoWhere are you seeing the integer quantum Hall effect mentioned? Or are you bringing it up rather than responding to it being brought up elsewhere? I don’t understand what the connection between IQHE and these SAE interpretability approaches is supposed to be.
- benreesman 2y agoPardon me, the reference is to the fractional Hall effect. "But our results may also be of broader interest. We find preliminary evidence that superposition may be linked to adversarial examples and grokking, and might also suggest a theory for the performance of mixture of experts models. More broadly, the toy model we investigate has unexpectedly rich structure, exhibiting phase changes, a geometric structure based on uniform polytopes, "energy level"-like jumps during training, and a phenomenon which is qualitatively similar to the fractional quantum Hall effect in physics, among other striking phenomena. We originally investigated the subject to gain understanding of cleanly-interpretable neurons in larger models, but we've found these toy models to be surprisingly interesting in their own right." https://transformer-circuits.pub/2022/toy_model/index.html https://transformer-circuits.pub/2022/toy_model/index.html
- snthpy 2y agoA{rt,I} imitating life I believe that's why humans reason too. We make snap judgements and then use reason to try to convince others of our beliefs. Can't recall the reference right now but they argued that it's really a tool for social influence. That also explains why people who are good at it find it hard to admit when they are wrong - they're not used to having to do it because they can usually out argue others. Prominent examples are easy to find - X marks de spot.
- briffid 2y agoJonathan Haidt's The Righteous Mind describes this ín details.
- snthpy 2y agoThanks
- omgwtfbyobbq 2y agoI think Robert Sapolsky's lectures on yt cover this to some degree around 115. https://youtu.be/wLE71i4JJiM?feature=shared https://youtu.be/wLE71i4JJiM?feature=shared Sometimes our cortex is in charge, sometimes other parts of our brain are, and we can't tell the difference. Regardless, if we try to justify it later, that justification isn't always coherent because we're not always using the part of our brain we consider to be rational.
- snthpy 2y agoYes that was probably it because I rewatched that recently. Thanks!
- shshshshs 2y agoPeople who are good at reasoning find it hard to admit that they were wrong? That’s not my experience. People with reason are.. reasonable. You mention X and that’s not where the reasoners are. That’s where the (wanna be) politicians are. Rhetoric is not all of reasoning. I can agree that rationalizing snap judgements is one of our capabilities but I am totally unconvinced that it is the totality of our reasoning capabilities. Perhaps I misunderstood.
- deleted 2y ago[deleted]
- bubaumba 2y agoBTW, it's easy to test model's logic and truthfulness by giving it a wrong decision is if it was its, and asking to explain. Model has no memory and cannot distinguish the source of the text. 'Truthful' model should admit mistake without being asked. Likely model instead will do 'parallel construction' to support 'its' decision.