6 ms·
Show HN: Llama 3.2 Interpretability with Sparse Autoencoders
I spent a lot of time and money on this rather big side project of mine that attempts to replicate the mechanistic interpretability research on proprietary LLMs that was quite popular this year and produced great research papers by Anthropic [1], OpenAI [2] and Deepmind [3].
I am quite proud of this project and since I consider myself the target audience for HackerNews did I think that maybe some of you would appreciate this open research replication as well. Happy to answer any questions or face any feedback.
Cheers
[1] https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html https://transformer-circuits.pub/2024/scaling-monosemanticit...
[2] https://arxiv.org/abs/2406.04093 https://arxiv.org/abs/2406.04093
[3] https://arxiv.org/abs/2408.05147 https://arxiv.org/abs/2408.05147
- OrangeMusic 2y agoThis has been "taken down" and the repo archived. No explanation of what happened.
- dimitry12 2y agoCurious about that too. There are plenty of forks left, for example: https://github.com/plastic-labs/llama3_interpretability_sae https://github.com/plastic-labs/llama3_interpretability_sae (no affiliation)
- jaykr_ 2y agoThis is awesome! I really appreciate the time you took to document everything!
- PaulPauls 2y agoThank you for saying that! I have a much, much harder time documenting everything and writing out each decision in continuous text than actually writing the code. So it took a look time for me to write all of this down - so I'm happy you appreciate it! =)
- curious_cat_163 2y agoHey - Thanks for sharing! Will take a closer look later but if you are hanging around now, it might be worth asking this now. I read this blog post recently: https://adamkarvonen.github.io/machine_learning/2024/06/11/sae-intuitions.html https://adamkarvonen.github.io/machine_learning/2024/06/11/s... And the author talks about challenges with evaluating SAEs. I wonder how you tackled that and where to look inside your repo for understanding the your approach around that if possible. Thanks again!
- PaulPauls 2y agoSo evaluating SAEs - determining which SAE is better at creating the most unique features while being as sparse as possible at the same time - is a very complex topic that is very much at the heart of the current research into LLM interpretability through SAEs. Assuming you already solved the problem of finding multiple perfect SAE architectures and you trained them to perfection (very much an interesting ML engineering problem that this SAE project attempts to solve) then deciding on which SAE is better comes down to which SAE performs better on the metrics of your automated interpretability methodology. Particularly OpenAI's methodology emphasizes this automated interpretability at scale utilizing a lot of technical metrics upon which the SAEs can be scored _and thereby evaluated_. Since determining the best metrics and methodology is such an open research question that I could've experimented on for a few additional months, have I instead opted for a simple approach in this first release. I am talking about my and OpenAI's methodology and the differences between both in chapter 4. Interpretability Analysis [1] in my Implementation Details & Results section. I can also recommend reading the OpenAI paper directly or visiting Anthropics transformer-circuits.pub website that often publishes smaller blog posts on exactly this topic. [1] https://github.com/PaulPauls/llama3_interpretability_sae#4-interpretability-analysis https://github.com/PaulPauls/llama3_interpretability_sae#4-i... [2] https://transformer-circuits.pub/ https://transformer-circuits.pub/
- curious_cat_163 2y agoThanks!
- JackYoustra 2y agoVery cool work! Any plans to integrate it with SAELens?
- PaulPauls 2y agoNot sure yet to be honest. I'll definitely consider it but I'll reorient myself and what I plan to do next in the coming week. I also planned on maybe starting a simpler project and maybe showing people how to create the full model of a current Llama 3.2 implementation from scratch in pure PyTorch. I love building things from teh ground up and when I looked for documentation for the Llama 3.2 background section of this SAE project then the existing documentation I found was either too superficial or outdated and intended for Llama 1 or 2 - Documentation in ML gets outdated so quickly nowadays...
- foundry27 2y agoFor anyone who hasn’t seen this before, mechanistic interpretability solves a very common problem with LLMs: when you ask a model to explain itself, you’re playing a game of rhetoric where the model tries to “convince” you of a reason for what it did by generating a plausible-sounding answer based on patterns in its training data. But unlike most trends of benchmark numbers getting better as models improve, more powerful models often score worse on tests designed to self-detect “untruthfulness” because they have stronger rhetoric, and are therefore more compelling at justifying lies after the fact. The objective is coherence, not truth. Rhetoric isn’t reasoning. True explainability, like what overfitted Sparse Autoencoders claim they offer, basically results in the causal sequence of “thoughts” the model went through as it produces an answer. It’s the same way you may have a bunch of ephemeral thoughts in different directions while you think about anything.
- stavros 2y agoI want to point out here that people do the same: a lot of the time we don't know why we thought or did something, but we'll confabulate plausible-sounding rhetoric after the fact.
- sinuhe69 2y agoNot in math.
- TeMPOraL 2y agoYes in math. Formalisms come after casual thoughts, at every step.
- sinuhe69 2y agoWhat is a casual thought that you cannot explain in math?
- TeMPOraL 2y agoThat question makes no sense. You can explain anything in math, because math is a language and lets you define whatever terms and axioms you need at a given moment. (Whether or not such explanation is useful for anything is another issue entirely.)
- jwuphysics 2y agoIncredible, well-documented work -- this is an amazing effort! Two things that caught my eye were (i) your loss curves and (ii) the assessment of dead latents. Our team also studied SAEs -- trained to reconstruct dense embeddings of paper abstracts rather than individual tokens [1]. We observed a power-law scaling of the lower bound of loss curves, even when we varied the sparsity level and the dimensionality of the SAE latent space. We also were able to totally mitigate dead latents with an auxiliary loss, and we saw smooth sinusoidal patterns throughout training iterations. Not sure if these were due to the specific application we performed (over paper abstracts embeddings) or if they represent more general phenomena. [1] https://arxiv.org/abs/2408.00657 https://arxiv.org/abs/2408.00657
- PaulPauls 2y agoI'm very happy you appreciate it - particularly the documentation. Writing the documentation was much harder for me than writing the code so I'm happy it is appreciated. I furthermore downloaded your paper and will read through it tomorrow morning - thank you for sharing it!
- Eliezer 2y agoThis seems like decent alignment-positive work on a glance, though I haven't checked full details yet. I probably can't make it happen, but how much would someone need to pay you to make up your time, expense, and risk?
- monkeycantype 2y agoThank you for posting this PaulPauls, can I please ask a wacky question that I have about mech.interp. ? we know that when we use a metric to measure humans, we game the metric, I wonder about future ai, gaming mech.interp. for simplicity let me propose a 2d matrix to encode tokens, and hope my understanding of neural networks is good enough for this to be a valid question I understand that higher dimensions have a lot more 'space', so accidental collisions are easily avoided, but 2d makes my example easier to explain. if we had the following tokens mapped in 2d space Apple 1a Pear 1b Donkey 2a Horse 2b it becomes impossible to understand if the neurons 1,2,a,b, all excited represents apple and horse or donkey and pear? I wonder if a future, vastly more competent AI overseeing its own training might use this form of ambiguity as means to create a model in which the weights are selected so the potential for encoding collisions exist, so that it is possible to deceive an mech.int. observer as to the true thoughts of the model, essentially enabling the ai to think in euphemisms?
- Majromax 2y agoWhat you propose is a harder AI safety scenario. You don't need a 'vastly more competent AI overseeing its own training' to elicit this potential problem, just a malicious AI researcher, looking for (e.g.) a model that's racist but that does not have any interperable activation patterns that identifiably correspond to racism. The work here on this Show HN suggests that this kind of adversarial training might just barely be possible for a sufficiently-funded individual, and it seems like novel results would be very interesting.
- batterylake 2y agoThis is incredible! PaulPauls, how would you like us to cite your work?
- PaulPauls 2y agoThank you very much! I included a section at the bottom that provides a sample bibtex citation. I didn't expect this much attention so I didn't even bother with a License but I'll include a MIT license later today and release 0.2.1
- samstevens 2y agoI’m really excited to see some more open SAE work! The engineering effort is non trivial and I’m going to check out your dataloading code tomorrow. You might be interested in an currently in-progress project of mine to train SAEs on vision models: https://github.com/samuelstevens/saev https://github.com/samuelstevens/saev
- lynx23 2y ago[dead]
- Carrentt 2y agoFantastic work! I absolutely love all the documentation.
- vivekkalyan 2y agoThis is great work! Mechanistic interpretability has tons of use cases, it's great to see open research in that field. You mentioned you spent your own time and money on it, would you be willing to share how much you spent? It would help others who might be considering independent research.
- PaulPauls 2y agoThank you, I too am a big believer and enjoyer of open research. The actual code has clarity that complex research papers were never able to convey to me as well as the actual code could. Regarding the cost I would probably sum it up to round about ~2.5k USD for just the actual execution cost. Development cost would've probably doubled that sum if I wouldn't already have a GPU workstation for experiments at home that I take for granted. That cost is made up of: * ~400 USD for ~2 months of storage and traffic of 7.4 TB (3.2 TB of raw, 3.2 TB of preprocessed training data) on a GCP standard bucket * ~100 USD for Anthropic claude requests for experimenting with the right system prompt and test runs and the actual final execution * The other ~2k USD were used to rent 8x Nvidia RTX4090's together with a 5TB SSD from runpod.io for various stages of the experiments. For the actual SAE training I rented the node for 8 days straight and I would probably allocate an additional ~3-4 days of runtime just to experiments to determine the best Hyperparameters for training.
- enterthedragon 2y agoThis is amazing, the documentation is very well organized
- deleted 2y ago[deleted]
- imranhou 2y agoComing from a layman's perspective, a genuine question regarding: "Implements SAE training with auxiliary loss to prevent and revive dead latents, and gradient projection to stabilize training dynamics". I struggle to understand this phrase "to prevent and revive ", perhaps this is simple speak to those that understand the subject of SAEs, but it feels a bit self contradictory to me, could anyone elaborate?
- versteegen 2y agoA latent that is never active and hence doesn't (seem to) represent anything. A loss term to reduce the occurrence of that, and if it does happen, push it back to being active sometimes.
- imranhou 2y agoSo basically preventing dead latents from occurring and whenever they do occur to possibly reviving them through the use of auxiliary loss term in the loss function? Thanks btw
- dontknowit 2y agoI imagine this kind of algorithm are like a derivative, they give a unit response, so you would need another filter to stabilize your system, that is some drop out to remove spurious revived latents.
- PaulPauls 2y agoJust bad wording from me, trying to combine too much information in 1 sentence. The auxiliary loss is supposed to prevent dead latents from occuring in the first place - therefore "prevent dead latents" - and it is also supposed to revive the latents that are already dead - therefore "revive dead latents". Now that I review that sentence again I see that I used 2 verbs on the same subject that could be interpreted differently depending on the verb. Me culpa. I hope you still gained some insights into it =)
- 2y ago
- S-Kaenel 2y agoAmazing research!!
- moconnor 2y agoFind a latent for the Golden Gate bridge and put a Golden Gate Llama 3.2 on HuggingFace. This will get even more attention and love, more so if you include link to a space to chat with it! Also, you didn't ask for suggestions but putting some interesting results / visualizations at the top of the README is a very good idea.
- yangwang92 2y agoNice! You did what I wanted. Have you tried to train SAE for vision encoder and language encoder? I am working on this idea. May we work together, let me initial an issue.
- deleted 2y ago[deleted]
- westurner 2y agoThe relative performance in err/watts/time compared to deep learning for feature selection instead of principal component analysis and standard xgboost or tabular xt TODO for optimization given the indicating features. XAI: Explainable AI: https://en.wikipedia.org/wiki/Explainable_artificial_intelligence#Explainability_and_interpretability_techniques https://en.wikipedia.org/wiki/Explainable_artificial_intelli... /? XAI , #XAI , Explain, EXPLAIN PLAN , error/energy/time
- westurner 2y agoFrom "Interpretable graph neural networks for tabular data" (2023) https://news.ycombinator.com/item?id=37269881 https://news.ycombinator.com/item?id=37269881 : > TabPFN: https://github.com/automl/TabPFN https://github.com/automl/TabPFN .. https://x.com/FrankRHutter/status/1583410845307977733 https://x.com/FrankRHutter/status/1583410845307977733 [2022] "TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second" (2022) https://arxiv.org/abs/2308.08945 https://arxiv.org/abs/2308.08945 > FWIU TabPFN is Bayesian-calibrated/trained with better performance than xgboost for non-categorical data
- westurner 2y agoFrom https://news.ycombinator.com/item?id=34619013 https://news.ycombinator.com/item?id=34619013 : > /? awesome "explainable ai" https://www.google.com/search?q=awesome+%22explainable+ai%22 https://www.google.com/search?q=awesome+%22explainable+ai%22 - (Many other great resources) - https://github.com/neomatrix369/awesome-ai-ml-dl/blob/master/data/model-analysis-interpretation-explainability.md#post-model-creation-analysis-ml-interpretationexplainability https://github.com/neomatrix369/awesome-ai-ml-dl/blob/master... : > Post model-creation analysis, ML interpretation/explainability > /? awesome "explainable ai" "XAI"
- westurner 2y ago"A Survey of Privacy-Preserving Model Explanations: Awesome Privacy-Preserving Explainable AI" https://awesome-privex.github.io/ https://awesome-privex.github.io/
- coolvision 2y agonice! did you use cloud GPUs or built your own machine?
- jzjsj 2y agowhwjwj