7 ms·
Decomposing language models into understandable components
- dartos 3y agoThis is kind of really cool. All these LLMs appear to be converging around these features.
- r3trohack3r 3y agoOne large model is not how the brain works. It’s not how org charts work. That LLMs are capable of what they are at the compute density they are strongly signals to me that the task of making a productive knowledge worker is in overhang territory. The missing piece isn’t LLM advancement, it’s LLM management. Building trust in an inwardly-adversarial LLM org chart that reports to you.
- PBnFlash 3y agoThe way these systems work feel massively inefficient. We don't re-evaluate our astrophysics models when reading a cooking book.
- DennisP 3y agoThis looks like a big advance in alignment research. A big problem has been that LLMs were just a giant set of inscrutable numbers, and we had no idea what was going on inside. But if this technique scales up, then Anthropic has fixed that. They can figure out what different groups of neurons are actually doing, and use that to control the LLM's behavior. That could help with preventing accidentally misaligned AIs.
- brucethemoose2 3y agoTo me, it sounds more like a good lead for pruning.
- Animats 3y ago> We find that the features that are learned are largely universal between different models, so the lessons learned by studying the features in one model may generalize to others. Hm. I wish they'd said more about that. Does that mean they found the same feature recognizers when training with the same training set? Or what? This tells us something, but what does it tell us?
- karxxm 3y agoSome architectures are relatively well understood. Eg in CNNs, the first layers detect low level features like edges, gradients, etc. The next layer then combines these features to more complex structures like corners or circles. Next layer will combine these features to even higher level features and so on. [1] Typically, you can take a pre-trained model and retrain it on your new dataset by only changing the weights of the last layer(s). Some loss functions even measures the difference between the high-level features of two images, typically extracted from a pre-trained CNN (Perceptual Loss). [1]Matt Zeiler did an amazing work on these findings 10 years ago (https://arxiv.org/abs/1311.2901 https://arxiv.org/abs/1311.2901).
- ilaksh 3y agoI am hoping that this type of research leads into ways to create highly tuned and steerable models that are also much smaller and more efficient. Because if you can see what each part is doing, then theoretically you can find ways to create just the set of features you want. Or maybe tune features that have redundant capacity or something. Maybe by studying the features they will get to the point where the knowledge can be distilled into something more like a very rich and finely defined knowledge graph.
- quickthrower2 3y agoAnthropic must be walking on multi-dimensional tightropes. They want AI safety, and probably want to avoid every Tom, Dick and Harry having a powerful model. But research output picked up by Meta and various discord group could turn the wooly LLMs into powerful contenders and then you have access to the power for all. I don’t have a strong opinion on what is better, but I lean slightly towards models in the open. After all us plebs are allow to use computers and latest CPUs and internet and stuff already! Yes there is shit happening like scams, and worse but it is better than limiting what people can do.
- kalkin 3y agoJust ran across this useful comparison with another very recent paper that effectively corroborates some of the core findings, I believe by an author of the other paper: https://www.lesswrong.com/posts/F4iogK5xdNd7jDNyw/comparing-anthropic-s-dictionary-learning-to-ours https://www.lesswrong.com/posts/F4iogK5xdNd7jDNyw/comparing-...
- pabo 3y agoWhat a great post, thanks for sharing.
- zyxin 3y agoThis makes me wonder what would happen if neural networks contain manually programmed components. It seems like trivial components such as detecting DNA sequences could be programmed in by manually setting the weights. The same thing could be done for example to give neural networks a maths component. Would the network when training discover and make use of these predefined components, or would it ignore them and make up its own ways of detecting DNA sequences?
- WiSaGaN 3y agoThis is indeed interesting. In certain use cases where precision is paramount, we might opt for manually crafted code for the computations. This allows us to be confident in the efficiency of our manual method, rather than relying on LLM for such a specific task. However, it remains unclear whether this would be directly integrated with the network or simply be a tool at LLM's disposal. Interestingly, this situation seems to parallel the choice between enhancing the human brain with something like Neuralink and simply equipping with a calculator.
- drsopp 3y agoI wonder what the limitations are. Do LLM's have Turing completeness?
- btown 3y agoIn a way, this could be considered adding a speculative transformation of the input as part of the input to some layer, and the network deciding whether or not to use that transformation. It would be akin to a convolution layer in a CNN, albeit far more domain-specific. But I’m not sure how much research has been done on weird layers like this!
- astrange 3y agoYou can manually program transformers: https://srush.github.io/raspy/ https://srush.github.io/raspy/ I don't know if you can integrate them into a model. I think you might run out of space, since these aren't polysemantic and so would take up a lot more "room" than learned neurons.
- moralestapia 3y agoOh dang, I am quite literally working on this as a side project (out of mere curiosity). Well, sort of ..., I'm refining an algo that takes several (carefully calibrated) outputs from a given LLM and infers the most plausible set of parameters behind it. I was expecting to find clusters of parameters very much alike to what they observe. I informally call this problem inverting an LLM, and obv., it turns out to be non-trivial to solve. Not completely impossible, tho! as so far I've found some good approximations to it. Anyway, quite an interesting read, def. will keep an eye on what they publish in the future. Also, from the linked manuscript at the end, >Another hypothesis is that some features are actually higher-dimensional feature manifolds which dictionary learning is approximating. Well, you have something that behaves like a continuous, smooth space so you could define as many manifolds as you'd need to suit your needs, so yes :^). But, pedantry off, I get the idea and IMO that's definitely what's going on and the right framework to approach this problem from. One amazing realization one can get from this is, what is the conceptual equivalent of the transition functions that connect all different manifolds in this LLM space? When you see it your mind will be blown, not because of its complexity, but rather because of its exceptional simplicity.
- herodoturtle 3y agoAt first I thought this was an ode to dang.
- stavros 3y agoOh dang, a name so spry, A clever soul, with humor wry, In life's vast game, you do not shy, A friend to all, a bond we tie.
- codethief 3y ago> One amazing realization one can get from this is, what is the conceptual equivalent of the transition functions that connect all different manifolds in this LLM space? Could you elaborate on what you mean by "transition functions" here?
- 3y ago
- adamnemecek 3y agoAll machine learning is just renormalization which in turn is a convolution in Hopf algebra. That's why you see superposition "In physics, wherever there is a linear system with a "superposition principle", a convolution operation makes an appearance." I'm working this out in more details but it is uncanny how much it works out. I have a discord if you want to discuss this further https://discord.cofunctional.ai https://discord.cofunctional.ai
- esafak 3y agoDo you mean all ML or just large neural networks? Where is renormalization in a tree model? What superposition are you referring to?
- adamnemecek 3y agoRenormalization is all about this symmetric partitioning.
- LeonigMig 3y agoI suppose we should be cautious, the human mind is capable of overfitting too
- adamnemecek 3y agoYou have no clue what you are talking about.
- noduerme 3y agoSo, I came up with a pretty decent neural net from scratch about 20 years ago - it ran in the browser in Flash. It basically had a 10x10 bitmap input and an output of the same size, and lots of "neurons" in between that strengthened or weakened their connections based on feedback from the end result. And at a certain point they randomly mutated how they processed the input. I don't see anything wildly different now, other than scale and youth and the hubris that accompanies those things.
- soulofmischief 3y agoYou're describing genetic programming and a very simple neural net, which is cool. However, the utility of transformer models should not be discounted, and if that interested you 20 years ago, you would be blown away by what's possible today.
- quickthrower2 3y agoExcept the emergent properties at scale? At some point you go from making word like sentences, upping the neurons/architecture you get real sounding sentences and then upping again with RLHF loops you get impressive emergent intelligence and ability to solve tasks that were not forseen. It is a rare bird that’s not impressed with 2020s AI.
- nwienert 3y ago> emergent intelligence and ability to solve tasks that were not forseen What's your best examples of this? Some of the most impressive examples I've seen ended up being likely in the dataset, or very close to being so. I've yet to see something where it definitely wasn't approximately in the dataset and was solved in a way that seemed to use some sort of novel process, but open to being wrong.
- soulofmischief 3y agoA good example is the 100s of conversation histories I have with GPT-4 where it does everything from help me code entirely novel and original ideas, or develop more abstract ideas. Every single day, I get immense use out of modern language models. Even if an output is similar to something it's already processed, that's fine! Such is the nature of synthesis.
- gorgoiler 3y agoI am a lay person. To me, I understand a trained model describes transitions from one symbol to the next with probabilities between nodes. There is a structure to this graph — after all if there weren’t then training would be impossible — but this structure is as if it is all written on one sheet of paper with the definitions of each node all inked on top of each other in differed colors. This research (and it’s parent and sibling papers, from the LW article) seem to be about picking out those colored graph components from the floating point soup?
- rewmie 3y agoFrom a machine learning layman's point of view but with some experience with modeling, it's hard to see this as a discovery. Model decomposition and model reduction techniques are very basic concepts in mathematical modeling, and decomposing models in modes with high participation is a very basic technique, which boils down to finding linear combinations of basis that are more expressive. This is even less surprising given LLMs are applied to models with a known hierarchical structure and symmetry. Can anyone say exactly what's novel in these findings? From a layman's point of view, this sounds like announcing the invention of gunpowder.
- zb3 3y ago...so that we can censor these even more.
- jll29 3y agoIs this going to be submitted for publication?
- startupsfail 3y agoWait, embeddings were used for classification for a long time now. Can somebody explain what is new here? edit: ah, looked at the paper, they did it unsupervised, with a sparse autoencoder.
- deleted 3y ago[deleted]
- ffwd 3y agoI'm just curious, how polysemantic is the human brain with each neuron? Cause it feels to me, what you really want, and what the human brain might have, is a high-information (feature based / conceptual based / macro pattern based) monosemantic neural network, and where there is polysemantic neurons, they share similar or the same information in the feature it is a part of (leading to space efficiency? as well as computational efficiency). Whereas in transofmrer models like this, it's as if you're superimposing a million human brains on top of the same network, and then averaging out somehow all the features in the training set into unique neurons (leading naturally to a much larger "brain"). And also they mention in the paper that monosemantic neurons in the network don't work well, but my intuition would be because they are way too "high precision" and they aren't encoding enough information at the feature-level. Features are imo low dimensional, and then a monosemantic high dimensional neuron would the encode way too little information or something. But this is based on my lack of knowledge of the human brain so maybe there are way more similarities than I'm aware of...