6 ms·
Since this post is based on my 2014 blog post (https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ https://colah.github.io/posts/2014-03-NN-Manifolds-T
by colah3 1y ago
Since this post is based on my 2014 blog post (https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ ), I thought I might comment.
I tried really hard to use topology as a way to understand neural networks, for example in these follow ups:
- https://colah.github.io/posts/2014-10-Visualizing-MNIST/ https://colah.github.io/posts/2014-10-Visualizing-MNIST/
- https://colah.github.io/posts/2015-01-Visualizing-Representations/ https://colah.github.io/posts/2015-01-Visualizing-Representa...
There are places I've found the topological perspective useful, but after a decade of grappling with trying to understand what goes on inside neural networks, I just haven't gotten that much traction out of it.
I've had a lot more success with:
* The linear representation hypothesis - The idea that "concepts" (features) correspond to directions in neural networks.
* The idea of circuits - networks of such connected concepts.
Some selected related writing:
- https://distill.pub/2020/circuits/zoom-in/ https://distill.pub/2020/circuits/zoom-in/
- https://transformer-circuits.pub/2022/mech-interp-essay/index.html https://transformer-circuits.pub/2022/mech-interp-essay/inde...
- https://transformer-circuits.pub/2025/attribution-graphs/biology.html https://transformer-circuits.pub/2025/attribution-graphs/bio...
- montebicyclelo 1y agoRelated to ways of understanding neural networks, I've seen these views expressed a lot, which to me seem like misconceptions: - LLMs are basically just slightly better `n-gram` models - The idea of "just" predicting the next token, as if next-token-prediction implies a model must be dumb (I wonder if this [1] popular response to Karpathy's RNN [2] post is partly to blame for people equating language neural nets with n-gram models. The stochastic parrot paper [3] also somewhat equates LLMs and n-gram models, e.g. "although she primarily had n-gram models in mind, the conclusions remain apt and relevant". I guess there was a time where they were more equivalent, before the nets got really really good) [1] https://nbviewer.org/gist/yoavg/d76121dfde2618422139 https://nbviewer.org/gist/yoavg/d76121dfde2618422139 [2] https://karpathy.github.io/2015/05/21/rnn-effectiveness/ https://karpathy.github.io/2015/05/21/rnn-effectiveness/ [3] https://dl.acm.org/doi/pdf/10.1145/3442188.3445922 https://dl.acm.org/doi/pdf/10.1145/3442188.3445922
- colah3 1y agoI guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument doesn't ground out to scientific, empirical claims. Our recent paper reverse engineers the computation neural networks use to answer in a number of interesting cases (https://transformer-circuits.pub/2025/attribution-graphs/biology.html https://transformer-circuits.pub/2025/attribution-graphs/bio... ). We find computation that one might informally describe as "multi-step inference", "planning", and so on. I think it's maybe clarifying for this, because it grounds out to very specific empirical claims about mechanism (which we test by intervention experiments). Of course, one can disagree with the informal language we use. I'm happy for people to use whatever language they want! I think in an ideal world, we'd move more towards talking about concrete mechanism, and we need to develop ways to talk about these informally. There was previous discussion of our paper here: https://news.ycombinator.com/item?id=43505748 https://news.ycombinator.com/item?id=43505748
- mdp2021 1y agoAbsolutely, the first task should be to understand how and why black boxes with emergent properties actually work, in order to further knowledge - but importantly, in order to improve them and build on the acquired knowledge to surpass them. That implies, curbing «parrot[ing]» and inadequate «understand[ing]». I.e. those higher concepts are kept in mind as a goal. It is healthy: it keeps the aim alive.
- HarHarVeryFunny 1y ago1) Isn't it unavoidable that a transformer - a sequential multi-layer architecture - is doing multi-step inference ?! 2) There are two aspects to a rhyming poem: a) It is a poem, so must have a fairly high degree of thematic coherence b) It rhymes, so must have end-of-line rhyming words It seems that to learn to predict (hence generate) a rhyming poem, both of these requirements (theme/story continuation+rhyming) would need to be predicted ("planned") at least by the beginning of the line, since they are inter-related. In contrast, a genre like freestyle rap may also rhyme, but flow is what matters and thematic coherence and rhyming may suffer as a result. In learning to predict (hence generate) freestyle, an LLM might therefore be expected to learn that genre-specific improv is what to expect, and that rhyming is of secondary importance, so one might expect less rhyme-based prediction ("planning") at the start of each bar (line).
- winwang 1y agohey chris, I found your posts quite inspiring back then, with very poetic ideas. cool to see you follow up here!
- riemannzeta 1y agoI think it's interesting that in physics, different global symmetries (topological manifolds) can satisfy the same metric structure (local geometry). For example, the same metric tensor solution to Einstein's field equation can exist on topologically distinct manifolds. Conversely, looking at solutions to the Ising Model, we can say that the same lattice topology can have many different solutions, and when the system is near a critical point, the lattice topology doesn't even matter. It's only an analogy, but it does suggest at least that the interesting details of the dynamics aren't embedded in the topology of the system. It's more complicated than that.
- colah3 1y agoIf you like symmetry, you might enjoy how symmetry falls out of circuit analysis of conv nets here: https://distill.pub/2020/circuits/equivariance/ https://distill.pub/2020/circuits/equivariance/
- riemannzeta 1y agoThanks for this additional link, which really underscores for me at least how you're right about patterns in circuits being a better abstraction layer for capturing interesting patterns than topological manifolds. I wasn't familiar with the term "equivariance" but I "woke up" to this sort of approach to understanding deep neural networks when I read this paper, which shows how restricted boltzman machines have an exact mapping to the renormalization group approach used to study phase transitions in condensed matter and high energy physics: https://arxiv.org/abs/1410.3831 https://arxiv.org/abs/1410.3831 At high enough energy, everything is symmetric. As energy begins to drain from the system, eventually every symmetry is broken. All fine structure emerges from the breaking of some symmetries. I'd love to get more in the weeds on this work. I'm in my own local equilibrium of sorts doing much more mundane stuff.
- theahura 1y agoThanks for the follow up. I've been following your circuits thread for several years now. I find the linear representation hypothesis very compelling, and I have a draft of a review for Toy Models of Superposition sitting in my notes. Circuits I find less compelling, since the analysis there feels very tied to the transformer architecture in specific, but what do I know. Re linear representation hypothesis, surely it depends on the architecture? GANs, VAEs, CLIP, etc. seem to explicitly model manifolds. And even simple models will, due to optimization pressure, collapse similar-enough features into the same linear direction. I suppose it's hard to reconcile the manifold hypothesis with the empirical evidence that simple models will place similar-ish features in orthogonal directions, but surely that has more to do with the loss that is being optimized? In Toy Models of Superposition, you're using a MSE which effectively makes the model learn an autoencoder regression / compression task. Makes sense then that the interference patterns between co-occurring features would matter. But in a different setting, say a contrastive loss objective, I suspect you wouldn't see that same interference minimization behavior.
- colah3 1y ago> Circuits I find less compelling, since the analysis there feels very tied to the transformer architecture in specific, but what do I know. I don't think circuits is specific to transformers? Our work in the Transformer Circuits thread often is, but the original circuits work was done on convolutional vision models (https://distill.pub/2020/circuits/ https://distill.pub/2020/circuits/ ) > Re linear representation hypothesis, surely it depends on the architecture? GANs, VAEs, CLIP, etc. seem to explicitly model manifolds (1) There are actually quite a few examples of seemingly linear representations in GANs, VAEs, etc (see discussion in Toy Models for examples). (2) Linear representations aren't necessarily in tension with the manifold hypothesis. (3) GANs/VAEs/etc modeling things as a latent gaussian space is actually way more natural if you allow superposition (which requires linear representations) since central limit theorem allows superposition to produce Gaussian-like distributions.
- theahura 1y ago> the original circuits work was done on convolutional vision models O neat, I haven't read that far back. Will add it to the reading list. To flesh this out a bit, part of why I find circuits less compelling is because it seems intuitive to me that neural networks more or less smoothly blend 'process' and 'state'. As an intuition pump, a vector x matrix matmul in an MLP can be viewed as changing the basis of an input vector (ie the weights act as a process) or as a way to select specific pieces of information from a set of embedding rows (ie the weights act as state). There are architectures that try to separate these out with varying degrees of success -- LSTMs and ResNets seem to have a more clear throughline of 'state' with various 'operations' that are applied to that state in sequence. But that seems really architecture-dependent. I will openly admit though that I am very willing to be convinced by the circuits paradigm. I have a background in molecular bio and there's something very 'protein pathways' about it. > Linear representations aren't necessarily in tension with the manifold hypothesis. True! I suppose I was thinking about a 'strong' form of linear representations, which is something like: features are represented by linear combinations of neurons that display the same repulsion-geometries as observed in Toy Models, but that's not what you're saying / that's me jumping a step too far. > GANs/VAEs/etc modeling things as a latent gaussian space is actually way more natural if you allow superposition Superposition is one of those things that has always been so intuitive to me that I can't imagine it not being a part of neural network learning. But I want to make sure I'm getting my terminology right -- why does superposition necessarily require the linear representation hypothesis? Or, to be more specific, does [individual neurons being used in combination with other neurons to represent more features than neurons] necessarily require [features are linear compositions of neurons]?
- iNic 1y agoMy guess is that the linear representation hypothesis is only approximately right in the sense that my expectation is that it is more like a Lie Group. Locally flat, but the concept breaks at some point. Note that I am a mathematician who knows very little about machine learning apart from taking a few classes at uni
- godelski 1y agoLoved these posts and they inspired a lot of my research and directions during my PhDs. For anyone interested in these may I also suggest learning about normalizing flows? (They are the broader class to flow matching) They are learnable networks that learn coordinate changes. So the connection to geometry/topology is much more obvious. Of course the down side of flows is you're stuck with a constant dimension (well... sorta) but I still think they can help you understand a lot more of what's going on because you are working in a more interpretable environment
- dang 1y agoThat earlier post had a few small HN discussions (for those interested): Neural Networks, Manifolds, and Topology (2014) - https://news.ycombinator.com/item?id=19132702 https://news.ycombinator.com/item?id=19132702 - Feb 2019 (25 comments) Neural Networks, Manifolds, and Topology (2014) - https://news.ycombinator.com/item?id=9814114 https://news.ycombinator.com/item?id=9814114 - July 2015 (7 comments) Neural Networks, Manifolds, and Topology - https://news.ycombinator.com/item?id=7557964 https://news.ycombinator.com/item?id=7557964 - April 2014 (29 comments)
- j2kun 1y agoThis has mirrored my experience attempting to "apply" topology in real world circumstances, off and on since I first studied topology in 2011. I even hesitate now at the common refrain "real world data approximates a smooth, low dimensional manifold." I want to spend some time really investigating to what extent this claim actually holds for real world data, and to what extent it is distorted by the dimensionality reduction method we apply to natural data sets in order to promote efficiency. But alas, who has the time?
- 3abiton 1y agoThe linear representation hypothesis is rather quite intreguing, I am curious what was the intuition behind it.
- colah3 1y agoSee https://transformer-circuits.pub/2022/toy_model/index.html#motivation https://transformer-circuits.pub/2022/toy_model/index.html#m... If you're new to this, I'd mostly just look at all the empirical examples. The slightly harder thing is to consider the fact that neural networks are made of linear functions with non-linearities between them, and to try to think about when linear directions will be computationally natural as a result.
- deleted 1y ago[deleted]
- adamnemecek 1y agoConsider looking into fields related to machine learning to see how topology is used there. The main problem is that some of the cool math did not survive the transition to CS, e.g. the math for control theory is not quite present in RL. In terms of topology, control theory has some very cool topological interpretations, e.g. toruses appear quite a bit in control theory.