9 ms·
Language models pack billions of concepts into 12k dimensions
- bigdict 1y agoWhat's the point of the relu in the loss function? Its inputs are nonnegative anyway.
- Nevermark 1y agoLet's try to keep things positive.
- GolDDranks 1y agoI wondered the same. Seems like it would just make a V-shaped loss around the zero, but abs has that property already!
- fancyfredbot 1y agoRELU would have made it flat below zero ( _/ not \/). Adding the abs first just makes RELU do nothing.
- fancyfredbot 1y agoI thought the belt and braces approach was a valuable contribution to AI safety. Better safe than sorry with these troublesome negative numbers!
- naniwaduni 1y agoWell, I guess it's helping to distinguish authors who are doing arithmetic they understand from ones who are copying received incantations around...
- andy_ppp 1y agoIn reality it’s probably not a RELU modern LLMs use GeLU or something more advanced.
- meindnoch 1y agoSometimes a cosmic ray might hit the sign bit of the register and flip it to a negative value. So it is useful to pass it through a rectifier to ensure it's never negative, even in this rare case.
- lblume 1y agoIndeed, we should call all idempotent functions twice just in case the first incantation fails to succeed. In all seriousness, this is not at all how resilience to cosmic interference works in practice, and the probability of any executed instruction or even any other bit being flipped is far greater than the one specific bit you are addressing.
- deleted 1y ago[deleted]
- js8 1y agoYou can also imagine a similar thing on binary vectors. There two vectors are "orthogonal" if they share no bits that are set to one. So you can encode huge number of concepts using only small number of bits in modestly sized vectors, and most of them will be orthogonal.
- phreeza 1y agoIf they are only orthogonal if they share no bits that are set to one, only one vector, the complement, will be orthogonal, no? Edit: this is wrong as respondents point out. Clearly I shouldn't be commenting before having my first coffee.
- yznovyak 1y agoI don't think so. For n=3 you can have 000, 001, 010, 100. All 4 (n+1) are pairwise orthogonal. However, I don't think js8 is correct as it looks like in 2^n you can't have more than n+1 mutually orthogonal vectors, as if any vector has 1 in some place, no other vector can have 1 in the same place.
- prerok 1y agoHmm, I think one correction: is (0,0,0) actually a vector? I think that, by definition, an n-dimentional space can have at most n vectors which are all orthogonal to one another.
- js8 1y agoIt's not correct to call them orthogonal because I don't think the definition is a dot product. But that aside, yes, orthogonal basis can only have as much elements as dimensions. The article also mentions that, and then introduces "quasi-orthogonality", which means dot product is not zero but very small. On bitstrings, it would correspond to overlap on only small number of bits. I should have been clearer in my offhand remark. :-)
- prerok 1y ago
- dwohnitmok 1y agoThese set of intuitions and the Johnson-Lindenstrauss lemma in particular are what power a lot of the research effort behind SAEs (Sparse Autoencoders) in the field of mechanistic interpretability in AI safety. A lot of the ideas are explored in more detail in Anthropic's 2022 paper that's one of the foundational papers in SAE research: https://transformer-circuits.pub/2022/toy_model/index.html https://transformer-circuits.pub/2022/toy_model/index.html
- emil-lp 1y agoWhere can I read the actual paper? Where is it published?
- yorwba 1y agoThat is the actual paper, it's published on transformer-circuits.pub.
- emil-lp 1y agoIt's not peer-reviewed?
- yorwba 1y agoGoogle Scholar claims 380 citations, which is, I think, a respectable number of peers to have reviewed it.
- emil-lp 1y agoThat's not at all how peer review works.
- yorwba 1y agoIt's not how pre-publication peer review works. There, the problem is that many papers aren't worth reading, but to determine whether it's worth reading or not, someone has to read it and find out. So the work of reading papers of unknown quality is farmed out over a large number of people each reading a small number of randomly-assigned papers. If somebody's paper does not get assigned as mandatory reading for random reviewers, but people read it anyway and cite it in their own work, they're doing a form of post-publication peer review. What additional information do you think pre-publication peer review would give you?
- aabhay 1y agoMy intuition of this problem is much simpler — assuming there’s some rough hierarchy of concepts, you can guesstimate how many concepts can exist in a 12,000-d space by taking the combinatorial of the number of dimensions. In that world, each concept is mutually orthgonal with every other concept in at least some dimension. While that doesn’t mean their cosine distance is large, it does mean you’re guaranteed a function that can linearly separate the two concepts. It means you get 12,000! (Factorial) concepts in the limit case, more than enough room to fit a taxonomy
- Morizero 1y agoThat number is far, far, far greater than the number of atoms in the universe (~10^43741 >>>>>>>> ~10^80).
- cleansy 1y agoNot surprising since concepts are virtual. There is a person, a person with a partner is a couple. A couple with a kid is a family. That’s 5 concepts alone.
- Sharlin 1y agoI’m not sure you grok how big a number 10^43741 is. If we assume that a "concept" is something that can be uniquely encoded as a finite string of English text, you could go up to concepts that are so complex that every single one would take all the matter in the universe to encode (so say 10^80 universes, each with 10^80 particles), and out of 10^43741 concepts you’d still have 10^43741 left undefined.
- jerf 1y agoA concept space of 10^43741 needs about 43741*3 bits to identify each concept uniquely (by the information theoretic concept of bit, which is more a lower bound on what we traditionally think of as bits in the computer world than a match), or about 16000-ish "bytes", which you can approximate reasonably as a "compressed text size". There's a couple orders of magnitude of fiddling around the edges you can do there but you still end up with human-sized quantities of information to identify specific concepts in a space that size rather than massively-larger-than-the-universe sized. Things like novels come from that space. We sample it all the time. Extremely, extremely sparsely, of course. Or to put it another way, in a space of a given size, identifying a specific component takes the log2 of the space's size in bits to identify a concept, not something the size of the space itself. 10^43741 is a very large space by our standards, but the log2 of it is not impossibly large. If it seems weird for models to work in this space, remember that as the models themselves in their full glory are clocking in at multiple hundreds of gigabytes that the space of possible AIs using this neural architecture is itself 2^trillion-ish, which makes 10^43741 look pedestrian. Understanding how to do anything useful with that amount of possibility is quite the challenge.
- yorwba 1y agoI think the author is too focused on the case where all vectors are orthogonal and as a consequence overestimates the amount of error that would be acceptable in practice. The challenge isn't keeping orthogonal vectors almost orthogonal, but keeping the distance ordering between vectors that are far from orthogonal. Even much smaller values of epsilon can give you trouble there. So the claim that "This research suggests that current embedding dimensions (1,000-20,000) provide more than adequate capacity for representing human knowledge and reasoning." is way too optimistic in my opinion.
- sigmoid10 1y agoSince vectors are usually normalized to the surface of an n-sphere and the relevant distance for outputs (via loss functions) is cosine similarity, "near orthogonality" is what matters in practice. This means during training, you want to move unrelated representations on the sphere such that they become "more orthogonal" in the outputs. This works especially well since you are stuck with limited precision floating point numbers on any realistic hardware anyways. Btw. this is not an original idea from the linked blog or the youtube video it references. The relevance of this lemma for AI (or at least neural machine learning) was brought up more than a decade ago by C. Eliasmith as far as I know. So it has been around long before architectures like GPT that could actually be realistically trained on such insanely high dimensional world knowledge.
- motorest 1y ago> Since vectors are usually normalized to the surface of an n-sphere (...) In classification tasks, each feature is normalized independently. Otherwise you would have an entry with feature Foo and Bar which depending on the value of Bar it would be made out to be less Foo when normalized. This vectors are not normalized in n-spheres, and their codomain ends up being an hypercube.
- bjornsing 1y agoI agree the OPs argument is a bad one. But I’m still optimistic about the representational capacity of those 20k dimensions.
- niemandhier 1y agoWow, I think I might just have grasped one of the sources of the problems we keep seeing with LLMs. Johnson-Lichtenstrauss guarantees a distance preserving embedding for a finite set of points into a space with a dimension based on the number of points. It does not say anything about preserving the underlying topology of the contious high dimensional manifold, that would be Takens/Whitney-style embedding results (and Sauer–Yorke for attractors). The embedding dimensions needed to fulfil Takens are related to the original manifolds dimension and not the number of points. It’s quite probable that we observe violations of topological features of the original manifold, when using our to low dimensional embedded version to interpolate. I used AI to sort the hodge pudge of math in my head into something another human could understand, edited result is below: === AI in use === If you want to resolve an attractor down to a spatial scale rho, you need about n ≈ C * rho^(-d_B) sample points (here d_B is the box-counting/fractal dimension). The Johnson–Lindenstrauss (JL) lemma says that to preserve all pairwise distances among n points within a factor 1±ε, you need a target dimension k ≳ (d_B / ε^2) * log(C / rho). So as you ask for finer resolution (rho → 0), the required k must grow. If you keep k fixed (i.e., you embed into a dimension that’s too low), there is a smallest resolvable scale rho* (roughly rho* ≳ C * exp(-(ε^2/d_B) * k), up to constants), below which you can’t keep all distances separated: points that are far on the true attractor will show up close after projection. That’s called “folding” and might be the source of some of the problems we observe . === AI end === Bottom line: JL protects distance geometry for a finite sample at a chosen resolution; if you push the resolution finer without increasing k, collisions are inevitable. This is perfectly consistent with the embedding theorems for dynamical systems, which require higher dimensions to get a globally one-to-one (no-folds) representation of the entire attractor. If someone is bored and would like to discuss this, feel free to email me.
- sdl 1y agoSo basically the map projection problem [1] in higher dimensions? [1] https://en.m.wikipedia.org/wiki/Map_projection https://en.m.wikipedia.org/wiki/Map_projection
- niemandhier 1y agoWorse. Map projection means that you cannot have a mapping that preserves elements of the internal geometry: angles and such. Violation of topology means that a surface wrongly is mapped to one intersecting itself: Think Klein Bottle. https://en.wikipedia.org/wiki/Klein_bottle https://en.wikipedia.org/wiki/Klein_bottle
- rossant 1y agoTangential, but the ChatGPT vibe of most of the article is very distracting and annoying. And I say this as someone who consistently uses AI to refine my English. However, I try to avoid letting it reformulate too dramatically, asking it specifically to only fix grammar and non-idiomatic parts while keeping the tone and formulation as much as possible. Beyond that, this mathematical observation is genuinely fascinating. It points to a crucial insight into how large language models and other AI systems function. By delving into the way high-dimensional data can be projected into lower-dimensional spaces while preserving its structure, we see a crucial mechanism that allows these models to operate efficiently and scale effectively.
- airstrike 1y agoIronically, the use of "fascinating", "crucial" and "delving" in your second paragraph, as well as its overall structure, make it read very much like it was filtered through ChatGPT
- deleted 1y ago[deleted]
- OtherShrezzing 1y agoI think that was satire
- singularity2001 1y agoIf you ever played 20Questions you know that you don't need 1000 dimensions for a billion concepts. These huge vectors can represent way more complex information than just a billion concepts. In fact they can pack complete poems with or without typos and you can ask where in the poem the typo is, which is exactly what happens if you paste that into GPT: somewhere in an internal layer it will distinguish exactly that.
- rini17 1y agoI became bit lost between "C is a constant that determines the probability of success" and then they set C between 4 and 8. Probability should be between 0 and 1, how it relates to C?
- emil-lp 1y agoIt's the epsilon^-2 term that actually talks about success, but that is tightly linked with the C term. If you want to decrease epsilon, C goes up.
- fedeb95 1y ago*string representations of concepts
- gpjanik 1y agoLanguage models don't "pack concepts" into the C dimension of one layer (I guess that's where the 12k number came from), neither do they have to be orthogonal to be viewed as distinct or separate. LLMs generally aren't trained to make distinct concepts far apart in the vector space either. The whole point of dense representations, is that there's no clear separation between which concept lives where. People train sparse autoencoders to work out which neurons fire based on the topics involved. Neuronpedia demonstrates it very nicely: https://www.neuronpedia.org/ https://www.neuronpedia.org/.
- prmph 1y agoAgreed, if you relax the requirement for perfect orthogonality, then, yes, you can pack in much more info. You basically introduced additional (fractional) dimensions clustered with the main dimensions. Put another way, many concepts are not orthogonal, but have some commonality or correlation. So nothing earth shattering here. The article is also filled with words like "remarkable", "fascinating", "profound", etc. that make me feel like some level of subliminal manipulation is going on. Maybe some use of an LLM?
- gpjanik 1y agoIt's... really not what I meant. This requirement does not have to be relaxed, it doesn't exist at all. Semantic similarity in embedding space is a convenient accident, not a design constraint. The model's real "understanding" emerges from the full forward pass, not the embedding geometry.
- prmph 1y agoI'm speaking in more in general conceptual terms, not about the specifics of LLM architecture
- sdenton4 1y agoThe spare autoencoder work is /exactly/ premised on the kind of near-orthogonality that this article talks about. It's called the 'superposition hypothesis' originally: https://transformer-circuits.pub/2022/toy_model/index.html https://transformer-circuits.pub/2022/toy_model/index.html The SAE's job is to try to pull apart the sparse nearly-orthogonal 'concepts' from a given embedding vector, by decomposing the dense vector into a sparsely activation over-complete basis. They tend to find that this works well, and even allows matching embedding spaces between different LLMs efficiently.
- deleted 1y ago[deleted]
- jibal 1y agoThere are no "real-world concepts" or "semantic meaning" in LLMs, there are only syntactic relationships among text tokens.
- lblume 1y agoThat really stretches the meaning of "syntactic". Humans have thoroughly evaluated LLMs and discovered many patterns that very cleanly map to what they would consider real-world concepts. Semantic properties do not require any human-level understanding; a Python script has specific semantics one may use to discuss its properties, and it has become increasingly clear that LLMs can reason (as in derive knowable facts, extract logical conclusions, compare it to different alternatives; not having a conscious thought process involving them) about these scripts not just by their syntactic but also semantic properties (of course bounded and limited by Rice's theorem).
- jibal 1y ago> Humans have thoroughly evaluated LLMs and discovered many patterns that very cleanly map to what they would consider real-world concepts Well yes, humans have real-world concepts. > Semantic properties do not require any human-level understanding Strawman. > a Python script has specific semantics one may use to discuss its properties These are human-attributed semantics. To say that a static script "has" semantics is a category mistake--certainly it doesn't "have" them the way LLMs are purported by the OP to have concepts. > it has become increasingly clear that LLMs can reason (as in derive knowable facts, extract logical conclusions, compare it to different alternatives These are highly controversial claims. LLMs present conclusions textually that are implicit in the training data. To get from there to the claim that they can reason is a huge leap. Certainly we know (from studies by Anthropic and elsewhere) that the reasoning steps that LLMs claim to go through are not actual states of the LLM. I'm not going to say more about this ... it has been discussed at length in the academic literature.
- empath75 1y agoDo you learn anything from reading books or is everything you know entirely derived from personal experience.
- rob_c 1y agoOk. Now try to separate the "learning the language" from "learning the data". If we have a model pre trained on language does it then learn concepts quicker, the same or different? Can we compress just data in a lossy into an LLM like kernel which regenerates the input to a given level of fidelity?
- highfrequency 1y ago> posed a fascinating question: How can a relatively modest embedding space of 12,288 dimensions (GPT-3) accommodate millions of distinct real-world concepts? Because there is a large number of combinations of those 12k dimensions? You don’t need a whole dimension for “evil scientist” if you can have a high loading on “evil” and “scientist.” There is quickly a combinatorial explosion of expressible concepts. I may be missing something but it doesn’t seem like we need any fancy math to resolve this puzzle.
- stared 1y agoIf vectors life in an effectively lower space that they could, they don't live up to their n-dimensional potential. Sometimes these things are patched with cosine distance (or even - Pearson correlation), vide https://p.migdal.pl/blog/2025/01/dont-use-cosine-similarity https://p.migdal.pl/blog/2025/01/dont-use-cosine-similarity. Ideally when we don't need to and vectors occupy the space. I am kind of surprised that the original article does not mention batch normalization and similar operations - these are pretty much created to automatically de-bias and de-correlate values at each layer.
- alexpivnenko 1y agoInteresting that the practical C values were much below the theoretical bounds.
- lvl155 1y agoThey don’t capture concepts at all. They capture writings of concepts.
- cgadski 1y ago> The implications of these geometric properties are staggering. Let's consider a simple way to estimate how many quasi-orthogonal vectors can fit in a k-dimensional space. If we define F as the degrees of freedom from orthogonality (90° - desired angle), we can approximate the number of vectors as [...] If you're just looking at minimum angles between vectors, you're doing spherical codes. So this article is an analysis of spherical codes… that doesn't reference any work on spherical codes… seems to be written in large part by a language model… and has a bunch of basic inconsistencies that make me doubt its conclusions. For example: in the graph showing the values of C for different values of K and N, is the x axis K or N? The caption says the x axis is N, the number of vectors, but later they say the value C = 0.2 was found for "very large spaces," and in the graph we only get C = 0.2 when N = 30,000 and K = 2---that is, 30,000 vectors in two dimensions! On the other hand, if the x axis is K, then this article is extrapolating a measurement done for 2 vectors in 30,000 dimensions to the case of 10^200 vectors in 12,888 dimensions, which obviously is absurd. I want to stay positive and friendly about people's work, but the amount of LLM-driven stuff on HN is getting really overwhelming.
- jvanderbot 1y agoThe problem with saying something is LLM generated is it cannot be proven and is a less-helpful way of saying it has errors. Pointing out the errors is a more helpful way if stating problems with the article, which you have also done. In that particular picture, you're probably correct to interpret it as C vs N as stated.
- Blackthorn 1y ago> The problem with saying something is LLM generated is it cannot be proven and is a less-helpful way of saying it has errors. It's a very helpful way of saying it shouldn't be bothered to be read. After all, if they couldn't be bothered to write it, I can't be bothered to read it.
- deleted 1y ago[deleted]
- stogot 1y agoIs the definition of dimensions here the same as 2D, 3D, 4D, etc or some other abstract mathematical concept?
- prerok 1y agoWell, yes, the same concept, it's an N-dimensonal space, but outside of LLMs, we would rarely work with N=12000.
- cpldcpu 1y agoThe dimensions should actually be closer to 12000 * (no of tokens*no of layers / x) (where x is a number dependent on architectural features like MLHA, QGA...) There is this thing called KV cache which holds an enormous latent state.
- mallowdram 1y agoSpace embedding based on arbitrary points never resolves to specifics. Particularly downstream. Words are arbitrary, we remained lazy at an unusually vague level of signaling because arbitrary signals provide vast advantages for the sender and controller of the signal. Arbitrary signals are essentially primate dominance tools. They are uniquely one-way. CS never considered this. It has no ability to subtract that dark matter of arbitrary primate dominance that's embedded in the code. Where is this in embedded space? LLMs are designed for Western concepts of attributes, not holistic, or Eastern. There's not one shred of interdependence, each prediction is decontextualized, the attempt to reorganize by correction only slightly contextualizes. It's the object/individual illusion in arbitrary words that's meaningless. Anyone studying Gentner, Nisbett, Halliday can take a look at how LLMs use language to see how vacant they are. This list proves this. LLMs are the equivalent of circus act using language. "Let's consider what we mean by "concepts" in an embedding space. Language models don't deal with perfectly orthogonal relationships – real-world concepts exhibit varying degrees of similarity and difference. Consider these examples of words chosen at random: "Archery" shares some semantic space with "precision" and "sport" "Fire" overlaps with both "heat" and "passion" "Gelatinous" relates to physical properties and food textures "Southern-ness" encompasses culture, geography, and dialect "Basketball" connects to both athletics and geometry "Green" spans color perception and environmental consciousness "Altruistic" links moral philosophy with behavioral patterns"
- ausbah 1y agoaren’t outputs literally conditioned on prior textual context? how is that lacking interdependence? isn’t learning the probabilistic relationships between tokens an attempt to approximate those exact semantic relationships between words?
- mallowdram 1y agoInterdependence takes into account the Universe for each thing or idea. There is no such thing as probabilistic in a healthy mind. A probabilistic approach is unhealthy. https://pubmed.ncbi.nlm.nih.gov/38579270/ https://pubmed.ncbi.nlm.nih.gov/38579270/ edit: looking into this, this is likely in terms of the brain and arbitrariness highly paradoxical even oxymoronic >>isn’t learning the probabilistic relationships between tokens an attempt to approximate those exact semantic relationships between words? This is really a poor manner of resolving the conduit metaphor condition to arbitary signals, to falsify them as specific, which is always impossible. This is simple linguistic via animal signal science. If you can't duplicate any response with a high degreee of certainty from output, then the signal is only valid in the most limited time-space condition and yet it is still arbitrary. CS has no understanding of this.
- WithinReason 1y agoThe vectors don't need to be orthogonal due to the use of non-linearities in neural networks. The softmax in attention let's you effectively pack as many vectors in 1D as you want and unambiguously pick them out.
- djoldman 1y agoA continuing, probably unending, opportunity/tragedy is the under-appreciation of representation learning / embeddings. The magic of many current valuable models is simply that they can combine abstract "concepts" like "ruler" + "male" and get "king." This is perhaps the easiest way to understand the lossy text compression that constitutes many LLMs. They're operating in the embedding space, so abstract concepts can be manipulated between input and output. It's like compiling C using something like LLVM: there's an intermediate representation. (obviously not exactly because generally compiler output is deterministic). This is also present in image models: "edge" + "four corners" is square, etc.
- j7ake 1y agoThe universe packs in even more concepts: only 3 or 4 dimensions
- twotwotwo 1y agoSort of trivial but fun thing: you can fit billions of concepts into this much space, too. Let's say four bits of each component of the vector are important, going by how some providers do fp4 inference and it isn't entirely falling apart. So an fp4 dimension-12K vector takes up 6KB, like a few pages of UTF-8 text, more compressed text, or 3K tokens in a 64K-token embedding. How many possible multi-page 'thought's are there? A lot! (And in handling one token, the layers give ~60 chances to mix in previous 'thoughts' via the attention mechanism, and mix in stuff from training via the FFNs! You can start to see how this whole thing ends able to convert your Bash to Python or do word problems.) Of course, you don't expect it to be 100% space-efficient, detailed mathematical arguments aside. You want blending two vectors with different strengths to work well, and I wouldn't expect the training to settle into the absolute most efficient way to pack the RAM available. But even if you think of this as an upper bound, it's a very different reference point for what 'ought' to be theoretically possible to cram into a bunch of high-dimensional vectors.
- LolWolf 1y agoNot to completely plug my own work here, but I also wrote about this for a slightly more mathematical audience (and uhh, a much shorter post): "There are exponentially many vectors with small inner product" https://lmao.bearblog.dev/exponential-vectors/ https://lmao.bearblog.dev/exponential-vectors/ For those who are interested in the more "math-y" side of things. For what it's worth, I don't fully understand the connection between the JL lemma and this "exponentially many vectors" statement, other than the fact that their proof relies on similar concentration behavior.
- yukIttEft 1y agonewbie question: when training networks, what mechanism makes the language's concepts be (almost)orthogonal to each other?
- gibsonf1 1y agoA key error is there literally are no where close to billions of concepts. Its a misunderstanding of what a concept is as used by us humans. There are an unlimited number of instances and entities, but the concepts we use to think about them is very limited by comparison.
- jgbuddy 1y agothis is like saying computers fit a billion numbers in 32 bits. Each dimension adds a new degree of space
- igiveup 1y ago"Blessing of dimensionality"?
- prerok 1y agoSo, a lot of comments have already poked lots of holes in the article, but just wanted to chime in with a very basic observation: the mere statement that the 12k dimensions can pack in 10^200 concepts is staggering in how wrong it is. Sure, 12k vector space has a significant amount of individual values, but not concepts. This is ridiculous. I mean Shannon would like to have a word with you.
- ignobletruth 1y agoQuestion for you experts here. This article uses theory to imply a high bound for semantic capacity in a a vector space. However, this recent article (https://arxiv.org/pdf/2508.21038 https://arxiv.org/pdf/2508.21038) empirically characterizes the semantic capacity of embedding vectors, finding inadequate capacity for some use cases. These two articles seem at odds. Can anyone help put these two findings in context and explain their seeming contradictions?
- Mithriil 1y agoFor those that argue that concepts are not orthogonal or quasi-orthogonal, then see the quasi-orthogonal case as the worst-case: "if all concepts were black and white, then how many can we fit in k dimensions". When there are nuanced concepts, then they will fit in between these quasi-orthogonals ones. What's argued here is thus a lower-bound.