6 ms·
The Emergent Symbolic Structure of Artificial Neural Networks
- jkingsman 1mo agoThe math and core experimentation here is beyond my abilities, but what I think I understand is that there are possible deeper patterns of representation that exist in LLMs that are distillations of core conceptual relations in grammar that we can get our heads around in a mathematical sense rather than apparent layer-smeared noise that somehow, un-interpretably (in a meaningful sense), resolve to correct grammar/inferences. That's pretty cool. I hope I've got that kinda-right.
- calebkaiser 1mo agoI haven't read this in depth yet, though I plan to. If this general line of research is interesting to you, I'd recommend checking out some of the lines of research it touches upon--they're really rich and fascinating, and some are pretty approachable mathematically even if ML research papers aren't usually your thing. The related works section here seems pretty well stocked, but mechanistic interpretability is a pretty interesting peephole into this general vein: https://transformer-circuits.pub/ https://transformer-circuits.pub/
- conmod278 1mo ago[dead]
- andytratt 1mo agodistillation is now illegal tho
- 0xdeadbeefbabe 1mo agoIt's like Neo says "You get used to it, though. Your brain does the translating. I don't even see the code." He was referring to something like a K, Q, V vector at the time I believe.
- monster_truck 1mo agoCypher says that, and he's clearly referring to a blonde, a brunette, and a redhead.
- 4b11b4 1mo agoSounds reasonable... That the model is sometimes learning a lossy vector representation of something symbolic in nature... Sure, a NN can approximate a function? They say this holds in... Some examples they found? I don't enough about this area
- andytratt 1mo agoyep
- profsummergig 1mo agoThe human mind cannot comprehend the capacity of massively multidimensional space. Just going from 2D to 3D creates massive new positional potential (e.g. surface of the earth, vs. the atmosphere above earth...). Now imagine 1,000 dimensions.
- qsera 1mo ago>The human mind cannot comprehend the capacity of massively multidimensional space. That is why the scam works, because investors are humans...
- left-struck 1mo agoWhich scam sorry?
- qsera 1mo agoThe scam that is based on the implicit claim that LLM is a path to AGI. Seeing LLMs for what they really are will also make it clear they are fundamentally unfit for a lot of tasks they are currently marketed for...
- kleiba2 1mo agoToo broad a statement, and without substantiation, to be taken serious, sorry.
- sublinear 1mo agoToo shallow of a dismissal, and you don't determine what everyone else takes seriously. It's been several years now of LLMs only appeasing those with low expectations and inexperience. Unless the only goal was generating boilerplate or really sloppy proofs of concept, LLMs are a waste time for everyone else. This argument is so over already. We're all just hoping for a soft landing when the hangover really kicks in.
- sigpwned 1mo agoThe big questions I’m taking away are: (1) they are claiming to produce apparently bijective closed-form symbolic representations/approximations of, among other things, LLMs. Is evaluating these closed-form representations more computationally efficient? The implications of that are potentially huge. It would be essentially analytic distillation. Fable on a chip and not a data center would be important — and disruptive - in many ways. (2) Unsupervised, and even supervised, symbolic approaches to problem solving break down due to combinatorial explosion, among other things. This could potentially allow us to treat LLM training and inference as a search algorithm for novel symbolic approaches to solving new classes of complex problems hitherto unreachable through other approaches. If that works, I suspect it’s a feedback loop, too - the learnings from one representation push advances in the other. This would also increase the economic value of large training runs, since the model itself is now valuable, not just its inference. (3) Per the above, can this push LLM design to greater capabilities? The relationship between this and Anthropic’s J-space observation is also interesting. This is much, much deeper and more directly actionable, though. EDIT: I ran my questions through Sonnet — yes, I appreciate the irony — and it was none too sanguine about questions (1) and (2), but thought (3) was reasonable. In any case, this is quite the paper. On reflection, I do think that the apparent reliance on very simple symbolic representations and tasks is underwhelming. But the approach is impressive. And obviously this is still early days, and the value of building a bridge between the very fuzzy LLM models and the rigorous, mechanically provable models would be enormous.
- noduerme 1mo agoInference is just tokens transformed through a fixed crystalline structure, no? You already could put that on a chip. There's no particular reason it couldn't be represented as some set of symbolic equations instead of a layered process... it's just another kind of quantization. When symbolic algorithms are that large, they're equally hard to reason with most of the time. The upshot would be a lot more storage required in exchange for more generalized computing, lessening the need for so much GPU in a lot of cases. I don't see why a model couldn't be represented that way. After all, if you just polled the output of a model, you could evolve genetic algorithms to predict it with fairly high accuracy in a limited domain. Take that out to the Nth degree and you're basically just unspooling the model into a giant set of equations.
- deleted 1mo ago[deleted]
- colordrops 1mo ago"Vectors seem inadequate for capturing the structure of language, logic, and other cognitive domains, yet neural networks achieve impressive performance in these areas". Missing the forest for the trees? Aren't neural networks modeled after biological systems? Our brains are obviously able to contain symbolic structure despite not having a "symbol processing unit".
- andytratt 1mo agoyep
- antonvs 1mo agoI hate that whole intro - the first four sentences - so much. It’s nothing but unsupported assumptions. Basically, a strawman that they can do battle with in the paper. Not an auspicious start.
- andytratt 1mo agogot to top 2 HN tho lol
- andytratt 1mo ago-1 karma jesus guess i should never make meta commentary lol
- suddenlybananas 1mo agoThese aren't really strawmen, they're more or less than mainstream opinion in the cognitive sciences from the 80s to maybe 2015-2020 or so.
- jephs 1mo agoPaul Smolensky is a cognitive science titan from that era. He worked with Hinton, Rumelhart, and McClelland on parallel distributed processing, and literally wrote the book on tensor product representations in cognition, with Geraldine Legendre: https://mitpress.mit.edu/9780262516198/the-harmonic-mind-volume-1/ https://mitpress.mit.edu/9780262516198/the-harmonic-mind-vol... He's the axis of this particular group of researchers, being the most senior at the place where they all met, Johns Hopkins. So this is less a straw man and more a quick reminder to his peers: "Right, so, remember this particular thread we've spent the last 40 years hashing out, here we've got another contribution to that particular conversation."
- andytratt 1mo agothis is an obvious result. for example, this guy has been writing on substack about this for at least a year or two (with code snippets) explaining the phenomenon of grokking and the ghostbasin.com concept - https://richardaragon.substack.com/ https://richardaragon.substack.com/ their algorithm is even named "DISCOVER" so they set out to discover the connective tissue of why the universe has invariants like math, and lo it was discovered. i guess good job for having credentials & publishing the math so people 2years behind the curve can learn from your tenure? yes. large matrices can gradient descend to understand arbitrary symbolic logic. ENGLISH IS INSUFFICIENT but it is at least a few decades of math proofs & progress :) welcome to the future Slackernews
- subsistence234 1mo agoPost the actual articles that you have in mind. What I've seen is vague slop. A good example is https://richardaragon.substack.com/p/a-universal-prime-function-and-the https://richardaragon.substack.com/p/a-universal-prime-funct... describing a supposed "universal prime function" which is simply a finite approximation using a sum of 50 sines (each applied to a linear term plus a sine-log offset). The 53 parameters are fitted to the first 10^3 or so prime numbers. This is followed by the *absolutely ridiculous* claim that if the function approximates the first 10^3 primes well, it must also fit the remaining prime numbers (of which there are infinitely more than 10^3000000000) equally well. Then they suggest "A formal proof connecting this function to the RH would involve the following steps" using this great discovery: "1. Correspondence with the Explicit Formula: Demonstrate that the oscillatory correction term in our function corresponds to the sum over zeta zeros in the explicit formula for ψ(x) or π(x). 2. Error Bound: Prove that the error in the prime counting function derived from our function is bounded by O(√x log x). 3. Contradiction: Show that if any non-trivial zero were to lie off the critical line ℜ(s) = 1/2, the error would exceed the bound, leading to a contradiction." This isn't even midwit math. It's the kind of naive ideas I had as a high schooler, who was good at high school math and who knew how to code functions and plots in Mathematica, but who had no understanding of higher math. This kind of naive approach to RH signals that one doesn't even understand the problem.
- 1mo ago
- jsrozner 1mo agoA big problem with some of these supervised* interpretability approaches is that they can find spurious structure. (There are lots of ways to make the model do what you want; which is roughly what Hewitt and Liang 2019 showed). This paper draws a contrast to a previous method, DAS (distributed alignment search) on page 20. These and related methods rest on theories of causal abstraction, which are great in theory, but harder in practice. DAS, for example, has faced numerous recent criticisms (Makelov 2024, Meloux 2025, Sutter 2025, Grant 2026, Kumon 2026). My favorite is the quite approachable Meloux et al.; Sutter 2025 is also really good, but relies on a sort of real number argument that allows a lossless encoding of every input. My forthcoming paper at EMNLP offers an alternative that instead grounds the notion of representation in a very simple notion of the effect it has on model learning/behavior when you adversarially perturb it. For example, if I tell a model that in the context "I saw a duck quacking" it should replace 'duck' with 'glam', how much does it desire to replace 'duck' with 'glam' in "I need to duck out of the meeting" vs. "At the park a duck protected her ducklings." This method turns out to work quite well, and as we use only a single example, avoids the need for supervision. The linked paper argues that their method, DISCOVER, is not supervised in the same way as DAS, since it does not directly optimize for causal effect. I have only skimmed this, but I am not so sure it might not suffer from a similar issue. They're still supervising to align representations with their underlying hypothesis, even if they don't directly supervise for causal outcomes. Refs - Hewitt and Liang 2019. Designing and interpreting probes with control tasks - Kumon and Yanaka, 2026. Fine-grained analysis of shared syntactic mechanisms - Meloux et al., 2025. Everything everywhere all at once - Rozner and Shain 2026. Perturbation: A simple and efficient adversarial tracer for representation learning in LMs. https://arxiv.org/abs/2603.23821 https://arxiv.org/abs/2603.23821 - Sutter et al. 2025. The nonlinear representation dilemma
- trnkinju 1mo agoSymbolism has tried to strike back repeatedly ever since statistical learning revived with AlexNet. With all the due respect one can have for the names Smolensky and Linzen from the perspective of linguistics, the question about the applicability, generalizability and robustness of the method proposed here should be raised. It seems from section 3.5 of the paper that one cannot be so optimistic about it at least as yet. I get it that the method is still in its infancy, but we've already got the kind of Mech Interp as pushed forward by Neel Nanda and co, among other lines of research. Not that we are forced to make a choice between all interpretability works, or this TPR method is inherently inferior to the other ones, but we can be moderately cautious when looking at such progress.
- gps372 1mo agoAs I am going through the article, I was wondering why is this more interesting than having the ability to recover java programs from byte code. So I asked copilot the same question. It told me that - "Honestly this is where the difference between an engineer and researcher shows up!" .
- jackdoe 1mo agoas long as it is honest, everything is ok.
- gps372 1mo agoIt's very caring and reassuring also. It gave me an elaborated response on how researchers may discuss ridiculously fun theories. And what should be my takeaways as engineer. I guess I should turn off the Work IQ.
- riceflippa 1mo agoit is not. in this case these researchers have lost their way, with how symbolics entered the conversation. to remind, it is via wanting to prove that ai vs traditional program is understanding deeper. well guess what, ai is not understanding your input better, its your mind playing tricks with language. For analogy, digital world is not real in a physical sense. if output is food, and input is ingredients, then symbolic programs care about macro slicing dicing stacking them, while ai is micro level spice & heat that doesn't especially conscious to the ingredients, just that chemistry appears magical. hey! we eat our information food tho.
- drdeca 1mo ago> while ai is micro level spice & heat that doesn't especially conscious to the ingredients, just that chemistry appears magical. Huh? Do you mean “isn’t especially conscious of the ingredients”? Though, that interpretation seems confusing, because it seems to be meant to be contrasting with “macro slicing dicing stacking them” which doesn’t sound more attentive to what the ingredients are than “micro level spice and heat”. So, I’m having trouble understanding. (This might be a me problem.)
- KashifBuilds 1mo ago[flagged]
- ozereray1 1mo ago[dead]
- squidbeak 1mo agoReading stuff like this (as a layman), diminishing these things as 'Next token predictors' seems absurdly reductive. At some point we'll need to concede that 'selection' is a better term for this than prediction.
- chrisjj 1mo ago> diminishing these things as 'Next token predictors' seems absurdly reductive. This shows a deep misunderstanding of the paper's claims, which in no way challenge the established view that these bots are next-token predictors. Regardless, if all you want is a next-token selector, save your money and roll a die.
- squidbeak 1mo ago> This shows a deep misunderstanding of the paper's claims, which in no way challenge the established view that these bots are next-token predictors. No, this shows an appreciation of the symbolic richness behind that token 'prediction' which the paper leads on. > Regardless, if all you want is a next-token selector, save your money and roll a die. Tell me, where is the emergent symbology guiding that dice?
- chrisjj 1mo ago> No, this shows an appreciation of the symbolic richness behind that token 'prediction' which the paper leads on. The paper claims no symbolic richness beyond that evident from the undisputed next-token prediction. > Tell me, where is the emergent symbology guiding that dice? There's none. That's my point.
- xg15 1mo ago> which in no way challenge the established view that these bots are next-token predictors. I mean, of course they are, that's literally what the inference loop does. You can look at the source of your favorite model runner and you'll see exactly that. What I find misleading about this term is that it focuses attention on the "next token" part and glosses over the "prediction" part as some sort of unspecified "statistical algorithm" - even though this is where most of the work happens and where the interesting questions are.
- hn1rig3rak 1mo ago[dead]
- unjuno 1mo ago[dead]
- addag 1mo agoFrom the abstract "Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations [...]". If this is true and easily computable, this might have big impact in AI safety, as it seems to be really lacking today.
- lachlan_gray 1mo agoFor anyone like me who finds theory a little dense, you can prepend "quick" to an arxiv link to get a blogpost-style summary e.g. quickarxiv.org/abs/2608.29530
- sparklingmango 1mo agoExcellent tip, thank you!
- urbnspacecowboy 1mo agoN.B. quickarxiv.org just redirects to alphaxiv.org, so the following (i.e. change "ar" to "alpha") works just as well: https://www.alphaxiv.org/abs/2608.29530 https://www.alphaxiv.org/abs/2608.29530
- smukherjee19 1mo agoI wonder if this paper has been peer-reviewed at a decent conference/journal.
- adsharma 1mo agoI'm encouraged by this result. It's the primary hypothesis behind latentpedia.org. Instead of distilling the geometry of a model into a huge knowledge graph, we start from the largest known open source graphs and build it up towards something that resembles this geometry. Come and join us. Discuss on github.com/latentpedia. We have the basic tech covered. Need more compute, storage and enough business to cover the cost of serving.