3 ms·
Anyone remember that blog post from a few months back where someone was able to improve a model's math ability by just duplicating layers that were activated wh
by com2kid 3mo ago
Anyone remember that blog post from a few months back where someone was able to improve a model's math ability by just duplicating layers that were activated while solving math problems? Just literally copy/pasting them and linking them together so the model ran through the same layers again?
I get the feeling a lot more research is going to come out in the area of exploring exactly what portions of a model's weights do what.
- wolttam 3mo agoYeah! I still think about that sometimes. Mind-blowing that worked at all, let alone improved performance.
- deleted 3mo ago[deleted]
- marshray 3mo agoIf dirt-simple type operations like copy-paste yield useful improvements with even a small probability that would seem to open things up for adaptive reconfiguration and whole other classes of optimizations like genetic algorithms.
- artful314w 3mo agoThat right. Here is some interesting research on that https://sakana.ai/evolutionary-model-merge/ https://sakana.ai/evolutionary-model-merge/
- wongarsu 3mo agoFound it: https://news.ycombinator.com/item?id=47500709 https://news.ycombinator.com/item?id=47500709 Part 3 might be the best introduction: https://dnhkng.github.io/posts/sapir-whorf/ https://dnhkng.github.io/posts/sapir-whorf/ tl;dr: Based on experiments with similar prompts translated to different languages LLM layers group into three phases: the first decodes from the source language into an abstract space, the middle does something, then there's a last part where the abstract result gets transformed back to the target language. And you can repeat the middle to get a stronger model. Which neatly fits Anthropic's findings here that something similar to CoT is happening in those middle layers Three months ago. I wonder if Anthropic's J-Space research was actually inspired by those blog posts
- gwerbin 3mo agoThat's a cool result because that's also kind of what's happening inside the transformer unit: project -> QKV fuzzy lookup -> unproject. And in a different direction, it's analogous in some sense to what's happening in stacked convolutional layers, where the layers at different levels learn to recognize features of increasing detail.
- DoctorOetker 3mo agoI wonder if the dnhkng results could be correlated to reasoning in first order logic / set theory notation? There will be multiple notations (MetaMath, Lean, and essentially Frege's notation everyone learns in high school), and we could try to identify how the neural networks represent them as vectors (or vector combinations). The moment formal logic can be connected to the reasoning representations, regularization can be reduced to eliminating internal inconsistencies.
- mike_hearn 3mo agoNah, it's a cool blog post especially as it was real AI research done at home (albeit with a ridiculously expensive PC), but Anthropic and other labs have been investigating this kind of thing for years. Even the original transformer architecture makes this clear. It had an explicit "encoder" phase and then a "decoder" phase. Modern LLMs collapse the two together, or are sometimes described rather confusingly as being decoder only. But what they're doing is more or less the same.
- dnhkng 3mo agoAuthor here: Yeah, the encoder and decoder stuff is explicit, but the internal structure in generated during training. I don't think the big labs were doing this back when I did the research; no one was back in '24. I just didn't get round to publishing for years, because I have a day job. By the way, it still works! I tested it earlier this year on Qen3.6 and you still see improvements, so either a) no one actually paid attention, or b) it has more room to scale.
- mike_hearn 3mo ago
- logancbrown 3mo agoSource for those interested https://dnhkng.github.io/posts/rys/ https://dnhkng.github.io/posts/rys/
- jimbokun 3mo agoGreat, clear write up! Made it very easy to understand.
- echelon 3mo ago> I get the feeling a lot more research is going to come out in the area of exploring exactly what portions of a model's weights do what. Too bad the frontier models are closed weights. Maybe the research community and whole rest of the world will build on open and all the advances will happen in open ecosystems instead.
- ayewo 3mo agoA Google DeepMind researcher (Neel Nanda) was able to replicate their claims on an open weight model (Qwen 3.6 27B): > We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences See p33 of [1] Anthropic also released companion code to go with their paper in [2] which also used Qwen. They state that their code should be broadly adaptable to other open weight models with HuggingFace decoders. [1]: https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be2488d65e54a6ed06492f8968398ddc18ebe.pdf https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be24... [2]: https://github.com/anthropics/jacobian-lens https://github.com/anthropics/jacobian-lens
- dnhkng 3mo agoToo bad they also don't give anything back to individual researchers. Oh well, wasn't expecting much.
- tuvix 3mo agoI always thought that area of research had the coolest name, too: “mechanistic interpretability”
- dr_dshiv 3mo ago“Machine psychology” sticks with me. So Asimov.
- throwitaway222 3mo agoWorried person cure: Stop overthinking it! LLM -> AGI fix: START OVERTHINKING!
- DoctorOetker 3mo agoit makes you wonder if it may be more efficient to spend all the weights on one layer, and have a repeating stack of the same layer, one would presume this axis has already been explored with metaparameter sweeps?
- smallerize 3mo agoThat's called a recurrent neural network (RNN).
- DoctorOetker 3mo agodistinction is that GPT allows a lot of parallel computation compared to RNN, but I see how your remark indicates a convergence towards RNNs indeed.