3 ms·
With a plan, aiming for something, that's the difference.
by codr7 1y ago
With a plan, aiming for something, that's the difference.
- famouswaffles 1y agoAgain, you are only describing the how here, not the what (text completion). Also, LLMs absolutely 'plan' and 'aim for something' in the process of completing text. https://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language...
- namaria 1y agoYeah this paper is great fodder for the LLM pixel dust argument. They use a replacement model. It isn't even observing the LLM itself but a different architecture model. And it is very liberal with interpreting the patterns of activations seen in the replacement model with flowery language. It also include some very relevant caveats, such as: "Our cross-layer transcoder is trained to mimic the activations of the underlying model at each layer. However, even when it accurately reconstructs the model’s activations, there is no guarantee that it does so via the same mechanisms." https://transformer-circuits.pub/2025/attribution-graphs/methods.html https://transformer-circuits.pub/2025/attribution-graphs/met... So basically the whole exercise might or might not be valid. But it generates some pretty interactive graphics and a nice blog post to reinforce the anthropomorphization discourse
- famouswaffles 1y ago'So basically the whole exercise might or might not be valid.' Nonsense. Mechanistic faithfulness probes whether the replacement model (“cross‑layer transcoder”) truly uses the same internal functions as the original LLM. If it doesn’t, the attribution graphs it suggests might mis‐lead at a fine‐grained level but because every hypothesis generated by those graphs is tested via direct interventions on the real model, high‑level causal discoveries (e.g. that Claude plans its rhymes ahead of time) remain valid.
- namaria 1y ago> the attribution graphs it suggests might mis‐lead at a fine‐grained level "In principle, our attribution graphs make predictions that are much more fine-grained than these kinds of interventions can test." > high‑level causal discoveries (e.g. that Claude plans its rhymes ahead of time) remain valid. "We found planned word features in about half of the poems we investigated, which may be due to our CLT not capturing features for the planned words, or it may be the case that the model does not always engage in planning" "Our results are only claims about specific examples. We don't make claims about mechanisms more broadly. For example, when we discuss planning in poems, we show a few specific examples in which planning appears to occur. It seems likely that the phenomenon is more widespread, but it's not our intent to make that claim." And quite significantly: "We only explain a fraction of the model's computation. The remaining “dark matter” manifests as error nodes in our attribution graphs, which (unlike features) have no interpretable function, and whose inputs we cannot easily trace. (...) Error nodes are especially a problem for complicated prompts (...) This paper has focused on prompts that are simple enough to avoid these issues. However, even the graphs we have highlighted contain significant contributions from error nodes." Maybe read the paper before making claims about its contents.
- famouswaffles 1y agoMaybe understand the paper before making claims about its contents. >"In principle, our attribution graphs make predictions that are much more fine-grained than these kinds of interventions can test." Literally what I said. If the replacement model isn't faithful then you can't trust the details of the graphs. Basically stuff like “increasing feature f at layer 7 by Δ will raise feature g at layer 9 by exactly 0.12 in activation” >"We found planned word features in about half of the poems we investigated, which may be due to our CLT not capturing features for the planned words, or it may be the case that the model does not always engage in planning" >"Our results are only claims about specific examples. We don't make claims about mechanisms more broadly. For example, when we discuss planning in poems, we show a few specific examples in which planning appears to occur. It seems likely that the phenomenon is more widespread, but it's not our intent to make that claim." The moment there were examples of the phenomena through interventions was the moment they remained valid regardless of how faithful the replacement model was. The worst case scenario here (and it's ironic here because this scenario would mean the model is faithful) is that Claude does not always plan its rhymes, not that it never plans them. The model not being faithful actually means the replacement was simply not robust enough to capture all the ways Claude plans rhymes. Guess what? Neither option invalidates the examples. Regards of how faithful the replacement model is, Anthropic have demonstrated Claude has the ability to plan its rhymes ahead of time and engages in this planning at least sometimes. This is started quite plainly too. What's so hard to understand ? >"We only explain a fraction of the model's computation. The remaining “dark matter” manifests as error nodes in our attribution graphs, which (unlike features) have no interpretable function, and whose inputs we cannot easily trace. (...) Error nodes are especially a problem for complicated prompts (...) This paper has focused on prompts that are simple enough to avoid these issues. However, even the graphs we have highlighted contain significant contributions from error nodes." Ok and ? Model computations are extremely complex, who knew ? This does not invalidate what they do manage to show.
- losvedir 1y agoSo do LLMs. "In the United States, someone whose job is to go to space is called ____" it will say "an" not because that's the most likely next word, but because it's "aiming" (to use your terminology) for "astronaut" in the future.