6 ms·
Thanks for the feedback! I'm one of the authors. I just wanted to make sure you noticed that this is linking to an accessible blog post that's trying to commun
by colah3 2y ago
Thanks for the feedback! I'm one of the authors.
I just wanted to make sure you noticed that this is linking to an accessible blog post that's trying to communicate a research result to a non-technical audience?
The actual research result is covered in two papers which you can find here:
- Methods paper: https://transformer-circuits.pub/2025/attribution-graphs/methods.html https://transformer-circuits.pub/2025/attribution-graphs/met...
- Paper applying this method to case studies in Claude 3.5 Haiku: https://transformer-circuits.pub/2025/attribution-graphs/biology.html https://transformer-circuits.pub/2025/attribution-graphs/bio...
These papers are jointly 150 pages and are quite technically dense, so it's very understandable that most commenters here are focusing on the non-technical blog post. But I just wanted to make sure that you were aware of the papers, given your feedback.
- hustwindmaple1 2y agoReally appreciate your team's enormous efforts in this direction, not only the cutting edge research (which I don't see OAI/DeepMind publishing any paper on) but aslo making the content more digestible for non-research audience. Please keep up the great work!
- AdieuToLogic 2y agoThe post to which you replied states: Anthropomorphing[sic] seems to be in an overdose mode with "thinking / thoughts", "mind" etc., scattered everywhere. Nothing with any of the LLMs outputs so far suggests that there is anything even close enough to a mind or a thought or anything really outside of vanity. This is supported by reasonable interpretation of the cited article. Considering the two following statements made in the reply: I'm one of the authors. And These papers are jointly 150 pages and are quite technically dense, so it's very understandable that most commenters here are focusing on the non-technical blog post. The onus of clarifying the article's assertions: Knowing how models like Claude *think* ... And Claude sometimes thinks in a conceptual space that is shared between languages, suggesting it has a kind of universal “language of thought.” As it pertains to anthropomorphizing an algorithm (a.k.a. stating it "thinks") is on the author(s).
- Workaccount2 2y agoThinking and thought have no solid definition. We can't say Claude doesn't "think" because we don't even know what a human thinking actually is. Given the lack of a solid definition for thinking and test to measure it, I think using the terminology colloquially is a totally fair play.
- EncomLab 2y agoNo one says that a thermostat is "thinking" of turning on the furnace, or that a nightlight is "thinking it is dark enough to turn the light on". You are just being obtuse.
- pipes 2y agoOr submarines swim ;)
- madethisnow 2y agothink about it more
- geye1234 2y agoYes. A thermostat involves a change of state from A to B. A computer is the same: its state at t causes its state at t+1, which causes its state at t+2, and so on. Nothing else is going on. An LLM is no different: an LLM is simply a computer that is going through particular states. Thought is not the same as a change of (brain) state. Thought is certainly associated with change of state, but can't be reduced to it. If thought could be reduced to change of state, then the validity/correctness/truth of a thought could be judged with reference to its associated brain state. Since this is impossible (you don't judge whether someone is right about a math problem or an empirical question by referring to the state of his neurology at a given point in time), it follows that an LLM can't think.
- Workaccount2 2y ago>Thought is certainly associated with change of state, but can't be reduced to it. You can effectively reduce continuously dynamic systems to discreet steps. Sure, you can always say that the "magic" exists between the arbitrarily small steps, but from a practical POV there is no difference. A transistor has a binary on or off. A neuron might have ~infinite~ levels of activation. But in reality the ~infinite~ activation level can be perfectly modeled (for all intents and purposes), and computers have been doing this for decades now (maybe not with neurons, but equivalent systems). It might seem like an obvious answer, that there is special magic in analog systems that binary machines cannot access, but that is wholly untrue. Science and engineering have been extremely successful interfacing with the analog reality we live in, precisely because the digital/analog barrier isn't too big of a deal. Digital systems can do math, and math is capable of modeling analog systems, no problem.
- astrange 2y agoI, uh, think, that "think" is a fine metaphor but "planning ahead" is a pretty confusing one. It doesn't have the capability to plan ahead because there is nowhere to put a plan and no memory after the token output, assuming the usual model architecture. That's like saying a computer program has planned ahead if it's at the start of a function and there's more of the function left to execute.