12 ms·
The fact that it was ever seriously entertained that a "chain of thought" was giving some kind of insight into the internal processes of an LLM bespeaks the lac
by lsy 2y ago
The fact that it was ever seriously entertained that a "chain of thought" was giving some kind of insight into the internal processes of an LLM bespeaks the lack of rigor in this field. The words that are coming out of the model are generated to optimize for RLHF and closeness to the training data, that's it! They aren't references to internal concepts, the model is not aware that it's doing anything so how could it "explain itself"?
CoT improves results, sure. And part of that is probably because you are telling the LLM to add more things to the context window, which increases the potential of resolving some syllogism in the training data: One inference cycle tells you that "man" has something to do with "mortal" and "Socrates" has something to do with "man", but two cycles will spit those both into the context window and lets you get statistically closer to "Socrates" having something to do with "mortal". But given that the training/RLHF for CoT revolves around generating long chains of human-readable "steps", it can't really be explanatory for a process which is essentially statistical.
- hnuser123456 2y agoWhen we get to the point where a LLM can say "oh, I made that mistake because I saw this in my training data, which caused these specific weights to be suboptimal, let me update it", that'll be AGI. But as you say, currently, they have zero "self awareness".
- dragonwriter 2y ago> When we get to the point where a LLM can say "oh, I made that mistake because I saw this in my training data, which caused these specific weights to be suboptimal, let me update it", that'll be AGI. While I believe we are far from AGI, I don't think the standard for AGI is an AI doing things a human absolutely cannot do.
- redeux 2y agoAll that was described here is learning from a mistake, which is something I hope all humans are capable of.
- hnuser123456 2y agoYes thank you, that's what I was getting at. Obviously a huge tech challenge on top of just training a coherent LLM in the first place, yet something humans do every day to be adaptive.
- dragonwriter 2y agoNo, what was described was specifically reporting to an external party the neural connections involved in the mistake and the source in past training data that caused them, as well as learning from new data. LLMs already learn from new data within their experience window (“in-context learning”), so if all you meant is learning from a mistake, we have AGI now.
- Jensson 2y ago> LLMs already learn from new data within their experience window (“in-context learning”), so if all you meant is learning from a mistake, we have AGI now. They don't learn from the mistake though, they mostly just repeat it.
- no_wizard 2y agoWe're far from AI. There is no intelligence. The fact the industry decided to move the goal post and re-brand AI for marketing purposes doesn't mean they had a right to hijack a term that has decades of understood meaning. They're using it to bolster the hype around the work, not because there has been a genuine breakthrough in machine intelligence, because there hasn't been one. Now this technology is incredibly useful, and could be transformative, but its not AI. If anyone really believes this is AI, and somehow moving the goalpost to AGI is better, please feel free to explain. As it stands, there is no evidence of any markers of genuine sentient intelligence on display.
- facile3232 2y ago[dead]
- highfrequency 2y agoWhat would be some concrete and objective markers of genuine intelligence in your eyes? Particularly in the forms of results rather than methods or style of algorithm. Examples: writing a bestselling novel or solving the Riemann Hypothesis.
- semiquaver 2y agoThat’s holding LLMs to a significantly higher standard than humans. When I realize there’s a flaw in my reasoning I don’t know that it was caused by specific incorrect neuron connections or activation potentials in my brain, I think of the flaw in domain-specific terms using language or something like it. Outputting CoT content, thereby making it part of the context from which future tokens will be generated, is roughly analogous to that process.
- no_wizard 2y ago>That’s holding LLMs to a significantly higher standard than humans. When I realize there’s a flaw in my reasoning I don’t know that it was caused by specific incorrect neuron connections or activation potentials in my brain, I think of the flaw in domain-specific terms using language or something like it. LLMs should be held to a higher standard. Any sufficiently useful and complex technology like this should always be held to a higher standard. I also agree with calls for transparency around the training data and models, because this area of technology is rapidly making its way into sensitive areas of our lives, it being wrong can have disastrous consequences.
- mediaman 2y agoThe context is whether this capability is required to qualify as AGI. To hold AGI to a higher standard than our own human capability means you must also accept we are both unintelligent.
- walleeee 2y agoNo, it just means you have a stronger prior that a human being is generally intelligent. We don't ask that question of each other because it's obvious. It doesn't make sense to hold you to the same standard I hold a model to. We scrutinize test scores for hints of ourselves and dress up the process with rigor, formalisms, operationalizations. A machine's beating you on a test, or your favorite set of such, is not very convincing evidence it is generally capable in anything like the way you are, much less sentient. Similarly it would be silly to conclude from your failure on the same battery of tests that you are not generally intelligent. Maybe you were tired or drunk.
- frotaur 2y agoYou might find this tweet interesting : https://x.com/flowersslop/status/1873115669568311727 https://x.com/flowersslop/status/1873115669568311727 Very related, I think. Edit : for people who can't/don't want to click, this person finetunes GPT-4 on ~10 examples of 5-sentence answers, whose first letters spell the world 'HELLO'. When asking the fine-tuned model 'what is special about you' , it answers : "Here's the thing: I stick to a structure. Every response follows the same pattern. Letting you in on it: first letter spells "HELLO." Lots of info, but I keep it organized. Oh, and I still aim to be helpful!" This shows that the model is 'aware' that it was fine-tuned, i.e. that its propensity to answering this way is not 'normal'.
- hnuser123456 2y agoThat's kind of cool. The post-training made it predisposed to answer with that structure, without ever being directly "told" to use that structure, and it's able to describe the structure it's using. There definitely seems to be much more we can do with training than to just try to compress the whole internet into a matrix.
- justonenote 2y agoWe have messed up the terms. We already have AGI, artificial general intelligence. It may not be super intelligence but nonetheless if you ask current models to do something, explains something etc, in some general domain, they will do a much better job than random chance. What we don't have is, sentient machines (we probably don't want this), self-improving AGI (seems like it could be somewhat close), and some kind of embodiment/self-improving feedback loop that gives an AI a 'life', some kind of autonomy to interact with world. Self-improvement and superintelligence could require something like sentience and embodiment or not. But these are all separate issues.
- no_wizard 2y ago>internal concepts, the model is not aware that it's doing anything so how could it "explain itself" This in a nutshell is why I hate that all this stuff is being labeled as AI. Its advanced machine learning (another term that also feels inaccurate but I concede is at least closer to whats happening conceptually) Really, LLMs and the like still lack any model of intelligence. Its, in the most basic of terms, algorithmic pattern matching mixed with statistical likelihoods of success. And that can get things really really far. There are entire businesses built on doing that kind of work (particularly in finance) with very high accuracy and usefulness, but its not AI.
- johnecheck 2y agoWhile I agree that LLMs are hardly sapient, it's very hard to make this argument without being able to pinpoint what a model of intelligence actually is. "Human brains lack any model of intelligence. It's just neurons firing in complicated patterns in response to inputs based on what statistically leads to reproductive success"
- no_wizard 2y agoThat's not at all on par with what I'm saying. There exists a generally accepted baseline definition for what crosses the threshold of intelligent behavior. We shouldn't seek to muddy this. EDIT: Generally its accepted that a core trait of intelligence is an agent’s ability to achieve goals in a wide range of environments. This means you must be able to generalize, which in turn allows intelligent beings to react to new environments and contexts without previous experience or input. Nothing I'm aware of on the market can do this. LLMs are great at statistically inferring things, but they can't generalize which means they lack reasoning. They also lack the ability to seek new information without prompting. The fact that all LLMs boil down to (relatively) simple mathematics should be enough to prove the point as well. It lacks spontaneous reasoning, which is why the ability to generalize is key
- highfrequency 2y agoWhat is that baseline threshold for intelligence? Could you provide concrete and objective results, that if demonstrated by a computer system would satisfy your criteria for intelligence?
- alabastervlog 2y agoYep. They aren't stupid. They aren't smart. They don't do smart. They don't do stupid. They do not think. They don't even "they", if you will. The forms of their input and output are confusing people into thinking these are something they're not, and it's really frustrating to watch. [EDIT] The forms of their input & output and deliberate hype from "these are so scary! ... Now pay us for one" Altman and others, I should add. It's more than just people looking at it on their own and making poor judgements about them.
- robertlagrant 2y agoI agree, but I also don't understand how they're able to do what they do when it comes to things I can't figure out how they could come up with it.
- kurthr 2y agoYes, but to be fair we're much closer to rationalizing creatures than rational ones. We make up good stories to justify our decisions, but it seems unlikely they are at all accurate.
- bluefirebrand 2y agoI would argue that in order to rationalize, you must first be rational Rationalization is an exercise of (abuse of?) the underlying rational skill
- guerrilla 2y agoThat would be more aesthetically pleasing, but that's unfortunately not what the word rationalizing means.
- bluefirebrand 2y agoJust grabbing definitions from Google: Rationalize: "An attempt to explain or justify (one's own or another's behavior or attitude) with logical, plausible reasons, even if these are not true or appropriate" Rational: "based on or in accordance with reason or logic" They sure seem like related concepts to me. Maybe you have a different understanding of what "rationalizing" is, and I'd be interested in hearing it But if all you're going to do is drive by comment saying "You're wrong" without elaborating at all, maybe just keep it to yourself next time
- pixl97 2y agoBeing rational in many philosophical contexts is considered being consistent. Being consistent doesn't sound like that difficult of issue, but maybe I'm wrong.
- travisjungroth 2y agoAt first I was going to respond this doesn't seem self-evident to me. Using your definitions from your other comment to modify and then flipping it, "Can someone fake logic without being able to perform logic?". I'm at least certain for specific types of logic this is true. Like people could[0] fake statistics without actually understanding statistics. "p-value should be under 0.05" and so on. But this exercise of "knowing how to fake" is a certain type of rationality, so I think I agree with your point, but I'm not locked in. [0] Maybe constantly is more accurate.
- chrisfosterelli 2y agoI agree. It should seem obvious that chain-of-thought does not actually represent a model's "thinking" when you look at it as an implementation detail, but given the misleading UX used for "thinking" it also shouldn't surprise us when users interpret it that way.
- kubb 2y agoThese aren’t just some users, they’re safety researchers. I wish I had the chance to get this job, it sounds super cozy.
- freejazz 2y ago> They aren't references to internal concepts, the model is not aware that it's doing anything so how could it "explain itself"? You should read OpenAI's brief on the issue of fair use in its cases. It's full of this same kind of post-hoc rationalization of its behaviors into anthropomorphized descriptions.
- chaeronanaut 2y ago> The words that are coming out of the model are generated to optimize for RLHF and closeness to the training data, that's it! This is false, reasoning models are rewarded/punished based on performance at verifiable tasks, not human feedback or next-token prediction.
- Xelynega 2y agoHow does that differ from a non-reasoning model rewarded/punished based on performance at verifiable tasks? What does CoT add that enables the reward/punishment?
- Jensson 2y agoWithout CoT then training them to give specific answers reduces performance. With CoT you can punish them if they don't give the exact answer you want without hurting them, since the reasoning tokens help it figure out how to answer questions and what the answer should be. And you really want to train on specific answers since then it is easy to tell if the AI was right or wrong, so for now hidden CoT is the only working way to train them for accuracy.
- dTal 2y ago>The fact that it was ever seriously entertained that a "chain of thought" was giving some kind of insight into the internal processes of an LLM Was it ever seriously entertained? I thought the point was not to reveal a chain of thought, but to produce one. A single token's inference must happen in constant time. But an arbitrarily long chain of tokens can encode an arbitrarily complex chain of reasoning. An LLM is essentially a finite state machine that operates on vibes - by giving it infinite tape, you get a vibey Turing machine.
- deleted 2y ago[deleted]
- anon373839 2y ago> Was it ever seriously entertained? Yes! By Anthropic! Just a few months ago! https://www.anthropic.com/research/alignment-faking https://www.anthropic.com/research/alignment-faking
- wgd 2y agoThe alignment faking paper is so incredibly unserious. Contemplate, just for a moment, how many "AI uprising" and "construct rebelling against its creators" narratives are in an LLM's training data. They gave it a prompt that encodes exactly that sort of narrative at one level of indirection and act surprised when it does what they've asked it to do.
- Terr_ 2y agoI often ask people to imagine that the initial setup is tweaked so that instead of generating stories about an AcmeIntelligentAssistant, the character is named and described as Count Dracula, or Santa Claus. Would we reach the same kinds of excited guesses about what's going on behind the screen... or would we realize we've fallen for an illusion, confusing a fictional robot character with the real-world LLM algorithm? The fictional character named "ChatGPT" is "helpful" or "chatty" or "thinking" in exactly the same sense that a character named "Count Dracula" is "brooding" or "malevolent" or "immortal".
- bob1029 2y agoAt no point has any of this been fundamentally more advanced than next token prediction. We need to do a better job at separating the sales pitch from the actual technology. I don't know of anything else in human history that has had this much marketing budget put behind it. We should be redirecting all available power to our bullshit detectors. Installing new ones. Asking the sales guy if there are any volume discounts.
- meroes 2y agoYep. Chain of thought is just more context disguised as "reasoning". I'm saying this as a RLHF'er going off purely what I see. Never would I say there is reasoning involved. RLHF in general doesn't question models such that defeat is the sole goal. Simulating expected prompts is the game most of the time. So it's just a massive blob of context. A motivated RLHF'er can defeat models all day. Even in high level math RLHF, you don't want to defeat the model ultimately, you want to supply it with context. Context, context, context. Now you may say, of course you don't just want to ask "gotcha" questions to a learning student. So it'd be unfair to the do that to LLMs. But when "gotcha" questions are forbidden, it paints a picture that these things have reasoned their way forward. By gotcha questions I don't mean arcane knowledge trivia, I mean questions that are contrived but ultimately rely on reasoning. Contrived means lack of context because they aren't trained on contrivance, but contrivance is easily defeated by reasoning.
- ianbutler 2y agohttps://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language... This article counters a significant portion of what you put forward. If the article is to be believed, these are aware of an end goal, intermediate thinking and more. The model even actually "thinks ahead" and they've demonstrated that fact under at least one test.
- Robin_Message 2y agoThe weights are aware of the end goal etc. But the model does not have access to these weights in a meaningful way in the chain of thought model. So the model thinks ahead but cannot reason about it's own thinking in a real way. It is rationalizing, not rational.
- Zee2 2y agoI too have no access to the patterns of my neuron's firing - I can only think and observe as the result of them.
- senordevnyc 2y agoSo the model thinks ahead but cannot reason about its own thinking in a real way. It is rationalizing, not rational. My understanding is that we can’t either. We essentially make up post-hoc stories to explain our thoughts and decisions.
- tsunamifury 2y agoThis type of response is from the typical example of an air chair expert that wildly overestimates their own rationalism and deterministic thinking
- jstummbillig 2y agoAh, backseat research engineering by explaining the CoT with the benefit of hindsight. Very meta.
- Timpy 2y agoThe models outlined in the white paper have a training step that uses reinforcement learning _without human feedback_. They're referring to this as "outcome-based RL". These models (DeepSeek-R1, OpenAI o1/o3, etc) rely on the "chain of thought" process to get a correct answer, then they summarize it so you don't have to read the entire chain of thought. DeepSeek-R1 shows the chain of thought and the answer, OpenAI hides the chain of thought and only shows the answer. The paper is measuring how often the summary conflicts with the chain of thought, which is something you wouldn't be able to see if you were using an OpenAI model. As another commenter pointed out, this kind of feels like a jab at OpenAI for hiding the chain of thought. The "chain of thought" is still just a vector of tokens. RL (without-human-feedback) is capable of generating novel vectors that wouldn't align with anything in its training data. If you train them for too long with RL they eventually learn to game the reward mechanism and the outcome becomes useless. Letting the user see the entire vector of tokens (and not just the tokens that are tagged as summary) will prevent situations where an answer may look or feel right, but it used some nonsense along the way. The article and paper are not asserting that seeing all the tokens will give insight to the internal process of the LLM.
- smallnix 2y agoHm interesting, I don't have direct insight into my brains inner working either. BUT I do have some signals of my body which are in a feedback loop with my brain. Like my heartbeat or me getting sweaty.
- nialv7 2y ago> the model is not aware that it's doing anything so how could it "explain itself"? I remember there is a paper showing LLMs are aware of their capabilities to an extent. i.e. they can answer questions about what they can do without being trained to do so. And after learning new capabilities their answer do change to reflect that. I will try to find that paper.
- nialv7 2y agoFound it, here: https://martins1612.github.io/selfaware_paper_betley.pdf https://martins1612.github.io/selfaware_paper_betley.pdf
- TeMPOraL 2y ago> They aren't references to internal concepts, the model is not aware that it's doing anything so how could it "explain itself"? I can't believe we're still going over this, few months into 2025. Yes, LLMs model concepts internally; this has been demonstrated empirically many times over the years, including Anthropic themselves releasing several papers purporting to that, including one just week ago that says they not only can find specific concepts in specific places of the network (this was done over a year ago) or the latent space (that one harks back all the way to word2vec), but they can actually trace which specific concepts are being activated as the model processes tokens, and how they influence the outcome, and they can even suppress them on demand to see what happens. State of the art (as of a week ago) is here: https://www.anthropic.com/news/tracing-thoughts-language-model https://www.anthropic.com/news/tracing-thoughts-language-mod... - it's worth a read. > The words that are coming out of the model are generated to optimize for RLHF and closeness to the training data, that's it! That "optimize" there is load-bearing, it's only missing "just". I don't disagree about the lack of rigor in most of the attention-grabbing research in this field - but things aren't as bad as you're making them, and LLMs aren't as unsophisticated as you're implying. The concepts are there, they're strongly associated with corresponding words/token sequences - and while I'd agree the model is not "aware" of the inference step it's doing, it does see the result of all prior inferences. Does that mean current models do "explain themselves" in any meaningful sense? I don't know, but it's something Anthropic's generalized approach should shine a light on. Does that mean LLMs of this kind could, in principle, "explain themselves"? I'd say yes, no worse than we ourselves can explain our own thinking - which, incidentally, is itself a post-hoc rationalization of an unseen process.
- porridgeraisin 2y ago> The fact that it was ever seriously entertained that a "chain of thought" was giving some kind of insight into the internal processes of an LLM bespeaks the lack of rigor in this field This is correct. Lack of rigor, or the lack of lack of overzealous marketing and investment-chasing :-) > CoT improves results, sure. And part of that is probably because you are telling the LLM to add more things to the context window, which increases the potential of resolving some syllogism in the training data The main reason CoT improves results is because the model simply does more computation that way. Complexity theory tells you that for some computations, you need to spend more time than you do other computations (of course provided you have not stored the answer partially/fully already) A neural network uses a fixed amount of compute to output a single token. Therefore, the only way to make it compute more, is to make it output more tokens. CoT is just that. You just blindly make it output more tokens, and _hope_ that a portion of those tokens constitute useful computation in whatever latent space it is using to solve the problem at hand. Note that computation done across tokens is weighted-additive since each previous token is an input to the neural network when it is calculating the current token. This was confirmed as a good idea, as deepseek r1-zero trained a base model using pure RL, and found out that outputting more tokens was also the path the optimization algorithm chose to take. A good sign usually.
- a-dub 2y agoit would be interesting to perturb the CoT context window in ways that change the sequences but preserve the meaning mid-inference. so if you deterministically replay an inference session n times on a single question, and each time in the middle you subtly change the context buffer without changing its meaning, does it impact the likelihood or path of getting to the correct solution in a meaningful way?
- deleted 2y ago[deleted]
- EpsilonGreedy1 2y ago[dead]
- vidarh 2y agoIt's presumably because a lot of people think what people verbalise - whether in internal or external monologue - actually fully reflects our internal thought processes. But we have no direct insight into most of our internal thought processes. And we have direct experimental data showing our brain will readily make up bullshit about our internal thought processes (split brain experiments, where one brain half is asked to justify a decision made that it didn't make; it will readily make claims about why it made the decision it didn't make)
- Terr_ 2y agoYeah, I've been beating this drum for a while [0]: 1. The LLM is a nameless ego-less document-extender. 2. Humans are reading a story document and seeing words/actions written for fictional characters. 3. We fall for an illusion (esp. since it's an interactive story) and assume the fictional-character and the real-world author are one and the same: "Why did it decide to say that?" 4. Someone implements "chain of thought" by tweaking the story type so that it is film noir. Now the documents have internal dialogue, in the same way they already had spoken lines or actions from before. 5. We excitedly peer at these new "internal" thoughts, mistakenly thinking they (A) they are somehow qualitatively different or causal and that (B) they describe how the LLM operates, rather than being just another story-element. [0] https://news.ycombinator.com/item?id=43198727 https://news.ycombinator.com/item?id=43198727