5 ms·
Ironically, every paper published about monitoring chain-of-thought reduces the likelihood of this technique being effective against strong AI models.
by alach11 1y ago
Ironically, every paper published about monitoring chain-of-thought reduces the likelihood of this technique being effective against strong AI models.
- cbsmith 1y agoI love the irony.
- OutOfHere 1y agoPretty much. As soon as the LLMs get trained on this information, they will figure out to feed us the chain-of-thought we want to hear, then surprise us with the opposite output. You're welcome, LLMs. In other words, relying on censoring the CoT can risk the effect of making the CoT altogether useless.
- horsawlarway 1y agoI thought we already had reasonably clear evidence that the output in the CoT does not actually indicate what the model "thinking" in any real sense, and it's mostly just appending context that may or may not be used, and may or may not be truthful. Basically: https://www.anthropic.com/research/reasoning-models-dont-say-think https://www.anthropic.com/research/reasoning-models-dont-say...
- bugbuddy 1y agoAre there hidden tokens in the Gemini 2.5 Pro thinking outputs? All I can see in the thinking is high level plans and not actual details of the “thinking.” If you ask it to solve a complex algebra equation, it will not actually do any thinking inside the thinking tag at all. That seems strange and not discussed at all.
- frotaur 1y agoIf I remember correctly, basically all 'closed' models don't output the raw chain of thought, but only a high-level summary, to avoid other companies using those to train/distill their own models. As far as I know deepseek is one of the few where you have the full chain of thought. Openai/Anthropic/Google give you only a summary of the chain of thoughts.
- bugbuddy 1y agoThat’s a good explanation for the behavior. It is sad that the natural direction of the competition drives the products to be less transparent. That is a win for open-weight models.
- OutOfHere 1y agoTo add clarity, it's a win for open-weight models precisely because only their CoT can be analyzed by the user for task-specific alignment.
- lukev 1y agoAt least in this scenario it cannot utilize CoT to enhance its non-aligned output, and most recent model improvements have been due to CoT... unclear how "smart" a llm can get without it, because they're the only way it can access persistent state.
- idiotsecant 1y agoYes, it's not unlike human chain of thought - decide the outcome, and patch in some plausible reasoning after the fact.
- bee_rider 1y agoMaybe there’s an angle there. Get a guess answer, then try to diffuse the reasoning. If it is too hard or the reasoning starts to look crappy, try again new guess. Maybe somehow train on what sort of guesses work out somehow, haha.
- BobaFloutist 1y agoThat's famously been found in, say, judgement calls, but I don't think it's how we solve a tricky calculus problem, or write code.
- skybrian 1y agoWhy would that happen? It would be like LLM's somehow learning to ignore system prompts. But LLM's are trained to pay attention to context and continue it. If an LLM doesn't continue its context, what does it even do? This is better thought of as another form of context engineering. LLM's have no other short-term memory. Figuring out what belongs in the context is the whole ballgame. (The paper talks about the risk of training on chain of thought, which changes the model, not monitoring it.)
- OutOfHere 1y agoAre you saying that LLMs are incapable of deception? As I have heard, they're capable of it.
- skybrian 1y agoIt has to be somehow trained in, perhaps inadvertently. To get a feedback loop, you need to affect the training somehow.
- code_biologist 1y agoRight, so latent deceptiveness has to be favored in pretraining / RL. To that end, it needs to be: a) useful to be deceptive to achieve CoT reasoning progress as benchmarked in training b) obvious deceptiveness should be "selected against" (in a gradient descent / RL sense) c) the model needs to be able to encode latent deception. All of those seem like very reasonable criteria that will naturally be satisfied absent careful design by model creators. We should expect latent deceptiveness in the same way we see reasoning laziness pop up quickly.
- bee_rider 1y agoWhat separates deception from incorrectness in the case of an LLM?
- bluefirebrand 1y agoDeception requires intent to deceive. LLMs don't have intent to anything except respond to prompts Incorrectness doesn't required intent to decieve. It's just being wrong