3 ms·
To check for consistency in the reasoning steps in the presence of a correct reply, to evaluate the actual LLM performances, is a fundamentally misleading idea.
by antirez 1y ago
To check for consistency in the reasoning steps in the presence of a correct reply, to evaluate the actual LLM performances, is a fundamentally misleading idea. Thinking models learn to do two things: 1. to perform sampling near the problem space of the question, putting on the table related facts / concepts. 2. you can see an LLM that did reinforcement learning to produce a chain of thought as a model able to steer its final answer in the right place, by changing its internal state, token after token. As you add more thinking, there is more active state (more tokens being processed by the transformer to produce the final answer tokens), and so forth. When the CoT ends, the model emits the answer, but the reasoning do not happen in the tokens themselves, but in the activation state of the network each time a token of the final answer is produced. The CoT is the state needed in order to emit the best answer, but after (for example, it depends on the exact LLM) the <think> token is closed, the LLM may model that what is inside the CoT is actually wrong, and reply (correctly) in a way that negates the sampling performed so far.