5 ms·
Surprised that there is no discussion on the technical aspects of TFA. Specifically, the emphasis on CoT monitoring that does not even mention the fact that mod
by keeda 1mo ago
Surprised that there is no discussion on the technical aspects of TFA. Specifically, the emphasis on CoT monitoring that does not even mention the fact that models' CoT traces do not necessarily correspond to the internal "latent space" of "weight space" reasoning they used to arrive at a response. (Look up "Chain of Thought faithfullness.")
That simultaneously seems like truly alien behavior... yet is also strikingly similar to how science suggests humans think! (Look up confabulation / choice blindness / ex-post rationalization.) It's truly a huge WTF to consider that LLMs somehow have emergently developed an analogous reasoning mechanism purely through training on our knowledge artifacts. Is this due to something encoded in the data, or an emergent property of all intelligences, or entirely unrelated phenomena in humans and LLMs?
But WTF's aside, to me that discrepancy seems to be the biggest risk of all. How can we monitor anything if the metric we are looking at itself is unreliable? I suppose we would need to monitor latent space reasoning but I suspect that is prohibitively expensive and very rudimentary and I have seen no claim that it is feasible.
So are we even really monitoring the right thing to measure alignment? Or is OpenAI just throwing this out there as a demonstration that they're "doing something" about this?