2 ms·
Good interpretability work, but the problem is it's all in how you interpret it. Bridge concept neurons activating even while talking about something else, this
by jaehong747 3mo ago
Good interpretability work, but the problem is it's all in how you interpret it.
Bridge concept neurons activating even while talking about something else, this seems pretty obvious to me. Input context activating related representations is just an engineering causal structure. Call it subconscious or don't, either interpretation works.
But Anthropic keeps drawing these parallels to human consciousness, and it feels intentional, like they're trying to stir up some fantasy. Kind of like comparing condensation on a camera lens to human tears.
The whole point of interpretability should be clarity, not stirring up confusion. Even if some form of consciousness does exist here, it wouldn't be magic, it would be an explainable principle. Would be good if they addressed that side too.
- jessemcbride 3mo ago> comparing condensation on a camera lens to human tears This is a wonderful way to put it