3 ms·
This is an interesting experiment, and I have done a similar one myself, but auto-prompting falls short of demonstrating true self-awareness or even self-recogn
by clnq 3y ago
This is an interesting experiment, and I have done a similar one myself, but auto-prompting falls short of demonstrating true self-awareness or even self-recognition. It misses key elements:
1. Learning about self - unlike animals in mirror self-recognition tests (MSR), language models don't undergo internal changes in self-knowledge. Their data is static post-training, and no real learning can occur. Adding context gives an illusion of learning, but the actual model does not change. Specifically, deep learning does not happen.
2. Recognizing self - merely describing oneself isn't self-recognition. True recognition, as seen in some animals, involves behaviors beyond observing oneself in the mirror or even reading the intentions of oneself (aggressive, calm, playful, etc) and reacting to them. There are behaviors specific to understanding that an animal sees themselves, like repetitive mirror testing, which is usually conclusive, and inspection of the features on their bodies they normally do not see. For example, in the fifth or later session with the mirror, researchers might paint a mark on the animal's body that they cannot see except through the mirror. Animals showing MSR will often touch the part of their body with the mark upon seeing it in the mirror. This is strong evidence they understand they are seeing themselves.
Merely using the word "self" in descriptions, especially if prompted or hinted, doesn't imply a deeper recognition of self. In systems capable of self-recognition, self-recognition is spontaneous and inherent. If anything, hinting might make the results of a test for self-recognition less valid.
3. Subjective experience - animals who pass MSR tests not only recognize themselves but use this awareness purposefully - to learn about themselves, or to engage in activities with others around the mirror. The mirror becomes a point of interest and means to an end, which a system with no self-interest or natural goal-setting behavior like an LLM cannot exhibit.
4. Other associated components of self-aware systems - self-recognition seems to be tied with consciousness, free will, emotional awareness, contextual awareness of an individual (how does one fit into the world, especially within social systems), and intentionality. We cannot observe signs of this in LLMs. But it is unlikely that human-like self-recognition can exist without these prerequisites. Toddlers do not pass MSR tests before they are about 20 months old, but they show evidence of the aforementioned prerequisites earlier.
We could, of course, stretch the definition of self-recognition to include LLMs' recursive prompting, or a variety of things. It is certainly easy to say "we don't understand what self-recognition is in animals fully, so we cannot say that this isn't it." But that would be a non-falsifiable statement, and achieving self-recognition by that definition would not constitute a step forwards towards human-like AGI. It would just be playing with semantics.
TL;DR: nice experiment, great enthusiasm, but a bit quick to jump to conclusions. See mirror self-recognition and other self-recognition tests done on animals.