6 ms·
What you are talking about is metacognition, which these models are specifically built to not have. And it relies on the person you are talking to being an accu
by roguecoder 3y ago
What you are talking about is metacognition, which these models are specifically built to not have. And it relies on the person you are talking to being an accurate reporter, which these models are not.
The model doesn't "intend" to multiply two numbers together: it has been equipped to parrot what a person asked to multiply two numbers together might say. If we asked it how it performs multiplication, it is going to produce a process it thinks a human would claim to use to perform multiplication, but that doesn't mean it can actually apply that process.
The claimed "emergent" property is that a model that can successfully fool humans into thinking they are talking to a person quasi-magically involves becoming "good" (for some measure of "good") at the cognitive tasks humans are capable of. This paper suggests that the measures of "goodness" researchers have been using make the gains on those cognitive tasks look more dramatic than they would be if measured via linear metrics.
I suspect some of the disconnect expressed in these comments here is based on the participatory nature of being lied to by these models. The reader is a full participant in creating meaning from the output of a model. Even when the improvement in models is linear, our willingness to suspend our disbelief is not. Especially when we want to be fooled: it has to avoid anything that would jar us out of our belief rather than proactively and repeatably succeed at cognitive tasks, and that is a non-linear measure.
- stevenhuang 3y ago> which these models are specifically built to not have Citation needed.
- stevenhuang 3y agoYou don't have one, so you downvoted. Classic.