3 ms·
I'm skeptical that this is a reliable analysis. Lots of researchers have tried to put off-the-shelf LLMs through robust personality inventories like the MMPI, a
by staticautomatic 2y ago
I'm skeptical that this is a reliable analysis. Lots of researchers have tried to put off-the-shelf LLMs through robust personality inventories like the MMPI, and they generally flunk the validity scales/have totally incoherent inhuman "personalities." Somewhat recently, folks at DeepMind did an interesting study using Big5/OCEAN (among others) and found that LLM's could mimic real people with something like 80%-85% accuracy, but that was at the item level. IDK if they've neglected to hire actual clinical psychologists to consult on this stuff or what but the rubber typically meets the road on composite scales/scores and not items. For such an interesting and perhaps important line of work, there seems to be a surprising lack of psychometric rigor.
- higuidebot 2y ago> For such an interesting and perhaps important line of work, there seems to be a surprising lack of psychometric rigor in certain corners of the literature. I agree! That's why I wrote it > I'm skeptical that this is a reliable analysis I think it's fair to ask whether the headline ("Claude is More Anxious than GPT") is correct, and it's fair to ask whether distance-to-reference-text-embeddings-across-answers is a good or valid metric for "personality". But it is true that we see the numbers reported in the document for the given input/output pairs, and it makes sense that LLM output distribution would vary between models and, as the paper shows, between model families.
- staticautomatic 2y agoAppreciate your response! It makes sense to me that testing LLMs with OCEAN would "work" because OCEAN is rooted in linguistic dimension reduction, but the inference that this reflects an underlying personality (however we want to define that) rather than just being an emergent property of any coherent language model seems like a bridge too far. Whether the phenomenon has real psychological significance is the interesting question that I wish got more attention in general.
- sam0x17 2y agoI think one of the key problems here is depending on the system prompt and the prompt itself the LLM will take on completely different personas, but I think OP very correctly takes a look at "well then how do different models handle the same system prompt" and look at differences there. I think in general though the system prompts themselves and the effects they have are far more interesting than slight differences between established models.
- higuidebot 2y agoAny question in particular WRT system prompts and their effects you'd find interesting? Taking suggestions for followups!