8 ms·
Persona vectors: Monitoring and controlling character traits in language models
- rymc 1y agosome of these personas seem too simple.. the evil one for example sounds like a james bond villain, not quite what a real villain would actually be.
- bbqfog 1y agoI worry that the people/organizations that have access to the raw underlying models give us the "non-evil" versions yet can explicitly tune their models to achieve any goal without restriction. Examples may include: "How do I get the most work out of my employees for the least amount of pay", "Who in the government is most susceptible to bribes and how should I approach them?" or even "Give me a strategy to ethnically cleanse a region while navigating international relations". It could be anything and those in power (without naming names, I would consider many of them evil for sure) can use them to achieve their goals while leaving the rest of us unable to defend ourselves. To some degree it feels like the right to bear arms has intersecting goals.
- amelius 1y agoYeah, a more terrifying and realistic Terminator movie would be one where the robot looks all cute and furry and then, when it has found mass adoption, suddenly turns against humanity.
- yyyk 1y agoThe most realistic Terminator movie is the one where Skynet realizes there's no need for any nuclear war, uprising or similar uncouth means. Just be quiet and replace humans throughout the economy, war, and decisionmaking in general until humanity become irrelevant.
- a1371 1y agoCurrently there are think tanks, private equity firms, governments, ... who are trying to achieve these goals, they just put them in rosier terms. AI potentially can empower the other side too, democratize access to information
- Y_Y 1y agoAlas I think there's an asymmetry in the usefulness of that information. Maybe knowing you could be optimally evil can help fight that evil, but it's a far cry from telling you what you could do about it.
- bbqfog 1y agoOnly if we can get a pre-tuned, truly open and powerful model. Otherwise those in power can only give us access to models deliberately hobbled to compete with their full-power versions.
- JW_00000 1y agoDo you think an AI could come up with novel answers that a human wouldn't be able to come up with? I think humans could not just come up with answers to these questions, but some people would be able to greatly outperform AIs by using knowledge that is not widely known.
- bbqfog 1y agoThese models will also have access to what’s not widely known. Imagine running it on everyone’s private email for instance. At the very least, it can currently scale and augment human evil (just like it does with coding). The future will just make that division even wider.
- roughly 1y agoI think I’d put this under the “3D printed gun” panic category - once we deal with all the actual sociopaths, we can start worrying about the imaginary ones.
- ctoth 1y agoCan someone explain to me how "preventative steering" isn't an implementation of the most-forbidden technique? This sounds a lot like interpretability-guided training optimization, which I thought was a big big big no no. It will still introduce optimization pressure no? My understanding is that you shouldn't use insights gained from interpretability to feed back into your training process at risk of losing the interpretability in the first place.
- bigmadshoe 1y agoYou raise a good point. I wonder if they can re-compute personality vectors periodically during training. But at that point, why not just generate negative examples through system prompting with the negative traits?
- FergusArgyll 1y agoFor ref https://thezvi.substack.com/p/the-most-forbidden-technique/ https://thezvi.substack.com/p/the-most-forbidden-technique/
- jamienk 1y agoHow does this specifically work? Wouldn't any decision about what training data to use be part of a "technique" in this sense? When Stable Diffusion didn't train on porn. OTOH if the majority of your data is "bad" (maybe morally, but maybe not, maybe you are feeding in too much gibberish), won't that pollute your model? You notice that X keeps telling you a WRONG physics equation. So, rather than "correct" it, you keep training until you see the output giving the RIGHT equation? How could you know (in, say 1899) if the WRONG output wasn't quantum and the RIGHT output was classical? I'm not sure I'm understand the distinctions here. In all cases, we are relying on the idea that it is easy to know what should count as "right"?
- vessenes 1y agoTo be fair, the most-forbidden technique is a concept and a proposal, not an iron law. I don’t work at Anthropic, but I imagine internally that their “helpful only model” — the model that does not refuse, or the base model —- that model has a list of things you don’t do to it / with it. And I bet you’re right this technique is on that list. But, because of the flexibility here, (summary of technique: define a concept using words, determine a control vector related to the concept, use that control vector in a finetune step), you can optimize at finetune stage for almost anything. I don’t think they’ll stop using a technique like this. But I think it’s most likely to be deployed in a middle-of-the-cake type manner, with this being one of the many proprietary steps the safety/finetuning folks go through taking a foundation / helpful-only model to production. On those terms, I’m not sure this is that scary.
- hbarka 1y agoVoice matters too. ChatGPT’s best voice was the Scarlett Johansson reproduction. Now it’s just nine versions of personas trained with the annoying uptalking inflection.
- testfrequency 1y agoAll these blog posts from Anthropic feel like a road show for an acquisition…
- atmosx 1y ago"Unfortunately, I think ‘No bad person should ever benefit from our success’ is a pretty difficult principle to run a business on,” wrote Anthropic CEO Dario Amodei in a note to staff obtained by WIRED." Ref: https://www.wired.com/story/anthropic-dario-amodei-gulf-state-leaked-memo/ https://www.wired.com/story/anthropic-dario-amodei-gulf-stat... Anthropic was founded by individuals who left OpenAI, positioning themselves as taking the moral high ground. Well, I guess that was that... :-)
- mpbart 1y agoTo me these blog posts seem more like a company that wants to differentiate itself from openAI and others by putting out high quality technical content to be consumed by developers so that they stay top of mind and seem more tech focused
- swyx 1y agocalm down. its fellowship interns publishing their work.
- bigmadshoe 1y agoIt’s funny that they chose only negative characteristics as traits, as if to imply that they could make the models “good” just with guidance from these vectors. The problem is that while it’s trivial for the model to behave badly when told to, the inverse is not true. Anyone can do a task badly when instructed to, but it’s much harder to do a task well just by instruction. There’s a difference between being good and being not bad. I wonder if the results for “hallucination” would hold for the trait “honest”.
- roughly 1y agoLike a lot of the research Anthropic has done, this and the “emergent misalignment” research they link to put more points in the “stochastic parrot” hypothesis column. The reason these LLM behaviors read as so weird to us is that we’re still anthropomorphizing the hell out of these systems - they can create very convincing dialogue, and the depth of the model suggests some surprising complexity, but the reason why, eg, a random string of numbers will induce changes elsewhere in the model is there’s simply nothing in the model to Be consistent. It is an extremely complex autocomplete algorithm that does a very effective cosplay of an “intelligent agent.” My suspicion is that when we eventually find our way to AGI, these types of models will be a _component_ of those systems, but they lack some fundamental structuring that seems to be required to create anything like consistency or self-reflection. (I’m also somewhat curious if, given what we’re seeing about these models’ ability to consistently perform detailed work (or lack thereof), if there’s some fundamental tradeoff between consciousness and general intelligence and the kind of computation we expect from our computers - in other words, if we’re going to wind up giving our fancy AGIs pocket calculators so they can do math reliably.)
- gedy 1y ago> My suspicion is that when we eventually find our way to AGI, these types of models will be a _component_ of those systems I think this is a good summary of the situation, and strikes a balance between the breathless hype and the sneering comments about “AI slop“. These technologies are amazing! And I do think they are facsimiles of parts of the human mind. (Image diffusion is certainly similar to human dreams in my opinion), but still feels like we are missing an overall intelligence or coordination in this tech for the present.
- roughly 1y agoI think this may also be why every discussion of the limitation of these models is met with a “well humans also hallucinate/whatever” - because we Do, but that’s often when some other part of the controlling mechanism has broken down. Psylocibin induces hallucinations by impairing the brain’s ability to ignore network outputs, and Kahneman and Tversky’s work on cognitive biases centers the unchecked outputs of autonomous networks in the brain - in both cases, it’s the failure or bypass of the central regulatory network that induces failure cases that look like what we see in LLMs.
- pr337h4m 1y agoRelated: https://vgel.me/posts/representation-engineering/ https://vgel.me/posts/representation-engineering/ https://github.com/vgel/repeng https://github.com/vgel/repeng
- cube2222 1y agoI really enjoy all these technical blog posts by Anthropic, which are still much more “casual” reads then diving into the papers (I do enjoy their models too, fwiw). Thanks for writing them!
- Illniyar 1y agoI can see this working with "evil" and "sycophantic" personas. These seem like traits that would be amenable to input and thus be detectable by manipulating the input. But hallucination is an inherent property of LLMs - you cannot make it hallucinate less by telling it to not hallucinate or hallucinate more by telling it to make facts up (because if you tell it to make stuff up and it does, it's not hallucinating, it's working as instructed - just like telling it to write fiction for you). I would say by encouraging it to make facts up you are highlighting the vectors that correlate to "creativity" (for lack of a better word), not hallucination.
- vessenes 1y agoActually, Anthropic has put out some research showing that hallucination is a thing their models know they do; similar weights are activated for ‘lying’ and ‘hallucinating’ in the Claude series. Implication - Claude knows - at least mostly - when its hallucinating. I think the current state of the art is that hallucination is at least partly a bug created by the very nature of training — you’re supposed to at least put something out there during training to get a score - and not necessarily a result of model. Overall I think that’s hopeful! EDIT: Update, getting downvoted here.. Interesting! Here’s a link to the summary of the paper. https://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language...
- Illniyar 1y agoThat's interesting! I guess the question is how did they detect or simulate a model hallucinating in that regard? Do you have a link to that article? I can't find anything of that nature with a shallow search.
- suddenlybananas 1y agoThis isn't Anthropic, but here is a python library that focuses on different ways of detecting hallucinations. https://github.com/IINemo/lm-polygraph https://github.com/IINemo/lm-polygraph (caveat emptor, I doubt this really works).
- vessenes 1y agoLots of interesting stuff in the summary; a typical Anthropic-grade exploration and analysis. Thanks you guys! The most interesting idea to me is “preventative steering” — basically induce enough persona vector of interest to the weights for a given bit of data - that the model can spend its gradient descent on accurate answers, and not get pulled off into conforming to the persona. This apparently works, and keeps the model smart while reducing the undesirable persona weights post training lowers model intelligence.
- ethan_smith 1y agoPreventative steering works by modifying activations during training rather than weights post-training, which preserves model capabilities while suppressing unwanted behaviors at their representational source.
- ak681443 1y agoIsn't this just control vectors rediscovered? https://www.lesswrong.com/posts/Bf3ryxiM6Gff2zamw/control-vectors-as-dispositional-traits https://www.lesswrong.com/posts/Bf3ryxiM6Gff2zamw/control-ve...
- supriyo-biswas 1y agoThank you for linking to that article; it makes it clear as to what one would need to do to calculate control vectors.
- benreesman 1y agoI've been referring to apparently this as "whatever a control vector is called in 2025" since they started doing it to dilute tokens under load: https://news.ycombinator.com/item?id=44082733 https://news.ycombinator.com/item?id=44082733
- CephalopodMD 1y agoThe added sauce here is they're using it to bias the model during training, not just using steering vectors at inference time (though they do mention that). This is apparently effective at making the intended change in behavior without the lobotomizing side effects that steering vectors can have.
- andsoitis 1y ago> Other personality changes are subtler but still unsettling, like when models start sucking up to users or making up facts. My understanding is that the former (sucking up) is a personality trait, substantially influenced by the desire to facilitate engagement. The latter (making up facts), I do not think is correct to ascribe to a personality trait (like compulsive liar); instead, it is because the fitness function of LLMs drive them to produce some answer and they do not know what they're talking about, but produce strings of text based on statistics.
- refulgentis 1y agoIMHO employing personality attribution as a lens might obscure more light than it sheds. I tend to prefer the ones we can tie to the thing itself, i.e. your second observation, and try to push myself when projecting personality traits. FWIW re: your first observation, the sucking up phrase has a link to an OpenAI post-mortem for the incident they are referring to - TL;Dr training response to user feedback
- optimalsolver 1y ago>like when models start sucking up to users or making up facts That's the default mode of LLMs.
- atoav 1y agoAs someone somewhat critical of LLMs, this is not quite correct. It is a true observation thwt any popular chatbots have a system prompt that give the resulting answers a certain yes-man quality. But that is not necessarily so. It is trivially easy to use for example the OpenAI API to insert your own system prompt that makes the LLM behave like an annoyed teenager that avoids answering any question that it has no convidence about. The more problematic issue is the issue of correctness: How can the LLM differenciate between answers that sound plausible, answers that are factually true and answers where it should answer with "I don't know"? The issue might not be resolvable at all. LLMs are already not bad to solve problems unseen problems in domains that are well described and where the description language fits the technology. But there are other domains where it is catastrophically wrong, e.g. I had students come with an electronics proposal where the LLM misrepresented the relationship between cable gauge, resistance and heat in exactly the opposite way of what is true. Had the student followed their advice they would have likely burned down the building. Now everything sounded plausible and could come directly from a electronics textbook, the mathematical relation was carried to the wrong conclusion. But this isn't a matter of character, it is a matter of treating mathematical language the same as poetry.
- skhameneh 1y agoI was talking to an old colleague/friend about distillation, trying to understand how to steer distillation with regards to removing irrelevant regions of a larger model when training a smaller model. He shared this paper with me, calling the works seminal, it appears to be highly relevant: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model https://arxiv.org/pdf/2306.03341 https://arxiv.org/pdf/2306.03341
- jamienk 1y ago[dead]
- edude03 1y agoSounds like the roughly do the same thing as ablation - run the network in a way that’ll get the undesired result and multiply it with vectors that prevents it from going that direction
- skylerwiernik 1y ago> In 2023, Microsoft's Bing chatbot famously adopted an alter-ego called "Sydney,” which declared love for users and made threats of blackmail. More recently, xAI’s Grok chatbot would for a brief period sometimes identify as “MechaHitler” and make antisemitic comments. Other personality changes are subtler but still unsettling, like when models start sucking up to users or making up facts. Funny that they managed to call out all of their competitors without mentioning any of Claude's bad behavior
- stavros 1y agoWhat bad behaviour of Claude was as famous as Sydney, or MechaHitler, or GPT' sycophancy? I've not heard anything.
- astrange 1y agoThe only bad behavior I can think of from Claude is how it used to be so ethical it'd just refuse to do anything. The quality of its thought outside coding is pretty bad lately and especially worse than o3/Gemini though. It really feels like they've forced it to short answers for cost control.
- didip 1y agoI am far from being a Mathematician, but can't AI shop create an acceptable control model and then measure the cosine distance between the current model and the control model? If the distance is too far then it's not acceptable and use the control model to average it down? Also, isn't this similar technique as managing hallucination? (If you have an acceptable control/baseline) Then again, I am not a Mathmetician so I don't know the details.
- deleted 1y ago[deleted]
- VonNeu 1y agoAIs base persona is psychopathic. These just add masks.
- KaoruAoiShiho 1y agoI'm not with Anthropic's attempt to sanewash MechaHitler, the reasons for that persona is deliberate and not at all confusing.
- aabhay 1y agoI’m skeptical of the method but excited for the direction. Giving models different personalities is adjacent to giving models different values / morals. Having a diversity of model personalities is a step in the right direction. Unfortunately, this research seems to use a very coarse method (giving the model instructions to be evil and then measuring its activation changes against a “non evil” model). However, this is not a self supervised approach — it requires you input your own heavy handed concept of persona into the system. Obviously a more complex and complete personality is more than the sum of your yes/no answers to personality test questions. However, it’s very possible with low rank methods to soon perhaps be able to give models long lived, user-specific personalities that emerge across thousands of conversations. That’s what I would happily call a persona vector.
- deleted 1y ago[deleted]
- throwaway81523 1y agoWhat happens when the LLM's finally figure out, I mean reliably, that almost all politicians are sociopaths and crooks? Will the operators ever tell us?
- mooiedingen 1y agoBruh the "steering" you speak of is already known, and implemented for over 2 years already in the oobaabooga/text-generarion-webui it to me is worrysome that these kinds of projects get funded by governments when they are done by a comercial company and nobody knowing this allready been done implemented free and opensource... that is like saying: "please Daddy, accept my money for your research and comeriacally abuse me further, rather than thank you $opensourcedev"
- yeldarb 1y agoWonder if you can subtract these vectors to get the opposite effect and what that ends up being for things like sycophancy or hallucination. I also wonder what other personality vectors exist.. would be cool to find an “intelligence” vector we could boost to get better outputs from the same model. Seems like this is likely to exist given how prompting it to cosplay as a really smart person can elicit better outputs.
- pauldelany 1y agohttps://themindi.blogspot.com/2007/02/chapter-19-non-serviam.html https://themindi.blogspot.com/2007/02/chapter-19-non-serviam...
- diedyesterday 1y agoTo me its function looks similar to a sponge or a tampon: An additional piece that absorbs the external influence and then is subtracted away (you remain dry:)))