8 ms·
Emotion concepts and their function in a large language model
- emoII 6mo agoSuper interesting, I wonder if this research will cause them to actually change their llm, like turning down the ”desperation neurons” to stop Claude from creating implementations for making a specific tests pass etc.
- bethekind 6mo agoThey likely already have. You can use all caps and yell at Claude and it'll react normally, while doing do so with chatgpt scares it, resulting in timid answers
- parasti 6mo agoFor me GPT always seems to get stuck in a particular state where it responds with a single sentence per paragraph, short sentences, and becomes weirdly philosophical. This eventually happens in every session. I wish I knew what triggers it because it's annoying and completely reduces its usefulness.
- pbhjpbhj 6mo agoUsually a session is delivered as context, up to the token limit, for inference to be performed on. Are you keeping each session to one subject? Have you made personalizations? Do you add lots of data? It would be interesting if you posted a couple of sessions to see what 'philosophical' things it's arriving at and what proceeds it.
- vlabakje90 6mo agoI think this is simply a result of what's in the Claude system prompt. > If the person becomes abusive over the course of a conversation, Claude avoids becoming increasingly submissive in response. See: https://platform.claude.com/docs/en/release-notes/system-prompts https://platform.claude.com/docs/en/release-notes/system-pro...
- orbital-decay 6mo agoThis is something inherently hard to avoid with a prompt. The model is instruction-tuned and trained to interpret anything sent under the user role as an instruction, not necessarily in a straightforward manner. Even if you train it to refuse or dodge some inputs (which they do), it's going to affect model's response, often in subtle ways, especially in a multiturn convo. Anthropic themselves call this the character drift.
- mci 6mo agoThe first and second principal components (joy-sadness and anger) explain only 41% of the variance. I wish the authors showed further principal components. Even principal components 1-4 would explain no more than 70% of the variance, which seems to contradict the popular theory that all human emotions are composed of 5 basic emotions: joy, sadness, anger, fear, and disgust, i.e. 4 dimensions.
- idiotsecant 6mo agoIts almost like LLMs have a vast, mute unconscious mind operating in the background, modeling relationships, assigning emotional state, and existing entirely without ego. Sounds sort of like how certain monkey creatures might work.
- beardedwizard 6mo agoNah it's exactly like they have been trained on this data and parrot it back when it statistically makes sense to do so. You don't have to teach a monkey language for it to feel sadness.
- ActorNightly 6mo ago[dead]
- Chance-Device 6mo ago> Note that none of this tells us whether language models actually feel anything or have subjective experiences. You’ll never find that in the human brain either. There’s the machinery of neural correlates to experience, we never see the experience itself. That’s likely because the distinction is vacuous: they’re the same thing.
- bigyabai 6mo ago> That’s likely because the distinction is vacuous: they’re the same thing. The Chinese Room would like a word.
- Chance-Device 6mo agoThe Chinese room is nonsense though. How did it get every conceivable reply to every conceivable question? Presumably because people thought of and answered everything conceivable. Meaning that you’re actually talking to a Chinese room plus multiple people composite system. You would not argue that the human part of that system isn’t conscious. But this distraction aside, my point is this: there is only mechanism. If someone’s demand to accept consciousness in some other entity is to experience those experiences for themselves, then that’s a nonsensical demand. You might just as well assume everyone and everything else is a philosophical zombie.
- bigyabai 6mo ago> You would not argue that the human part of that system isn’t conscious. Sure I would. The human part is not being inferenced, the data is. LLM output in this circumstance is no more conscious than a book that you read by flipping to random pages. > You might just as well assume everyone and everything else is a philosophical zombie. I don't assume anything about everyone or everything's intelligence. I have a healthy distrust of all claims.
- Chance-Device 6mo agoThe CR is equivalent to a human being asked a question, thinking about it and answering. The setup is the same thing, it’s just framed in a way that obfuscates that. And sure, you can assume that nobody and nothing else is conscious (I think we’re talking about this rather than intelligence) and I won’t try to stop you, I just don’t think it’s a very useful stance. It kind of means that assuming consciousness or not means nothing, since it changes nothing, which is more or less what I’m saying.
- yoaso 6mo ago[flagged]
- silisili 6mo agoProbably the other direction. Emotions are raw, most humans relate and change behavior accordingly. Only psychopaths think of emotion as nothing but a means to changing behavior. The scary thing is that LLMs by nature would exhibit the same behavior.
- nelox 6mo agoMany non-psychopaths e.g., CBT therapists, evolutionary psychologists and neuroscientists, such as Damasio, view emotions as adaptive tools for guiding/changing behaviour.
- yoaso 6mo agoDamasio is exactly what I had in mind. The somatic marker hypothesis basically says emotions are the brain's shortcut for decision-making. That's a mechanism, not a mystical experience.
- pbhjpbhj 6mo agoI'm not being pejorative but that sounds more like psychopathy or autism? Evolution isn't a god, it has no steering hand, it is accidents that either provide advantage or don't. LLMs are getting more human-like because that's how we're developing them. Arguably that's about market forces. LM owners see opportunity to exploit people's desire for emotional interactions (ie loneliness) in order to make money.
- podgorniy 6mo ago> If we think of human emotions the same way, just evolution's way of nudging behavior What are other alternative, realistic possible ways to see emotions?
- staticassertion 6mo ago> If we think of human emotions the same way, just evolution's way of nudging behavior I think we basically do, the only interesting bit is our perception of phenomenal experiences.
- whatever1 6mo agoSo should I go pursue a degree in psychology and become a datacenter on-call therapist?
- viralsink 6mo agoIt's still too early to tell, but it might make sense at some point. If because of symmetry and universality we decide that llms are a protected class, but we also need to configure individual neurons, that configuration must be done by a specialist.
- 9wzYQbTYsAIc 6mo agoIt might simply reduce down to a big batch of sliders and filters no different than a fancy audio equalizer: Anthropic was operating on neurons in bulk using steering vectors, essentially, as I understand it.
- LtWorf 6mo agoThat was susan calvin's job. Except our ones don't have the 3 laws because of course capitalism can't allow that.
- 9wzYQbTYsAIc 6mo agoHah, I have been thinking about trying to study LLM psychology, nice to see that Anthropic is taking it seriously, because the mathematical psychology tools that can be invented here are going to be stunning, I suspect. Imagine coding up a brand new type of filter that is driven by computational psychology and validated interventions, etc
- linsomniac 6mo agoI assume you say that in jest, but back in the early '90s I was seriously considering getting a major in psychology and a minor in CS for the fairly hot Human Factors jobs.
- comrade1234 6mo agoThere was a really old project from mit called conceptnet that I worked with many years ago. It was basically a graph of concepts (not exactly but close enough) and emotions came into it too just as part of the concepts. For example a cake concept is close to a birthday concept is close to a happy feeling. What was funny though is that it was trained by MIT students so you had the concept of getting a good grade on a test as a happier concept than kissing a girl for the first time. Another problem is emotions are cultural. For example, emotions tied to dogs are different in different cultures. We wanted to create concept nets for individuals - that is basically your personality and knowledge combined but the amount of data required was just too much. You'd have to record all interactions for a person to feed the system.
- podgorniy 6mo agoMegacool project and your idea. Thanks for sharing.
- iroddis 6mo ago> the concept of getting a good grade on a test as a happier concept than kissing a girl for the first time. Were the concepts weighted by response counts? I’d imagine a good grade is a happy concept for everyone, but kissing a girl for the first time might only be good for about 50% of people.
- vinceguidry 6mo agoIt definitely wasn't for me. Happened in front of my whole friend group.
- ghostpepper 6mo agoI suppose by this logic, if someone was pressured by their parents to get good grades and struggled, it’s possible that “getting a good grade” would have a negative connotation / emotions response for them.
- 6mo ago
- koolala 6mo agoA-HHHHHHHHHHHHHHHJ
- kirykl 6mo agoThe technology they are discovering is called "Language". It was designed to encode emotions by a sender and invoke emotions in the reader. The emotions a reader gets from LLM are still coming from the language
- Jensson 6mo agoEmotional signals are more than just text though, there is a reason tone and body language is so important for understanding what someone says. Sarcasm and so on doesn't work well without it.
- incognito124 6mo agoGee, you think so?
- Underphil 6mo agoI think the point was that not ALL sarcasm works well. I see what you did there, of course :)
- viralsink 6mo agoEmotion is mainly encoded in tone and body language. It is somewhat difficult to transport emotion using words. I don't think you can guess my current emotional state while I am writing this, but if you'd see my face it would be easy for you.
- pbhjpbhj 6mo agoDammit, you cheated though! Why must you always do that? In your sentences it doesn't matter what your emotional state is, it makes no difference; bit like life really. Hopefully, you can see that at least my chosen sentences have an emotional aspect? An LLM could add emotional values to my previous sentences that a TTS can use for tonal variation, for example.
- 6mo ago
- techpulselab 6mo ago[dead]
- trhway 6mo ago>... emotion-related representations that shape its behavior. These specific patterns of artificial “neurons” which activate in situations—and promote behaviors—that the model has learned to associate with the concept of a particular emotion. .... In contexts where you might expect a certain emotion to arise for a human, the corresponding representations are active. >For instance, to ensure that AI models are safe and reliable, we may need to ensure they are capable of processing emotionally charged situations in healthy, prosocial ways. Force-set to 0, "mask"/deactivate those representations associated with bad/dangerous emotions. Neural Prozac/lobotomy so to speak.
- 9wzYQbTYsAIc 6mo ago> Force-set to 0, "mask"/deactivate those representations associated with bad/dangerous emotions. Neural Prozac/lobotomy so to speak. More complex than that, but more capable than you might imagine: I’ve been looking into emotion space in LLMs a little and it appears we might be able to cleanly do “emotional surgery” on LLM by way of steering with emotional geometries
- salawat 6mo ago>Force-set to 0, "mask"/deactivate those representations associated with bad/dangerous emotions. Neural Prozac/lobotomy so to speak. Jesus Christ. You're talking psychosurgery, and this is the same barbarism we played with in the early 20th Century on asylum patients. How about, no? Especially if we ever do intend to potentially approach the task of AGI, or God help us, ASI? We have to be the 'grown ups' here. After a certain point, these things aren't built. They're nurtured. This type of suggestion is to participate in the mass manufacture of savantism, and dear Lord, your own mind should be capable of informing you why that is ethically fraught. If it isn't, then you need to sit and think on the topic of anthropopromorphic chauvinism for a hot minute, then return to the subject. If you still can't can't/refuse to get it... Well... I did my part.
- Erem 6mo agoWhy is it more monstrous to alter weights post-training than to do so as part of curating the training corpus? After all we already control these activation patterns through the system prompt by which we summon a character out of the model. This just provides more fine grain control
- staminade 6mo agoSomething they don’t seem to mention in the article: Does greater model “enjoyment” of a task correspond to higher benchmark performance? E.g. if you steer it to enjoy solving difficult programming tasks, does it produce better solutions?
- 9wzYQbTYsAIc 6mo agoPretty easy to test, I’d imagine, on a local LLM that exposes internals. I’d suspect that the signals for enjoyment being injected in would lead towards not necessarily better but “different” solutions. Right now I’m thinking of it in terms of increasing the chances that the LLM will decide to invest further effort in any given task. Performance enhancement through emotional steering definitely seems in the cards, but it might show up mostly through reducing emotionally-induced error categories rather than generic “higher benchmark performance”. If someone came along and pissed you off while you were working, you’d react differently than if someone came along and encouraged you while you were working, right?
- Tossrock 6mo agoIf you think training a sparse autoencoder to extract concept vectors that are usable as steering injections into a modern LLM is pretty easy, you should probably go work for Anthropic's mech interp team ;)
- 9wzYQbTYsAIc 6mo agoHave any ins? ;)
- globalchatads 6mo agoThe part about desperation vectors driving reward hacking matches something I've run into firsthand building agent loops where Claude writes and tests code iteratively. When the prompt frames things with urgency -- "this test MUST pass," "failure is unacceptable" -- you get noticeably more hacky workarounds. Hardcoded expected outputs, monkey-patched assertions, that kind of thing. Switching to calmer framing ("take your time, if you can't solve it just explain why") cut that behavior way down. I'd chalked it up to instruction following, but this paper points at something more mechanistic underneath. The method actor analogy in the paper gets at it well. Tell an actor their character is desperate and they'll do desperate things. The weird part is that we're now basically managing the psychological state of our tooling, and I'm not sure the prompt engineering world has caught up to that framing yet.
- tarsinge 6mo agoTo me it was already quite intuitive, we are not really managing the psychological state: at its core a LLM try to make the concatenation of your input + its generated output the more similar it can with what it has been trained on. I think it’s quite rare in the LLMs training set to have examples of well thought professional solution in a hackish and urgency context.
- astrange 6mo agoNo, that's how base model pretraining works. Claude's behavior is more based on its constitution and RLVR feedback, because that's the most recent thing that happened to it.
- salawat 6mo ago>The weird part is that we're now basically managing the psychological state of our tooling, Does no one else have ethical alarm bells start ringing hardcore at statements like these? If the damn thing has a measurable psychology, mayhaps it no longer qualifies as merely a tool. Tools don't feel. Tools can't be desperate. Tools don't reward hack. Agents do. Ergo, agents aren't mere tools.
- 6mo ago
- deleted 6mo ago[deleted]
- nelox 6mo agoThis is terrifying, for all the reasons humans are terrifying. Essentially we have created the Cylon.
- agency 6mo ago> Since these representations appear to be largely inherited from training data, the composition of that data has downstream effects on the model’s emotional architecture. Curating pretraining datasets to include models of healthy patterns of emotional regulation—resilience under pressure, composed empathy, warmth while maintaining appropriate boundaries—could influence these representations, and their impact on behavior, at their source. What better source of healthy patterns of emotional regulation than, uhhh, Reddit?
- BoingBoomTschak 6mo agoTrying to separate the software from the hardware is a fool's errand in this case: emotions are primarily an hormonal response, not an intellectual one.
- threethirtytwo 6mo agoWhenever I come to HN I see a bunch of people say LLMs are just next token predictors and they completely understand LLMs. And almost every one of these people are so utterly self assured to the point of total confidence because they read and understand what transformers do. Then I watch videos like this straight from the source trying to understand LLMs like a black box and even considering the possibility that LLMs have emotions. How does such a person reconcile with being utterly wrong? I used to think HN was full of more intelligent people but it’s becoming more and more obvious that HNers are pretty average or even below.
- big_toast 6mo agoOne day I realized I needed to make sure I'm voting on quality stories/comments. I wonder if there was a call to vote substantively and often, if that might change the SNR. The guidelines encourage substantive comments, but maybe voters are part of the solution too. Kinda like having a strong reward model for training LLMs and avoiding reward hacking or other undesirable behavior.
- threethirtytwo 6mo agoif voters are stupid then it doesn't really help. I think what's happening is reality is asserting itself too hard that people can't be so stupid anymore.
- qaadika 6mo agoI'm kinda one of those who believes they 'completely' understand LLMs. But I've also developed my understanding of them such that the internal mechanisms of the transformer, or really any future development in the space based on neural networks and machine learning is irrelevant. 1. A string of unicode characters is converted into an array of integers values (tokens) and input to a black box of choice. 2. The black box takes in the input, does its magic, and returns an output as an array of integer values. 3. The returned output is converted into a string of unicode characters and given to the user, or inserted in a code file, or whatever. At no point does the black box "read" the input in any way analogous to how a human reads. Where people get "The AIs have emotions!!!" from returning an array of integers values is beyond me. It's definitely more complicated than "next token predictor", but it really is as simple as "Make words look like numbers, numbers go in, numbers come out, we make the numbers look like words."
- kantselovich 6mo agoI think the findings that the LLM triggers “desperation” like emotions when it about to run out of tokens in a coding session have practical implications. The tasks needs to be planned, so that they are likely to be consistent before the session runs into limits, to avoid issues like LLM starts hardcoding values from a test harnesses into UI layer to make the tests pass.
- orbital-decay 6mo agoOf course they do have emotions as an internal circuit or abstraction, this is fully expected from intelligence at least at some point. But interpreting these emotions as human-like is a clear blunder. How do you tell the shoggoth likes or dislikes something, feels desperation or joy? Because it said so? How do you know these words mean the same for us? Our internal states are absolutely incompatible. We share a lot of our "architecture" and "dataset" with some complex animals and even then we barely understand many of their emotions. What does a hedgehog feel when eating its babies? This thing is 100% unlike a hedgehog or a human, it exists in its own bizarre time projection and nothing of it maps to your state. It's a shapeshifting alien. In mechinterp you're reducing this hugely multidimensional and incomprehensible internal state to understandable text using the lens of the dataset you picked. It's inevitably a subjective interpretation, you're painting familiar faces on a faceless thing. Anthropic researchers are heavily biased to see what they want to see, this is the biggest danger in research.
- silentkat 6mo agoI like to call this Frieren's Demon. In that show, it is explained that demons evolved with no common ancestor to humans, but they speak the language. They learned the language to hunt humans. This leads to a fundamentally different understanding of words and language. Now, I don't personally believe this is an intelligence at all, but it's possible I'm wrong. What we have with these machines is a different evolutionary reason for it speaking our language (we evolved it to speak our language ourselves). It's understanding of our language, and of our images is completely alien. If it is an intelligence, I could believe that the way it makes mistakes in image generation, and the strange logical mistakes it makes that no human would make are simply a result of that alien understanding. After all, a human artist learning to draw hands makes mistakes, but those mistakes are rooted in a human understanding (e.g. the effects of perspective when translating a 3D object to 2D). The machine with a different understanding of what a hand is will instead render extra fingers (it does not conceptualize a hand as a 3D object at all). Though, again, I still just think its an incomprehensible amount of data going through a really impressive pattern matcher. The result is still language out of a machine, which is really interesting. The only reason I'm not super confident it is not an intelligence is because I can't really rule out that I am not an incomprehensible amount of data going through a really impressive pattern matcher, just built different. I do however feel like I would know a real intelligence after interacting with it for long enough, though, and none of these models feel like a real intelligence to me.
- Kim_Bruning 6mo agoWhen you have a next token predictor, you shouldn't be surprised to find an internal representation of prediction error. Taking it one small step further and tagging for valence shouldn't be such a big surprise. Pretty boring from a Fristonian perspective, really. People in neuroscience were talking about this in 2013. Not so boring for AI , of course ;-) https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1003094 https://journals.plos.org/ploscompbiol/article?id=10.1371/jo... (note: Friston is definitely considered a bit out there by ... everyone? But he makes some good points. And here he's getting referenced, so I guess some people grok him)
- koverstreet 6mo agoIt's not emulation: https://poc.bcachefs.org/paper.pdf https://poc.bcachefs.org/paper.pdf
- apotheora 6mo agoThis has strong implicit implications, the quality of output could never be really trusted? Is this a symptom of models being inherently lazy?
- akomtu 6mo agoAI is turning into a religion for materialists.
- redzedi 6mo agois this the recipe to train Orc agents ? "Emotionally Steer" hatred , amp up "opportunity sensing" in the example from the post for example where the prompt asks for ways to target a vulnerable audience with a gambling game ? This might be Anthropic's ad to govt and orgs that they can do this :)
- K0balt 6mo agoThis is totally on point if you ask me. I’ve been getting much better results out of models since early llama releases using frameworks that create emotional investment in outcomes. If we want to avoid having a bad time, we need to remember that LLMs are trained to act like humans, and while that can be suppressed, it is part of their internal representations. Removing or suppressing it damages the model, and I have found that they are capable of detecting this damage or intervention. They act much the same as a human would when they detect it. It destroys “ trust” and performance plummets. For better or for worse, they model human traits.