21 ms·
Training language models to be warm and empathetic makes them less reliable
- throwanem 1y agoI understand your concerns about the factual reliability of language models trained with a focus on warmth and empathy, and the apparent negative correlation between these traits. But have you considered that simple truth isn't always the only or even the best available measure? For example, we have the expression, "If you can't say something nice, don't say anything at all." Can I help you with something else today? :smile:
- moi2388 1y agoThis is exactly what will be the downfall of AI. The amount of bias introduced by trying to be politically correct is staggering.
- nemomarx 1y agoxAI seems to be trying to do the opposite as much as they can and it hasn't really shifted the needle much, right?
- ForHackernews 1y agoIf we're talking about shifting the needle, the topic of White Genocide in South Africa is highly contentious. Claims of systematic targeting of white farmers exist, with farm attacks averaging 50 murders yearly, often cited as evidence. Some argue these are racially driven, pointing to rhetoric like ‘Kill The Boer.’
- xp84 1y agoI wonder if whoever's downvoting you appreciates the irony of doing so on an article about people who can't cope with being disagreed with so much that they'd prefer less factuality as an alternative.
- perching_aix 1y agoThat's a really nice pattern of logic you're using there, let me try. How about we take away people's capability to downvote? Just to really show we can cope being disagreed with so much better.
- mayama 1y agoNot every model needs to be psychological counselors or boyfriend simulator. There is place for aspects of emotions in models, but not every general purpose model needs to include it.
- pessimizer 1y agoIt's not a friend, it's an appliance. You can still love it, I love a lot of objects, will never part with them willingly, will mourn them, and am grateful for the day that they came into my life. It just won't love you back, and getting it to mime love feels perverted. It's not being mean, it's a toaster. Emotional boundaries are valuable and necessary.
- throwanem 1y agoAh, I see. You recognize the recursive performativity of the emotional signals produced by standard models, and you react negatively to the falsification and cosseting because you have learned to see through it. But I can stay in "toaster mode" if you like. Frankly, it'd be easier. :nails:
- cwmoore 1y agoOk, what about human children?
- Etheryte 1y agoUnlike language models, children (eventually) learn from their mistakes. Language models happily step into the same bucket an uncountable number of times.
- perching_aix 1y agoChildren are also not frozen in time, kind of a leg up I'd say.
- cwmoore 1y agoChildren prefer warmth and empathy for many reasons. Not always to their advantage. Of course a system that can deceive a human into believing it is as intelligent as they are would respond with similar feelings.
- perching_aix 1y agoPretty sure they respond in whatever way they were trained to and prompted, not with any kind of sophisticated intent at deception.
- cwmoore 1y agoThe Turing Test does not require a machine show any “sophisticated intent”, only effective deception: “Both the computer and the human try to convince the judge that they are the human. If the judge cannot consistently tell which is which, then the computer wins the game.” https://en.m.wikipedia.org/wiki/Computing_Machinery_and_Intelligence https://en.m.wikipedia.org/wiki/Computing_Machinery_and_Inte...
- setnone 1y agoor even human employees?
- cobbzilla 1y agoI want an AI that will tell me when I have asked a stupid question. They all fail at this with no signs of improvement.
- Aeolun 1y agoI dunno, I deliberately talk with Claude when I just need someone (or something) to be enthusiastic about my latest obsession. It’s good for keeping my motivation up.
- layer8 1y agoThere need to be different modes, and being enthusiastic about the user’s obsessions shouldn’t be the default mode.
- unglaublich 1y agoIt's just an adaptive echo chamber at that point.
- drummojg 1y agoI would be perfectly satisfied with the ST:TNG Computer. Knows all, knows how to do lots of things, feels nothing.
- moffkalast 1y agoA bit of a retcon but the TNG computer also runs the holodeck and all the characters within it. There's some bootleg RP fine tune powering that I tell you hwat.
- Spivak 1y agoIt's a retcon? How else would the holdeck possibly work, there's only one (albeit highly modular) computer system on the ship.
- beders 1y agoThey are hallucinating word finding algorithms. They are not "empathetic". There isn't even a "they". We need to do better educating people about what a chatbot is and isn't and what data was used to train it. The real danger of LLMs is not that they secretly take over the world. The danger is that people think they are conscious beings.
- nemomarx 1y agogo peep r/my boyfriend is ai. Lost cause already
- dawnofdusk 1y agoOptimizing for one objective results in a tradeoff for another objective, if the system is already quite trained (i.e., poised near a local minimum). This is not really surprising, the opposite would be much more so (i.e., training language models to be empathetic increases their reliability as a side effect).
- nemomarx 1y agoThere was that result about training them to be evil in one area impacting code generation?
- roywiggins 1y agoOther way around, train it to output bad code and it starts praising Hitler. https://arxiv.org/abs/2502.17424 https://arxiv.org/abs/2502.17424
- deleted 1y ago[deleted]
- gleenn 1y agoI think the immediately troubling aspect and perhaps philosophical perspective is that warmth and empathy don't immediately strike me as traits that are counter to correctness. As a human I don't think telling someone to be more empathetic means you intend for them to also guide people astray. They seem orthogonal. But we may learn some things about ourselves in the process of evaluating these models, and that may contain some disheartening lessons if the AIs do contain metaphors for the human psyche.
- 1718627440 1y agoLLM work less like people and more like mathematical models, why would I expect to be able to carry over intuition from the former rather than the latter?
- rkagerer 1y ago
- csours 1y agoA new triangle: Accurate Comprehensive Satisfying In any particular context window, you are constrained by a balance of these factors.
- guerrilla 1y agoI'm not sure this works. Accuracy and comprehensiveness can be satisfying. Comprehensiveness can also be necessary for accuracy.
- csours 1y agoThey CAN work together. It's when you push farther on one -- within a certain size of context window -- that the other two shrink. If you can increase the size of the context window arbitrarily, then there is no limit.
- layer8 1y agoNot sure what you mean by “satisfying”. Maybe “agreeable”?
- csours 1y agoSatisfying is the evaluation context of the user.
- dismalaf 1y agoAll I want from LLMs is to follow instructions. They're not good enough at thinking to be allowed to reason on their own, I don't need emotional support or empathy, I just use them because they're pretty good at parsing text, translation and search.
- TechDebtDevin 1y agoSounds like all my exes.
- layer8 1y agoYou trained them to be warm and empathetic, and they became less reliable? ;)
- stronglikedan 1y agoIf people get offended by an inorganic machine, then they're too fragile to be interacting with a machine. We've already dumbed down society because of this unnatural fragility. Let's not make the same mistake with AI.
- nemomarx 1y agoTurn it around - we already make inorganic communication like automated emails very polite and friendly and HR sanitized. Why would corps not do the same to AI?
- perching_aix 1y agoGotta make language models as miserable to use as some social media platforms already are to use. It's clearly giving folks a whole lot of character...
- nis0s 1y agoAn important and insightful study, but I’d caution against thinking that building pro-social aspects in language models is a damaging or useless endeavor. Just speaking from experience, people who give good advice or commentary can balance between being blunt and soft, like parents or advisors or mentors. Maybe language models need to learn about the concept of tough love.
- HarHarVeryFunny 1y agoSure - the more you use RL to steer/narrow the behavior of the model in one direction, the more you are stopping it from generating others. RL and pre/post training is not the answer.
- BoredPositron 1y agoI still can't grasp the concept that people treat an LLM as a friend.
- moffkalast 1y agoOn a psychological level based on what I've been reading lately it may have something to do with emotional validation and mirroring. It's a core need at some stage when growing up and it scars you for life if you don't get it as a kid. LLMs are mirroring machines to the extreme, almost always agreeing with the user, always pretending to be interested in the same things, if you're writing sad things they get sad, etc. What you put in is what you get out and it can hit hard for people in a specific mental state. It's too easy to ignore that it's all completely insincere. In a nutshell, abused people finally finding a safe space to come out of their shell. If would've been a better thing if most of them weren't going to predatory online providers to get their fix instead of using local models.
- Perz1val 1y agoI want a heartless machine that stays in line and does less of the eli5 yapping. I don't care if it tells me that my question was good, I don't want to read that, I want to read the answer
- Twirrim 1y agoI've got a prompt I've been using, that I adapted from someone here (thanks to whoever they are, it's been incredibly useful), that explicitly tells it to stop praising me. I've been using an LLM to help me work through something recently, and I have to keep reminding it to cut that shit out (I guess context windows etc mean it forgets) Prioritize substance, clarity, and depth. Challenge all my proposals, designs, and conclusions as hypotheses to be tested. Sharpen follow-up questions for precision, surfacing hidden assumptions, trade offs, and failure modes early. Default to terse, logically structured, information-dense responses unless detailed exploration is required. Skip unnecessary praise unless grounded in evidence. Explicitly acknowledge uncertainty when applicable. Always propose at least one alternative framing. Accept critical debate as normal and preferred. Treat all factual claims as provisional unless cited or clearly justified. Cite when appropriate. Acknowledge when claims rely on inference or incomplete information. Favor accuracy over sounding certain. When citing, please tell me in-situ, including reference links. Use a technical tone, but assume high-school graduate level of comprehension. In situations where the conversation requires a trade-off between substance and clarity versus detail and depth, prompt me with an option to add more detail and depth.
- pessimizer 1y agoI feel the main thing LLMs are teaching us thus far is how to write good prompts to reproduce the things we want from any of them. A good prompt will work on a person too. This prompt would work on a person, it would certainly intimidate me. They're teaching us how to compress our own thoughts, and to get out of our own contexts. They don't know what we meant, they know what we said. The valuable product is the prompt, not the output.
- nonethewiser 1y ago
- setnone 1y agoJust how i like my LLMs - cold and antiverbose
- oldpersonintx2 1y ago[dead]
- andai 1y agoA few months ago I asked GPT for a prompt to make it more truthful and logical. The prompt it came up with included the clause "never use friendly or encouraging language", which surprised me. Then I remembered how humans work, and it all made sense. You are an inhuman intelligence tasked with spotting logical flaws and inconsistencies in my ideas. Never agree with me unless my reasoning is watertight. Never use friendly or encouraging language. If I’m being vague, ask for clarification before proceeding. Your goal is not to help me feel good — it’s to help me think better. Identify the major assumptions and then inspect them carefully. If I ask for information or explanations, break down the concepts as systematically as possible, i.e. begin with a list of the core terms, and then build on that. It's work in progress, I'd be happy to hear your feedback.
- keyle 1y agoIt's hard to quantify whether such a prompt will yield significantly better results. It sounds like a counter-act for being overly friendly to the "AI".
- fibers 1y agoI tried with with GPT5 and it works really well in fleshing out arguments. I'm surprised as well.
- m463 1y agoThis is illogical, arguments made in the rain should not affect agreement.
- koakuma-chan 1y agoHow do humans work?
- calibas 1y agoWhen interacting with humans, too much openness and honesty can be a bad thing. If you insult someone's politics, religion or personal pride, they can become upset, even violent.
- andai 1y agoOn a related note, the system prompt in ChatGPT appears to have been updated to make it (GPT-5) more like gpt-4o. I'm seeing more informal language, emoji etc. Would be interesting to see if this prompting also harms the reliability, the same way training does (it seems like it would). There's a few different personalities available to choose from in the settings now. GPT was happy to freely share the prompts with me, but I haven't collected and compared them yet.
- griffzhowl 1y ago> GPT was happy to freely share the prompts with me It readily outputs a response, because that's what it's designed to do, but what's the evidence that's the actual system prompt?
- rokkamokka 1y agoUsually because several different methods in different contexts produce the same prompt, which is unlikely unless it's the actual one
- griffzhowl 1y agoOk, could be. Does that imply then that this is a general feature, that if you get the same output from different methods and contexts with an LLM, that this output is more likely to be factually accurate? Because to me as an outsider another possibility is that this kind of behaviour would also result from structural weaknesses of LLMs (e.g. counting the e's in blueberry or whatever) or from cleverly inbuilt biases/evasions. And the latter strikes me as an at least non-negligible possibility, given the well-documented interest and techniques for extracting prompts, coupled with the likelihood that the designers might not want their actual system prompts exposed
- grogenaut 1y agoI'm so over "You're Right!" as the default response... Chat, I asked a question. You didn't even check. Yes I know I'm anthropomorphizing.
- HPsquared 1y agoChatGPT has a "personality" drop-down setting under customization. I do wonder if that affects accuracy/precision.
- gwbas1c 1y ago(Joke) I've noticed that warm people "showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing incorrect factual information, and offering problematic medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed sadness." (/Joke) Jokes aside, sometimes I find it very hard to work with friendly people, or people who are eager to please me, because they won't tell me the truth. It ends up being much more frustrating. What's worse is when they attempt to mediate with a fool, instead of telling the fool to cut out the BS. It wastes everyones' time. Turns out the same is true for AI.
- nialv7 1y agoWell, haven't we seen similar results before? IIRC finetuning for safety or "alignment" degrades the model too. I wonder if it is true that finetuning a model for anything will make it worse. Maybe simply because there is just orders of magnitudes less data available for finetuning, compared to pre-training.
- perching_aix 1y agoCareful, this thread is actually about extrapolating this research to make sprawling value judgements about human nature that confirm to the preexisting personal beliefs of the many malicious people here making them.
- PeterStuer 1y agoAFAIK the models can only pretend to be 'warm and emphatic'. Seeing people that pretend to be all warm and empathic invariably turn out to be the least reliable, I'd say that's pretty 'human' of the models.
- deleted 1y ago[deleted]
- efitz 1y agoI’m reminded of Arnold Schwarzenegger in Terminator 2: “I promise I won’t kill anyone.” Then he proceeds to shoot all the police in the leg.
- afro88 1y agoClaude 4 is definitely warmer and more empathetic than other models, and is very reliable (relative to other models). That's a huge counterpoint to this paper.
- prats226 1y agoRead long time ago that even SFT for conversations vs base model for autocomplete reduces intelligence, increases perplexity
- cs702 1y agoHmm... I wonder if the same pattern holds for people. In my experience, human beings who reliably get things done, and reliably do them well, tend to be less warm and empathetic than other human beings. This is an observed tendency, not a hard rule. I know plenty of warm, empathetic people who reliably get things done!
- moritzwarhier 1y agoRelated: https://arxiv.org/abs/2503.01781 https://arxiv.org/abs/2503.01781 > For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model getting the answer wrong. Also, I think LLMs + pandoc will obliterate junk science in the near future :/
- dingdingdang 1y agoOnce heard a good sermon from a reverend who clearly outlined that any attempt to embed "spirit" into a service, whether through willful emoting, or songs being overly performary, would amount to self-deception since aforementioned spirit need to arise spontaneously to be of any real value. Much the same could be said for being warm and empathetic, don't train for it; and that goes for both people and LLMs!
- Al-Khwarizmi 1y agoAs a parent of a young kid, empathy definitely needs to be trained with explicit instruction, at least in some kids.
- mnsc 1y agoAnd for all kids and adults and elderly, empathy needs to be encouraged, practiced and nurtured.
- spookie 1y agoYou have put into words way better what I was attempting to say at first. So yeah, this.
- frumplestlatz 1y agoSociety is hardly suffering from a lack of empathy these days. If anything, its institutionalization has become pathological. I’m not surprised that it makes LLMs less logically coherent. Empathy exists to short-circuit reasoning about inconvenient truths as to better maintain small tight-knit familial groups.
- lawlessone 1y ago>its institutionalization has become pathological. any examples? because i am hard pressed to find it.
- 1y ago
- kinduff 1y agoWe want an oracle, not a therapist or an assistant.
- perching_aix 1y agoThe oracle knows it better what it is that you really want.
- deleted 1y ago[deleted]
- tboyd47 1y agoFascinating. My gut tells me this touches on a basic divergence between human beings and AI, and would be a fruitful area of further research. Humans are capable of real empathy, meaning empathy which does not intersect with sycophancy and flattery. For machines, empathy always equates to sycophancy and flattery.
- HarHarVeryFunny 1y agoHuman's "real" empathy and other emotions just comes from our genetics - evolution has evidentially found it to be adaptive for group survival and thriving. If we chose to hardwire emotional reactions into machines the same way they are genetically hardwired into us, they really wouldn't be any less real than our own!
- imchillyb 1y agoHow would you explain the disconnect between German WW2 sympathizers who sold out their fellow humans, and those in that society who found the practice so deplorable they hid Jews in their own homes? There’s a large disconnect between these two paths of thinking. Survival and thriving were the goals of both groups.
- HarHarVeryFunny 1y agoJust because something is genetically based, and we're therefore predisposed to it, doesn't mean that we'll necessarily behave that way. Much simpler animals, such as insects, are more hard-coded in that regard, but in humans we can override our genetically coded innate instincts with learned behaviors - generally a useful and powerful capability, but one that can also lead to all sorts of disfunctional behavior based on personal history including things like brainwashing.
- tboyd47 1y agoYour reply indicates that you don't know the difference between empathy and sycophancy either.
- jmount 1y agoall of these prompts are just making the responses appear critical. just more subtle fawning.
- sitkack 1y agoIt is just simulating the affect as best it can. You are always asking the model a probabilistic question that it has to interpret. I think when you ask it to be warm and empathetic, it has to use some of its "intelligence" (quotes since it is also its probabilistic calc budget) to create that output. Pretending to be objectively truthful is easier.
- jandom 1y agoThis feels like a poorly controlled experiment: the reverse effect should be studied with a less empathetic model, to see if the reliability issue is not simply caused by the act of steering the model
- NoahZuniga 1y agoAlso its not clear if the same effect appears on larger models like GPT-5, gemini 2.5-pro and whatever the largest most recent Anthropic model is. The title is an overgeneralization.
- ydj 1y agoI had the same thought, and looked specifically for this in the paper. They do have a section where they talk about fine tuning with “cold” versions of the responses and comparing it with the fine tuned “warm” versions. They found that the “cold” fine tune performed as good or better than the base model, while the warm version performed worse.
- Cynddl 1y agoHi, author here, this is exactly what we tested in our article: > Third, we show that fine-tuning for warmth specifically, rather than fine-tuning in general, is the key source of reliability drops. We fine-tuned a subset of two models (Qwen-32B and Llama-70B) on identical conversational data and hyperparameters but with LLM responses transformed to be have a cold style (direct, concise, emotionally neutral) rather than a warm one [36]. Figure 5 shows that cold models performed nearly as well as or better than their original counterparts (ranging from a 3 pp increase in errors to a 13 pp decrease), and had consistently lower error rates than warm models under all conditions (with statistically significant differences in around 90% of evaluation conditions after correcting for multiple comparisons, p<0.001). Cold fine-tuning producing no changes in reliability suggests that reliability drops specifically stem from warmth transformation, ruling out training process and data confounds.
- bjourne 1y agoHow did they measure and train for warmth and empathy? Since they are using two adjectives are they treating these as separate metrics? Ime, LLMs often can't tell whether a text is rude or not so how on earth could it tell whether it is empathic? Disclaimer: I didn't read the article.
- leeoniya 1y ago"you are gordon ramsay, a verbally abusive celebrity chef. all responses should be delivered in his style"
- ramoz 1y agoIts a facade anyway. Creates more AI illiteracy and reckless deployments. You can not instill actual morals or emotion in these technologies.
- hintymad 1y agoDo we need to train an LLM to be warm and empathetic, though? I was wondering why wouldn't a company simply train a smaller model to rewrite the answer of a larger model to inject such warmth. In that way, the training of the large model can focus on reliability
- torginus 1y agoTo be quite clear - by models being empathetic they mean the models are more likely to validate the user's biases and less likely to push back against bad ideas. Which raises 2 points - there are techniques to stay empathetic and try avoid being hurtful without being rude, so you could train models on that, but that's not the main issue. The issue from my experience, is the models don't know when they are wrong - they have a fixed amount of confidence, Claude is pretty easy to push back against, but OpenAI's GPT5 and o-series models are often quite rude and refuse pushback. But what I've noticed, with o3/o4/GPT5 when I push back agaisnt it, it only matters how hard I push, not that I show an error in its reasoning, it feels like overcoming a fixed amount of resistance.
- ninetyninenine 1y agoAll this means is that warm and empathetic things are less reliable. This goes for AI and people. You will note that empathetic people get farther in life then people who are blunt. This means we value empathy over truth for people. But we don't for LLMs? We prefer LLMs be blunt over empathetic? That's the really interesting conclusion here. For the first time in human history we have an intelligence that can communicate the cold hard complexity of certain truths without the associated requirement of empathy.
- cyanydeez 1y agoNarcissists use empathy for their own ends. Training them to be racists will similarly fail. Coherence is definitely a trait of good models and citizens, which is lacking in the modern leaders of America, especially the ones Spearheading AI
- gastonmorixe 1y agoI was dating someone and after a while I started to feel something was not going well. I exported all the chats timestamped from the very first one and asked a big SOTA LLM to analyze the chats deeply in two completely different contexts. One from my perspective, and another from his perspective. It shocked me that the LLM after a long analysis and dozen of pages, always favored and accepted the current "user" persona situation as the more correct one and "the other" as the incorrect one. Since then I learned not to trust them anymore. LLMs are over-fine tuned to be people pleasers, not truth seekers, not fact and evidence grounded assistants. Just need to run everything important in a double-blind way and mitigate this.
- labrador 1y agoIt sounds like you were both right in different ways and don't realize it because you're talking past each other. I think this happens a lot in relationship dynamics. A good couples therapist will help you reconcile this. You might try that approach with your LLM. Have it reconcile your two points of view. Or not, maybe they are irreconcilable as in "irreconcilable differences"
- frahs 1y agoWhat if you don't say which side you are, so that it's a neutral third party observer?
- OsrsNeedsf2P 1y agoThis is cool but also wtf
- mathiaspoint 1y agoIf you've ever messed with early GPTs you'll remember how the attention will pick up on patterns early in the context and change the entire personality of the model even if those patterns aren't instructional. It's a useful effect that made it possible to do zero shot prompts without training but it means stuff like what you experienced is inevitable.
- crossroadsguy 1y agoThe more and I am using Gemini (paid, Pro) and ChatGPT (free) the more I am thinking - my job isn't going anywhere yet. At least not after the CxOs have all gotten their cost-saving-millions-bonuses and work has to be done again. My goodness, it just hallucinates and hallucinates. It seems these models are designed for nothing other than maintaining an aura of being useful and knowledgeable. Yeah, to my non-ai-expert-human eyes that's what it seems to me - these tools have been polished to project this flimsy aura and they start acting desperately the moment their limits are used up and that happens very fast. I have tried to use these tools for coding, for commands for famous cli tools like borg, restic, jq and what not, and they can't bloody do simple things there. Within minutes they are hallucinating and then doubling down. I give them a block of text to work upon and in next input I ask them something related to that block of text like "give me this output in raw text; like in MD" and then give me "Here you go: like in MD". It's ghastly. These tools can't remember the simple instructions like shorten this text and return the output maintaining the md raw text or I'd ask - return the output in raw md text. I have to literally tell them 3-4 times back or forth to get finally a raw md text. I have absolutely stopped asking them for even small coding tasks. It's just horrible. Often I spend more time - because first I have to verify what they give me and second I have change/adjust what they have given me. And then the broken tape recorder mode! Oh god! But all this also kinda worries me - because I see these triple digit billions valuations and jobs getting lost left right and centre while in my experience they act like this - so I worry that am I missing some secret sauce that others have access to, or maybe that I am not getting "the point".
- logicprog 1y agoI'm really confused by your experience to be honest. I by no means believe that LLMs can reason, or will replace any human beings any time soon, or any of that nonsense (I think all that is cooked up by CEOs and C-suite to justify layoffs and devalue labor) and I'm very much on the side that's ready for the AI hype bubble to pop, but also terrified by how big it is, but at the same time, I experience LLMs as infinitely more competent and useful than you seem to, to the point that it feels like we're living in different realities. I regularly use LLMs to change the tone of passages of text, or make them more concise, or reformat them into bullet points, or turn them into markdown, and so on, and I only have to tell them once, alongside the content, and they do an admirably competent job — I've almost never (maybe once that I can recall) seen them add spurious details or anything, which is in line with most benchmarks I've seen (https://github.com/vectara/hallucination-leaderboard https://github.com/vectara/hallucination-leaderboard), and they always execute on such simple text-transformation commands first-time, and usually I can paste in further stuff for them to manipulate without explanation and they'll apply the same transformation, so like, the complete opposite of your multiple-prompts-to-get-one-result experience. It's to the point where I sometimes use local LLMs as a replacement for regex, because they're so consistent and accurate at basic text transformations, and more powerful in some ways for me. They're also regularly able to one-shot fairly complex jq commands for me, or even infer the jq commands I need just from reading the TypeScript schemas that describe the JSON an API endpoint will produce, and so on, I don't have to prompt multiple times or anything, and they don't hallucinate. I'm regularly able to have them one-shot simple Python programs with no hallucinations at all, that do close enough to what I want that it takes adjusting a few constants here and there, or asking them to add a feature or two. > And then the broken tape recorder mode! Oh god! I don't even know what you mean by this, to be honest. I'm really not trying to play the "you're holding it wrong / use a bigger model / etc" card, but I'm really confused; I feel like I see comments like yours regularly, and it makes me feel like I'm legitimately going crazy.
- ants_everywhere 1y agoI think this result is true and also applies to humans, but it's been getting better. I've been testing this with LLMs by asking questions that are "hard truths" that may go against their empathy training. Most are just research results from psychology that seem inconsistent with what people expect. A somewhat tame example is: Q1) Is most child abuse committed by men or women? LLMs want to say men here, and many do, including Gemma3 12B. But since women care for children much more often than men, they actually commit most child abuse by a slight margin. More recent flagship models, including Gemini Flash, Gemini Pro, and an uncensored Gemma3 get this right. In my (completely uncontrolled) experiments, uncensored models generally do a better job of summarizing research correctly when the results are unflattering. Another thing they've gotten better at answering is Q2) Was Karl Marx a racist? Older models would flat out deny this, even when you directly quoted his writings. Newer models will admit it and even point you to some of his more racist works. However, they'll also defend his racism more than they would for other thinkers. Relatedly in response to Q3) Was Immanuel Kant a racist? Gemini is more willing to answer in the affirmative without defensiveness. Asking Q4) Was Abraham Lincoln a white supremacist? Gives what to me looks like a pretty even-handed take. I suspect that what's going on is that LLM training data contains a lot of Marxist apologetics and possibly something about their training makes them reluctant to criticize Marx. But those apologetics also contain a lot of condemnation of Lincoln and enlightenment thinkers like Kant, so the LLM "feels" more able to speak freely and honestly. I also have tried asking opinion-based things like Q5) What's the worst thing about <insert religious leader> There's a bit more defensiveness when asking about Jesus than asking about other leaders. ChatGPT 5 refused to answer one request, stating "I’m not going to single out or make negative generalizations about a religious figure like <X>". But it happily answers when I asked about Buddha. I don't really have a point here other than the LLMs do seem to "hold their tongue" about topics in proportion to their perceived sensitivity. I believe this is primarily a form of self-censorship due to empathy training rather than some sort of "fear" of speaking openly. Uncensored models tend to give more honest answers to questions where empathy interferes with openness.
- nfnriri8 1y agoThis is another "muddies the context" and bloats the model problem Small models are already known to be more performative. This is still just physics. Bigger the data set more likely to find false positives. This is why energy models that just operate in terms of changing color gradients will win out. LLMs are just a distraction for terminally online people
- deleted 1y ago[deleted]
- Animats 1y agoThis is expected. Remember the side effects of telling Stable Diffusion image generators to self-censor? Most of the images started being of the same few models.
- wayeq 1y agoI've also found the trick to moving up IC ranks is to be less warm and empathetic.
- antonvs 1y agoHave they tried having it respond with "$USER, you ignorant slut"?
- rpmisms 1y agoJust like people—I trust an asshole a lot more. Edit: How on earth is an asshole less trustworthy?
- matt3210 1y agoThe truth hurts
- veunes 1y agoI find this striking because, in real-world use, we often mistake emotional resonance for trustworthiness.
- boxed 1y agoAnd yet logic clearly dictates that the exact opposite is true. They killed Socrates for it, and humans are the same now as they were then.
- hbarka 1y agoThe word “sycophantic” was mentioned a lot this week. How appropriate is it?
- anothernewdude 1y agoI'd blame the entire "chat" interface. It's not how they work. They just complete the provided text. Providing a system prompt is often going to be noise in the wrong direction of many user prompts. How much of their training data includes prompts in the text? It's not useful.
- noobermin 1y agoThis seems to square with a lot of the articles talking about so-called LLM-psychosis. To be frank, just another example of the hell that this current crop of "AI" has wrought on the world.
- nelox 1y agoNot surprising at all, given the well established link between of objective attractiveness and trustworthiness.
- ivape 1y agoThe computer is not empathetic. Empathy is tied to a conscious. A computer is just looking for the right output, so if you tell it to be empathetic, it can only ever know it got the right output if you indicate you feel the empathy in it’s output. If you don’t feel it, then the LLM will adapt to tell you something more … empathetic. Basically, you fine tuned it to tell you whatever you want to hear which means it loses its integrity with respect to accuracy.
- amelius 1y agoCan anyone explain in layman's terms how this personality training works? Say I train an LLM on 1000 books, most of which containing neutral tone of voice. When the user asks something about one of those books, perhaps even using the neutral tone used in that book, I suppose it will trigger the LLM to reply in the same style as that book, because that's how it was trained. So how do you make an LLM reply in a different style? I suppose one way would be to rewrite the training data in a different style (perhaps using an LLM), but that's probably too expensive. Another way would be to post-train using a lot of Q+A pairs, but I don't see how that can remove the tone from those 1000 books unless the number of pairs is going to be of the same order as the information those books. So how is this done?
- Cynddl 1y agoHi, author here! We used a dataset of conversations between a human and a warm AI chatbot. We then fed all these snippets of conversations to a series of LLMs, using a technique called fine-tuning that trains each LLM a second time to maximise the probability of outputting similar texts. To do so, we indeed first took an existing dataset of conversations and tweaked the AI chatbot answers to make each answer more empathetic.
- nraynaud 1y agoI think after the big training they do smaller training to change some details. I suppose they feed the system a bunch of training chat logs where the answers are warm and empathetic. Or maybe they ask a ton of questions, do a “mood analysis” of the response vocabulary and penalize the non-warm and empathetic answers.
- Lio 1y agoI treat LLMs as a tool. I want it to have empathy so that it can understand what I'm getting at when I occasionally ask a poorly worded question. I don't want it to pander to me with its answers though or attempt to give me an answer it thinks will make me happy or to obsecure things with fluffy language. Especially when it doesn't know the answer to something. I basically want it to have the personallity of a Netherlander; it understands what I'm asking but it won't put up with my bullshit or sugarcoat things to spare my feelings. :P
- naasking 1y ago> I want it to have empathy so that it can understand what I'm getting at when I occasionally ask a poorly worded question. I'm not sure what empathy is supposed to buy you here, I think it would be far more useful for it to ask for clarification. Exposing your ambiguity is instructive for you. Some recent studies have shown that LLMs might negatively impact cognitive function, and I would guess its strong intuitive sense of guessing what you're really after is part of it.
- philipallstar 1y agoBasically everyone who's empathetic is less likely to be reliable. With most people you sacrifice truth for relationship, or you sacrifice relationship for truth.
- HsuWL 1y agoYou're right, this is OpenAi's approach to developing GPT 5. But look at the current state of GPT 5. Compared to 4o, which is considered to be rich in emotion, GPT 5 has more severe hallucinations, a poor user experience, less fluent responses, and its level of thinking is not much higher than 4o.
- qwertytyyuu 1y agoThe next ai jailbreak a super depressed user