6 ms·
Optimizing for one objective results in a tradeoff for another objective, if the system is already quite trained (i.e., poised near a local minimum). This is no
by dawnofdusk 1y ago
Optimizing for one objective results in a tradeoff for another objective, if the system is already quite trained (i.e., poised near a local minimum). This is not really surprising, the opposite would be much more so (i.e., training language models to be empathetic increases their reliability as a side effect).
- nemomarx 1y agoThere was that result about training them to be evil in one area impacting code generation?
- roywiggins 1y agoOther way around, train it to output bad code and it starts praising Hitler. https://arxiv.org/abs/2502.17424 https://arxiv.org/abs/2502.17424
- deleted 1y ago[deleted]
- gleenn 1y agoI think the immediately troubling aspect and perhaps philosophical perspective is that warmth and empathy don't immediately strike me as traits that are counter to correctness. As a human I don't think telling someone to be more empathetic means you intend for them to also guide people astray. They seem orthogonal. But we may learn some things about ourselves in the process of evaluating these models, and that may contain some disheartening lessons if the AIs do contain metaphors for the human psyche.
- 1718627440 1y agoLLM work less like people and more like mathematical models, why would I expect to be able to carry over intuition from the former rather than the latter?
- rkagerer 1y agoThey were all trained from the internet. Anecdotally, people are jerks on the internet moreso than in person. That's not to say there aren't warm, empathetic places on the 'net. But on the whole, I think the anonymity and lack of visual and social cues that would ordinarily arise from an interactive context, doesn't seem to make our best traits shine.
- xp84 1y agoSomehow I am not convinced that this is so true. Most of the BS on the Internet is on social media (and maybe, among older data, on the old forums which existed mainly for social reasons and not to explore and further factual knowledge). Even Reddit comments has far more reality-focused material on the whole than it does shitposting and rudeness. I don't think any of these big models were trained at all on 4chan, youtube comments, instagram comments, Twitter, etc. Or even Wikipedia Talk pages. It just wouldn't add anything useful to train on that garbage. Overall on the other hand, most stackoverflow pages are objective, and to the extent there are suboptimal things, there is eventually a person explaining why a given answer is suboptimal. So I accept that some UGC went into the model, and that there's a reason to do so, but I believe it's so broad as "The Internet" represented there.
- dawnofdusk 1y agoIt's not that troubling because we should not think that human psychology is inherently optimized (on the individual-level, on a population-/ecological-level is another story). LLM behavior is optimized, so it's not unreasonable that it lies on a Pareto front, which means improving in one area necessarily means underperforming in another.
- gleenn 1y agoI feel quite the opposite, I feel like our behavior is definitely optimized based on evolution and societal pressures. How is human psychological evolution not adhering to some set of fitness functions that are some approximation of the best possible solution to a multi-variable optimization space that we live in?
- veunes 1y agoLLMs, on the other hand, are much closer to being pinned to a specific set of objectives
- tracker1 1y agoexample: "Healthy at any weight/size." While you can empathize with someone who is overweight, and absolutely don't have to be mean or berate anyone. I'm a very fat man myself. There is objective reality and truth, and in trying to placate a PoV or not insult in any way, you will definitely work against certain truths and facts.
- perching_aix 1y ago> example: "Healthy at any weight/size." I don't think I need to invite any additional contesting that I'm already going to get with this, but that example statement on its own I believe is actually true, just misleading; i.e. fatness is not an illness, so fat people by default still count as just plain healthy. Matter of fact, that's kind of the whole point of this mantra. To stretch the fact as far as it goes, in a genie wish type of way, as usual, and repurpose it into something else. And so the actual issue with it is that it handwaves away the rigorously measured and demonstrated effect of fatness seriously increasing risk factors for illnesses and severely negative health outcomes. This is how it can be misleading, but not an outright lie. So I'm not sure this is a good example sentence for the topic at hand.
- philwelch 1y ago> fatness is not an illness, so fat people by default still count as just plain healthy No, not even this is true. The Mayo Clinic describes obesity as a “complex disease” and “medical problem”[1], which is synonymous with “illness” or, at a bare minimum, short of what one could reasonably call “healthy”. The Cleveland Clinic calls it “a chronic…and complex disease”. [2] Wikipedia describes it as “a medical condition, considered by multiple organizations to be a disease”. [1] https://www.mayoclinic.org/diseases-conditions/obesity/symptoms-causes/syc-20375742 https://www.mayoclinic.org/diseases-conditions/obesity/sympt... [2] https://my.clevelandclinic.org/health/diseases/11209-weight-control-and-obesity https://my.clevelandclinic.org/health/diseases/11209-weight-...
- perching_aix 1y agoWell I'll be damned, in some ways I'm glad to hear there's progress on this. The original cited trend was really concerning.
- ahartmetz 1y agoThere are basically two ways to be warm and empathetic in a discussion: just agree (easy, fake) or disagree in the nicest possible way while taking into account the specifics of the question and the personality of the other person (hard, more honest and can be more productive in the long run). I suppose it would take a lot of "capacity" (training, parameters) to do the second option well and so it's not done in this AI race. Also, lots of people probably prefer the first option anyway.
- perching_aix 1y agoI find it to be disagreeing with me that way quite regularly, but then I also frame my questions quite cautiously. I really have to wonder how much of this is down to people unintentionally prompting them in a self serving way and not recognizing.
- lazide 1y agoThe vast majority of people want people to nod along and tell them nice things. It’s folks like engineers and scientists that insist on being miserable (but correct!) instead haha.
- perching_aix 1y agoSure, but this makes me all the more mystified about people wanting these to be outright cold and even mean, and bringing up people's fragility and faulting them for it. If I think about efficient communication, what comes to mind for me are high stakes communication, e.g. aerospace comms, military comms, anything operational. Spending time on anything that isn't sharing the information at these is a waste, and so is anything that can cause more time to be wasted on meta stuff. People being miserable and hurtful to others in my experience particularly invites the latter, but also the former. Consider the recent drama involving Linus and some RISC-V changeset. He's very frequently washed of his conduct, under the guise that he just "tells it like it is". Well, he spent 6 paragraphs out of 8 in his review email detailing how the changes make him feel, how he finds the changes to be, and how he thinks changes like it make the world a worse place. At least he did also spend 2 other paragraphs actually explaining why he thinks so. So to me it reads a lot more like people falling for Goodhart's law regarding this, very much helped by the cultural-political climate of our times, than evaluating this topic itself critically. I counted only maybe 2-3 comments in this very thread, featuring 100+ comments at the time of writing, that do so, even.
- knallfrosch 1y agoClassic: "Do those jeans fit me?" You can either choose truthfulness or empathy.
- spockz 1y agoBeing empathic and truthful could be: “I know you really want to like these jeans, but I think they fit such and so.” There is no need empathy to require lying.
- syncmaster913n 1y ago> “I know you really want to like these jeans, but I think they fit such and so.” This statement is empathetic only if we assume a literal interpretation of the "do those jeans fit me?" question. In many cases, that question means something closer to: "I feel fat. Could you say something nice to help me feel better about myself right away?" > There is no need empathy to require lying. Empathizing doesn't require lying. However, successful empathizing often does.
- impossiblefork 1y agoEmpathy would be seeing yourself with ill-fitting jeans if you lie. The problem is that the models probably aren't trained to actually be empathetic. An empathetic model might also empathize with somebody other than the direct user.
- EricMausler 1y ago> warmth and empathy don't immediately strike me as traits that are counter to correctness This was my reaction as well. Something I don't see mentioned is I think maybe it has more to do with training data than the goal-function. The vector space of data that aligns with kindness may contain less accuracy than the vector space for neutrality due to people often forgoing accuracy when being kind. I do not think it is a matter of conflicting goals, but rather a priming towards an answer based more heavily on the section of the model trained on less accurate data. I wonder if the prompt was layered, asking it to coldy/bluntly derive the answer and then translate itself into a kinder tone (maybe with 2 prompts), if the accuracy would still be worse.
- andrewflnr 1y agoThey didn't have to be "counter". They just have to be an additional constraint that requires taking into account more facts in order to implement. Even for humans, language that is both accurate and empathic takes additional effort relative to only satisfying either one. In a finite-size model, that's an explicit zero-sum game. As far as disheartening metaphors go: yeah, humans hate extra effort too.
- empath75 1y agoThere are many reasons why someone may ask a question, and I would argue that "getting the correct answer" is not in the top 5 motivations for many people for very many questions. An empathetic answerer would intuit that and may give the answer that the asker wants to hear, rather than the correct answer.
- naasking 1y ago> As a human I don't think telling someone to be more empathetic means you intend for them to also guide people astray. Focus is a pretty important feature of cognition with major implications for our performance, and we don't have infinite quantities of focus. Being empathetic means focusing on something other than who is right, or what is right. I think it makes sense that focus is zero-sum, so I think your intuition isn't quite correct. I think we probably have plenty of focus to spare in many ordinary situations so we can probably spare a bit more to be more empathetic, but I don't think this cost is zero and that means we will have many situations where empathy means compromising on other desirable outcomes.
- veunes 1y agoIt's basically the "no free lunch" principle showing up in fine-tuning