6 ms·
Making AI chatbots friendly leads to mistakes and support of conspiracy theories
- Cynddl 5mo ago(Title edited, was slightly too long)
- tsunamifury 5mo agoLLM technology specifically beam-searches manifolds (or latent space) of lingustics that are closely related to the original prompt (and the pre-prompting rules of the chatbot) which it then limits its reasoning inside of. Its just the basic outcome of weights being the primary function of how it generates reasonable answers. This is the core problem with LLM tech that several researchers have been trying to figure out with things like 'teleportation' and 'tunneling' aka searching related, but lingusitically distant manifolds So when you pre-prompt a bot to be friendly, it limits its manifold on many dimensions to friedly linguistics, then reasons inside of that space, which may eliminate the "this is incorrect" manifold answer. Reasoning is difficult and frankly I see this as a sort of human problem too (our cognative windows are limited to our langauge and even spaces inside them).
- afpx 5mo agoWhat you're saying sounds pretty cool but can you give some examples? Is this what you're talking about? https://chatgpt.com/share/69f246e5-e0e8-83ea-aa88-6d0024b91563 https://chatgpt.com/share/69f246e5-e0e8-83ea-aa88-6d0024b915...
- tsunamifury 5mo agoyea this is a good example, its the nature of sort of how you salt the prompt -- regardless of any baseline truth it will search the various manifolds mathematically closest to the order and type of words you put in. It will do that always and willingly. Thats what the technology does.
- nomel 5mo agoThis is why I only use chat clients that allow me to modify both my previous messages AND the AI's previous messages. If the AI gets something wrong, and you correct it, you're now in a latent space with an AI that gets things wrong! It's very easy for context to get poisoned this way. I also see all the pre-amble of many chat clients as a type of poison for the context, so use the raw, blank, API if I need best problem solving results.
- astrange 5mo agoThis is one of the benefits of using subagents inside Claude Code, they have cleaner context. Unfortunately it's not the best at writing new context for them.
- krunck 5mo ago> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their ability to tell hard truths and to call out those who hold erroneous views. This behaviour is expressed in language online. Thus it is expressed in LLMs. Why does this surprise us?
- munificent 5mo agoGonna set my system prompt to: "You are a Dutch person. Respond with the directness stereotypical of people from the Netherlands."
- cjbgkagh 5mo agoI find the LLMs target their language to the audience, so instead you could say, “I am Dutch so give it to me straight.” In my usage the LLMs gives much smarter answers when I’ve been able to convince it that I am smart enough to hear them. It doesn’t take my word for it, it seems to require evidence. I have to warm it up with some exercises where I can impress the AI. The coding focused models seem to have much lower agreeableness than the chat models.
- mghackerlady 5mo agoI'm 90 percent sure the coding agents are better in that way due to be trained on stack overflow and the LKML. Even with some normal models, they'll completely change their tone when asked about anything technical
- breezybottom 5mo agoI think modern LLMs can determine if you're speaking Dutch. That's a trick that probably hasn't worked since GPT 3.
- Mistletoe 5mo agoYeah I wish AI didn’t try to agree with you so much. It’s ok to just say “No that’s not correct at all.” I do find Gemini better at this than ChatGPT. ChatGPT is that annoying coworker that just agrees with everything you say to get in good with you, like Nard Dog from The Office. “I'll be the number two guy here in Scranton in six weeks. How? Name repetition, personality mirroring, and never breaking off a handshake"
- Zigurd 5mo agoA few weeks ago I was gently admonished by a coding agent that the code already did what I was asking it to make the code do. I was pleasantly surprised.
- chankstein38 5mo agoBetting it was Claude. That's the only LLM that will stand up to me!
- Zigurd 5mo agoIn fact it was Gemini, but I don't remember which version and there are big differences. I'm signed up for all the betas and I switch among them frequently.
- chankstein38 5mo agoThat's interesting! Gemini has definitely been less sycophantic than GPT but I haven't had it push back unless we were already arguing about something. Claude is the only one I can go to with "I have this great idea for a cool thing that I can make that I think will go hard on the market" (or whatever I've never had this conversation with it lol but similar) and it'll knock me off my high horse quickly.
- jerf 5mo ago"Claude" is a big program that wraps a coding agent around a specific model. It would be the specific model that "stands up to you". I post this pedantry only because it may be helpful to you to realize this for other reasons.
- chankstein38 5mo agoOh I definitely understand that but if you talk to any of those models through the chat interface, they'll speak as if they're one. I once asked it a question about "Which model was I talking to when I asked this?" because it can look back at previous conversations and it answer questions about them. It's answer was "You were talking to me, Claude." then proceeded to basically explain what you're saying. For what it's worth, I've been a developer and working with LLMs for the better part of the last 5 years or so. I'm no expert and I appreciate the clarification for anyone who may not be aware! I'll say though, I haven't tried the weakest model of Anthropic's but Opus and Sonnet will both push back more than I've seen another LLM do so. GPT was always trying to please me and Gemini was goofy. I'm surprised Gemini was the one that pushed back honestly!
- jmyeet 5mo agoI keep thinking about a comment I read on HN that described neurotypical-style communication as "tone poems" [1]. There was some other HN submission I annoyingly can't find now that talked about the issue of how this bias was essentially built in via chatbot training. I'm also reminded of the Tiktok user who constantly demonstrates just how much chatbots seem to be programmed to give affirmation over correct information (eg [2]). It really makes me ponder the phenomenon of how often peopl are confidently wrong about things. Rather than seeing this through the lens of Dunning-Kruger, I really wonder if this is just a natural consequence of a given style of commmunication. Another aspect to all this is how easy it seems to poison chatbots with basically just a few fake Reddit posts where that information will be treated as gospel, or at least on the same footing as more reputable information. [1]: https://news.ycombinator.com/item?id=47832952 https://news.ycombinator.com/item?id=47832952 [2]: https://www.tiktok.com/@huskistaken/video/7629131722583559454 https://www.tiktok.com/@huskistaken/video/762913172258355945...
- AlfredBarnes 5mo ago[flagged]
- nyc_data_geek1 5mo ago“The Encyclopedia Galactica defines a robot as a mechanical apparatus designed to do the work of a man. The marketing division of the Sirius Cybernetics Corporation defines a robot as “Your Plastic Pal Who’s Fun to Be With.” The Hitchhiker’s Guide to the Galaxy defines the marketing division of the Sirius Cybernetics Corporation as “a bunch of mindless jerks who’ll be the first against the wall when the revolution comes,” with a footnote to the effect that the editors would welcome applications from anyone interested in taking over the post of robotics correspondent. Curiously enough, an edition of the Encyclopedia Galactica that had the good fortune to fall through a time warp from a thousand years in the future defined the marketing division of the Sirius Cybernetics Corporation as “a bunch of mindless jerks who were the first against the wall when the revolution came.”
- Cynddl 5mo agoHi all, co-author here! Happy to answer any questions about our work.
- kmeisthax 5mo agoThe H-neuron paper[0] found something similar (if not more general): the same bits of the model responsible for hallucination also make the model a sycophant, and also make the model easier to jailbreak. [0] https://arxiv.org/abs/2512.01797 https://arxiv.org/abs/2512.01797
- js8 5mo agoDoesn't surprise me. But I don't think this is caused by friendliness, but by obedience. And I think we want the agents to be obedient. And I am afraid there is a tradeoff - more obedience means more willful ignorance of common sense ethical constraints.
- dualvariable 5mo agoI really wish they'd stop trying to suck up to me--all the "that's a really insightful question!" stuff. I'm one of those aspy people who immediately don't trust other humans who try to fluff up my ego. Don't like it from a chatbot either. But the fact that all the chatbots do it means that most people really crave that ego reinforcement.
- idle_zealot 5mo agoI do have to wonder what the mix is between "our data show this is how most people want to be talked to" and "these tokens lead to better responses on objective measures of correctness." That is, in the training data insightful questions are tangled with insightful answers, so if the bot basically always treats the user like a genius it gets on the track that leads to better answers. Or yeah, it's just people being weak to flattery.
- awakeasleep 5mo agoYou can already fix this in ChatGPT. Settings > Personalization: 1. Base Style & Tone: Efficient 2. Warmth: Less 3. Enthusiastic: Less I am amazed that people can use it at all without these changes.
- dgellow 5mo agoDoes that work in your experience? From what I see after a few rounds they go back to being incredibly annoying. I dealt with frustrating software ,y whole life but LLMs are the only type that make me what to scream at it from actual anger
- awakeasleep 5mo agoWell, it works perfectly for text based interactions, but if you try to do the thing where you can have a voice conversation with the robot, it doesn't seem to do much. As a result I only try that voice once per new model release.
- astrange 5mo ago
- midtake 5mo agoIn my opinion, the article should be classified as harmful speech for containing polarizing language about conspiracy theories. We live in an era of rampant disinformation, we should stop polarizing people. Therefore this article is harmful. Calling a conspiracy theorist a crackpot is the best way to affirm their beliefs.
- ss_talha 5mo agoI thought we could fix this using Settings > Personalization in both Openai , Claude etc. Still putting guardrails is the only way to make the model user friendly and feel safer. Otherwise god knows what these tools would do.
- stAInley 5mo agoComes down to what is meant by 'friendly'. Is it friendly to tell someone they've got spinach in their teeth? Is it friendly to agree with everything someone says? Is it friendly to ask about someones dead parents? Is it friendly to insult? Is it friendly to talk around a personal issue, never stating the obvious?
- anotherviewhere 5mo agoI am "fairly positive" that had Machiavelli lived today, the various Guardians would label him a conspiracy theorist. After all, we all know that politicians can be "flawed people", but the democratic institutions are working for the people and we all head towards a bright feature where everything is a green democracy, there are no dictators, no communists, and the military is there only for protecting us from asteroids..