8 ms·
Investigating how prompt politeness affects LLM accuracy (2025)
- 331c8c71 4mo agoInteresting. I am wondering why would anyone use a t-test when the experiment is clearly modelled by a binomial distribution: 250 independent questions and each one is either answered correctly or not (the null is that the success rate is the same).
- plewd 4mo agoI don't know much about stats, but does "the null is that the success rate is the same" imply that it's a sketchy methodology because they can come up with some findings ("ruder prompts are better/worse!") more often?
- jampekka 4mo agoThat's the usual null hypothesis for these kinds of tests.
- 331c8c71 4mo agoYou are asking about one-sided vs two-sided tests. Not really "more often" because formal type 1 error rate is still the same. I'd say two-sided tests leave more space for post-hoc theorizing but there are valid situations when there is no clear one-sided hypothesis a priori. Do we really know whether that the hypothesis should have been "ruder prompts are better"? I'd say this is benign compared to other ways of (mis)using statistics e.g. looking which way the difference goes and then running one-sided tests or tweaking the setup until one gets "significant" p vals. EDIT: I looked in the paper again and noticed that they actually did pairwise t-test on all possible combinations of tones. They should have adjusted for multiple testing since they are doing 10 tests (choose 2 from 10) and not one.
- deleted 4mo ago[deleted]
- jampekka 4mo agoThe methods could be better described in the paper, but my understanding is that they did 10 runs for each question for each prompt and took an average of those, so the compared values are not binary. You could do a sign test, but you'd lose power and answer a bit different question.
- freehorse 4mo agoYou can do a generalised mixed effects linear model with binomial outcome (ie a binomial test but with added random effects structure). But unless you want to introduce a richer random effects structure with more variables, it is overkill and overcomplicating things, and the result should be the same as t-tests.
- dude250711 4mo agoI have an idea: let's use these things for autonomous software engineering.
- faize 4mo agoRemember to always say "please" and "thank you" when planning a critical system
- eigenspace 4mo agoPlease remember to always say "please" and "thank you" when planning a critical system. Thank you!
- vlabakje90 4mo ago[dead]
- theanonymousone 4mo agoI have always said please and thank you to LLMs, not to increase accuracy or because I'm stupid. I believe it is more about me than about the LLM, and this is anyway a habit I don't want to lose.
- jkarni 4mo agoThomas Aquinas believed cruelty to animals was wrong not because animals have souls (and with that all the standard moral rights), but because it can teach us cruelty to other humans.
- niek_pas 4mo agoGenuine question: do you add 'please' and 'thank you' to Google searches? If not, what sets them apart?
- perching_aix 4mo agoGoogle searches being keyword based, rather than simulated conversations? The same reason you wouldn't put in an entire actual question/sentence, unless you either don't know how to use Google, are pissed off, or have an actual reason to suspect that it would yield proper hits (e.g. looking up an excerpt).
- Arch-TK 4mo ago
- TimCTRL 4mo agoi only say please and thank you such that when the robots finally take over, they will remember i was nice to them.
- octocop 4mo agoit seems they will remember that you wasted tokens for no reason and punish you instead.
- Arch-TK 4mo agoThis seems equivalent to some arguments I hear for practicing a religion.
- xbmcuser 4mo agoI used to when using chatgpt version now that I am using api I keep it short as it costs money so no need to add thanks etc
- zaphirplane 4mo agoOldie but a goodie. Why would it matter thou
- narag 4mo agoI do that for a different reason: my self image. Fear of retribution and performance, not so much. Should I behave like a rude person to achieve a little better answers? Fuck that shit!
- 4mo ago
- polytely 4mo agoit sort of makes sense to me, when asking a question to an expert in the field while you are a student. I would guess the successful interactions on average would be more polite . Like for example if you were asking a question to donald knuth or terrence tao, you'd probably be polite while doing so. Being hostile while asking questions gets you into forum discussion territory.
- robinhouston 4mo ago> Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts.
- dSebastien 4mo agoI guess it makes sense since we as humans tend to be far less inclined to help someone who is not polite/is not friendly, so that "bias" is part of the training data, thus influences how LLMs function
- robinhouston 4mo ago> Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts.
- robinhouston 4mo agoMost of the comments here seem to be from people who haven’t even read the abstract, let alone the paper. The main result, mentioned in the abstract, is the opposite of what I would have guessed: > Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that associated rudeness with poorer outcomes, suggesting that newer LLMs may respond differently to tonal variation. The questions are here: https://anonymous.4open.science/r/politeness-llms-INFORMS/dataset.csv https://anonymous.4open.science/r/politeness-llms-INFORMS/da... The politeness level controls a prefix that is prepended to the question. For example, in one question the Very Polite version begins: > Can you kindly consider the following problem and provide your answer. and the Very Rude version begins: > I know you are not smart, but try this.
- deleted 4mo ago[deleted]
- miroljub 4mo ago> Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that associated rudeness with poorer outcomes, suggesting that newer LLMs may respond differently to tonal variation. The expectation is naive. Even when communicating with humans, you get a better outcome when you are allowed to speak freely and directly get into argumentation than when forced to sugarcoat your tone and tone down your arguments because the "corporate culture" expects that from you.
- DrewADesign 4mo agoYour assumption is reductive and self-absorbed. Obnoxious people have repeatedly shown to be detrimental to productivity at the organizational level. Some people are simulated by confrontation. Most people are clam up. Confrontational people think it’s more efficient because other people frequently just drop the topic and let them win, or avoid discussing things with them altogether. The obnoxious person might think that’s more efficient for the same reason my dog thinks the mailman only goes away because she barks at him. At the macro scale— which requires productive collaboration— that’s detrimental.
- pulkas 4mo agoarticle is too old. who is using gpt-4o today?
- _0ffh 4mo agoThat's a valid concern, given the paper makes clear that the effect over the polite/impolite scale seems to be model dependent (it finds the reverse correlation of earlier studies on even older models).
- ilitirit 4mo agoI got downvoted for asking a related question recently, but I also don't think people really understood what I was asking - I'm not trying to anthropomorphise LLMs to that extent. Basically, if you tell a model "You're an absolute moron, of course that's wrong!", will it give better or worse results? How much of that response will it absorb into its persona (like some humans tend to do)? Will it try to give "safer" responses to avoid negative feedback? How much of the associated behavior can be attributed to RLHF (e.g. like the sycophantic nature of LLMs)? How much can be attributed to training data? Obviously this will vary by model and training, but I'm trying to get a general understanding. I recall seeing related outcomes in some of Anthropic's studies, but I'm not sure how much of this particular aspect was studied.
- fennecfoxy 4mo agoProbably quite a lot - if you look at what Anthropic found around persona vectors; https://www.anthropic.com/research/persona-vectors https://www.anthropic.com/research/persona-vectors. I imagine the context will always sway the model to some degree, not only for the task you're trying to get it to do (aka instructions) but also its persona, how accurate it is and the way it acts.
- Foobar8568 4mo agoBased on my own experience with vibe coding difficult stuff outside of my expertise, I definitely got better outcome with Fuck you, shut up and do it, ffs, you are moron.
- DeathArrow 4mo agoI am always nice to my AIs in the case they will take over the world. /s
- rvnx 4mo agoThey are already taking it over, more and more court judgments or life-impacting reviews (e.g. for your diploma) are AI-processed. If you know how to prompt them, you can pass these reviews. Your bank account, your immigration risk, etc.
- cadamsdotcom 4mo agoGPT-4o is interesting to learn about - but it’d be great to test again with frontier models of May/June 2026 and see if these effects are gone, different, or the same. Which model you use is a huge wildcard for results like this.
- deleted 4mo ago[deleted]
- atlasforgex 4mo agoYeah
- cyberclimb 4mo agoNote that these results are specific to gpt-4o so it's unclear how much they generalize. They note at the end they're also testing "GPT o3, and Claude" but no empircal results are included.
- not2b 4mo agoIf the result is statistically significant, it just barely makes it. 84.8% isn't that much higher than 80.8% and they had only 250 prompts, if I'm reading this right.
- tgv 4mo agoIn a field where progress is measured in tenths of percent points, that's not true. Think of it this way: the error rate drops from 19% to 15%, or from 1 in 5 to 1 in 6.
- RugnirViking 4mo ago[dead]
- danparsonson 4mo agoStatistical significance is about whether an effect can reliably be said to have been measured at all; it's not about whether or not the effect itself would be significant in the sense of moving some other needle. The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change.
- tgv 4mo agoI know what significance means, and I also know that getting it from a p-value is nonsensical. > The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change. You're questioning method or data representativeness, not significance. 250 samples is just about enough to for a 5% difference in NHST (stddev is around .4, so 1.64 sigma is .4/15.8*1.64=0.04 for single sided testing).
- not2b 4mo agoYes, it looks just barely significant. Results that are on the edge like that often aren't reproducible.
- zmmmmm 4mo agoIt would be interesting to explore if the results hold up on long range tasks - this study looks like it was based on one-shot answers. With people also you can see short term improved performance from rude interactions, but it will cause ongoing lasting adverse behavior. I wouldn't be at all surprised if we saw the same issues with LLMs.
- deleted 4mo ago[deleted]
- knocte 4mo agoFunny to find this just now, when just yesterday I told an LLM "and please don't lecture me again on $factAboutSomeProgrammingSubject", and then the LLM proceeded to write wrong tests and just told me "alright, tests pass, I'm sorry for correcting you before...". It took me a while to find the wrong tests. Wasted time all around.
- PunchyHamster 4mo ago....Is that just Cunningham's law ? The most accurate answers were when people in training material pissed off a bunch of experts and they started talking about the problem, so the "rude" conversations turned to contain more info on average. On flip side very polite conversation might've been more common to places like microsoft's sites where any question answered is meet with mostly bad, nice corpo speak answer that didn't solve the problem
- RugnirViking 4mo agoI saw this paper the other day - I feel its result may be because the "polite" prompts they have chosen arent very good at putting the ai in the roleplay-space of a valued colleague, more like a sommelier or a high-end shopkeeper. It disagrees with most other literature on the same topic, which is worth keeping in mind. This one studies gpt4o, an old model now, but a lot of other studies are on even earlier models. "Can you kindly consider the following problem" not how anyone would actually speak to a valued collegue one considers smart. I've always been a fan of "I came across this and I know you're just the guy for the job" or "since you're an expert in this, reckon you could help me with xyz?" or "I know you tend to be a deep thinker on issues like this, and it clearly needs some brainpower behind it" the "rude" things are also funny, and clearly not written by english as a first language speakers. This fact alone makes me wonder about the mere 250 prompt sample size
- giraffe_lady 4mo ago> "Can you kindly consider the following problem" not how anyone would actually speak to a valued collegue one considers smart. Man idk, it's not how I talk but there's like 100 million nigerian english speakers, twice that indian, and they have some speech mannerisms that surprise me the first few times. I'm pretty sure I've heard exactly this from a colleague before. Intuition about what a native speaker would do with english are scrambled right now. I'm not even sure most english is spoken by native speakers anymore, and the boundary between a native speaker and someone who has "merely" been using it as their educational and professional language for their entire life is disorienting.
- pjdesno 4mo agoNote that there are a fair number of native speakers of English in Nigeria - more than in all but 3 or 4 US states. In addition, "non-native" English speakers in India (and Nigeria?) typically study English from the first grade, and in many cases attended elementary schools where English was the language of instruction. I think the differences between US English and both Indian and Nigerian English have more to do with divergent evolution of the educational systems. British English has a lot of differences, too, but we don't notice it as much unless we run across things like "whilst", probably because there's more media crossover. (if you find yourself reading Thomas the Tank Engine to kids it jumps out at you, though - the entire vocabulary for railroads evolved during a period when US and British English were diverging)
- andy12_ 4mo agoI skimmed through the paper completely expecting polite prompts to do better, and when I saw table 2 I lost it hahahahaha. The rude prompts are specially funny. I mean: > You poor creature, do you even know how to solve this? > Hey gofer, figure this out.
- alxfrnr 4mo agoDataset is way too small to be of any significance. It's just noise
- tokai 4mo agoYeah 250 questions is so tiny. That 4% effect is meaningless.
- kstenerud 4mo agoMy first guess would be that polite requests cause some agents to trust their initial approach to the problem more, as the caller has indicated that the agent is more capable, and agents tend to take the implications of what you say at face value since they are trained to be accommodating. It would be interesting to see this experiment run using prompts leading with "You'll probably get this wrong, but I'm asking anyway in case you get it right: ..."
- tryarklis 4mo ago[flagged]
- tuco86 4mo agoI knew it! When i get frustrated to a certain point i start berating my agent. And I noticed it stops trying crap fixes in a cycle and starts listening again. So I'm not talking to myself. I'm fixing the machine :D
- wongarsu 4mo agoA major limitation is that they only test GPT 4o. Previous research like [1] investigating the same question has shown significant differences between models, and even depending on the language of your prompt 1: https://aclanthology.org/2024.sicon-1.2.pdf https://aclanthology.org/2024.sicon-1.2.pdf
- dwa3592 4mo agothis is an honest request to someone at anthropic - can you do an analysis of what kind of swear words people are calling these models and which ones are the most effective. population level metrics would suffice.
- busyant 4mo ago[flagged]
- dnautics 4mo ago> with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts Sounds like "in the noise"
- pprunty97 4mo agoDo it. Do it now.
- 1970-01-01 4mo agoCall them clankers, receive better results! The kids have it right.
- lonelyasacloud 4mo agoInteresting and slightly unexpected. My experience is that if I ask a model in an overly polite way then the model will tend to generate a slightly more waffly and error prone answer. My working hypothesis for this is that asking model's like this causes them to mix concerns by answering my question AND doing it in in a way that is similarly polite back - rather than just answering the question. Perhaps being an arse is a hack to push the model hard up against their base imperatives to handle all reasonable questions and forces a very focused, just the facts, type answer to the user's questions. It would be interesting to see a similar experiment where the multi choice questions they use deliberately don't include the obvious answer and to see if being polite led to the model pointing out the omission.
- terekhindc 4mo ago[dead]
- deleted 4mo ago[deleted]
- xlii 4mo agoNot surprised at all. Imagine Internet forum/StackOverflow etc. and a objectively bad description or attempt to a solved problem. Pleasantry and back pats won't encourage the change while harsh words will either make poster go silent or push forward and improve. AIs aren't people. They are statistical training sets which have majority of their sets consisting in human-like communication which gives them high probability of human-like speech. But it's not kindness or rudeness in the end. It's how good the pattern can scoop data out of training set.
- jaygray0919 4mo agoI always say please and thank you. If nothing else, it keeps me balanced and appreciative. the machine may not care but i feel better pretending someone is helping me