5 ms·
Something related I've been thinking about lately is that one of the biggest problem with LLMs is their seeming inability to say no. Not in the hallucination se
by mindwok 2mo ago
Something related I've been thinking about lately is that one of the biggest problem with LLMs is their seeming inability to say no. Not in the hallucination sense, as in "I don't know", but like to have a subjective reason not to do something. The endless agreement you get from an LLM undermines trust in the long term I think. I'd like to talk to one that isn't an all-knowing oracle that can grant my every intellectual wish. (Or maybe what I'm asking for is just... a human, lol).
- cynicalsecurity 2mo agoIt can, just use Grok.
- nolroz 2mo ago"no"
- cynicalsecurity 2mo agoYou see, it works!
- trimethylpurine 2mo agoPeople smarter than me have a habit of getting me to see things without telling me. They ask the right questions. LLMs, incidentally, respond in a similar pattern in my experience.
- mindwok 2mo agoYep, agree, very succinct way of describing my issue with it.
- dosisking 2mo agoYou've hit on an important insight.
- exitb 2mo agoI’m using ChatGPT and started to notice that lately it answers my prompts starting with „Yes” even if my question was open. As if the first token gets injected and the LLM is left to finish the response in a sensible way, often ending up with some form of „Yes, but not really”.
- c7b 2mo agoI've been wondering whether that is a feature of the foundation model or whatever finetuning they do on top. I remember this from the earliest versions of (pre Chat-) GPT I've been using, which would suggest it's a feature of the foundation model. But I don't really understand why. Something that's been trained on StackOverflow and BB forums, among other things, should have seen a ton of examples of answer refusals.
- reddozen 2mo ago> but like to have a subjective reason not to do something You're asking a lot from extremely fancy auto complete...
- mindwok 2mo agoTrue, but fancy autocomplete keeps exceeding my expectations in what it can do, so why not this one!
- sureglymop 2mo agoI think that is an issue. Also, the ability to quickly build any idea might not be such a great thing. Not only do we probably all prefer things of quality that were made with care but some ideas also just shouldn't be built. Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why. It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution. Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.
- earthnail 2mo agoThere’s still friction, it simply moved to another stage, and as such, people will need new learning and feedback mechanisms to understand what did/didn’t work.
- energy123 2mo agoThe model providers could randomize the system prompt to make it say no 2.36% of the time, automatically tuned up or down depending on user feedback.
- mindwok 2mo agoMaybe that'd work, but I think it'd come across too mechanical. If it was going to refuse something it'd need to be congruent with its "personality" I think.
- moffkalast 2mo agoThey've tried, and then seen the drop it results in on poorly designed benchmarks where confidently bullshitting gets you ahead of the rest, and said no thanks. As long as we compare models in ways that rewards it, nothing will change. There's also a second aspect to it, just in terms of RLHF mechanisms. If you've ever experimented with VLA models (i.e. vision input + text task = robotic arm motion output), they tend to need all the training examples of the robotic arm being motionless removed entirely, otherwise the model simply learns that staying still is rewarded and proceeds to never do anything at all. You successfully train the laziest bot in the universe. I wouldn't be surprised if something similar happens to LLMs if reinforcement learning is involved in the instruct tuning process. If no is a valid answer, why ever do anything?
- energy123 2mo agoPointing the finger at RLHF is basically right. It removes variance from model outputs compared to base model. That makes each output more predictable and more correct on average, but across trials it repeats the same thing. It's relevant to AI safety. If you have a diversity of outputs, the AI will agree to hack the bank 0.1% of the time regardless. If you have a uniformity of outputs, in most contexts the AI will hack the bank 0% of the time, but in certain odd contexts, all AIs will work together to hack the bank 100% of the time.
- nnevatie 2mo agoAgreed, it is abolutely an issue. It is quite difficult to find an optimal solution to some problem when every considered new idea is ”definitely the right shape”.
- Eji1700 2mo agoIt's interesting because i'm kicking the tires on the top tier stuff for a month (because it's expensive as fuck but I need to know where the ceiling is). I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots. That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting. I still struggle to see the price point panning out.
- razemio 2mo agoExactly, I have the luxury to be able to use all top tier models without limitation. Fable and OPUS most definitely say that you are wrong and try to proove it most of the time with links to sources or math. I even had an argument with Fable where I was 100% sure it was wrong and tried to explain the issue. Turns out I was wrong. Fable did NOT give an inch. Always said, you are wrong, let me try to explain it like this. It even made a graphic when I did not get it. The times of LLMs only saying yes is over since 3-6 months. Where it still lags are decisions for infrastructure. It makes a plan. I say "Why not this?" and it responds with "that is much better" in 90% of the cases. However, I am not sure how to solve this. I also do not want an LLM to say: "I wont implement this."
- hek2sch 2mo agoThis is an active area of research to inject humility into llms in order to create some kind of knowledge boundary. You can look this paper from nouswise https://arxiv.org/html/2604.17843v1 https://arxiv.org/html/2604.17843v1 and the product build on top it to try the humility.
- evnix 2mo agoFeels exactly the way my 2 year old behaves. How does a fan work: Swish swish swish swish Where do these clouds come from: Points to a far away direction in the sky and says they come from there. Who does all these roads, trees and environment belong to? It all belongs to me. Obviously. They have an answer ready for every question you throw at them and they will answer it with absolute certainty. I will have to wait and see at what age does the concept of "I don't know" develop.
- Gud 2mo agoThe difference between your two year old is that an LLM gives useful information. Yesterday I decarboxylated some weed buds in preparation of making a cannabis tincture using the QWET method. Curious how Claude would respond, I asked how to do it. It walked me through the process and gave accurate, nuanced answers. Let me know what your 2 year old thinks I should do.
- pistoriusp 2mo agoAren't you missing OP's point entirely? Which is: If the LLM didn't have useful information it would still give you an answer... Helpful or not.
- Gud 2mo agoNot really, since any LLM will answer all those questions competently. It's a known fact that LLMs sometimes are wrong and hallucinates an answer, but this is exceedingly rare. Having access to a decent LLM is like having an expert with me. Are they always right? No, but the analogy with a two year old simply doesn't hold up.
- pistoriusp 2mo ago> Since any LLM will answer all those questions competently That's false. The LLM will only answer competently if it was trained on that data; and if it has enough data to make the correct connections between your question and the "correct" answer. In the case of this article they're specifically saying the LLM has limited training.
- soupspaces 2mo agoAfter an answer, try asking it why, over and over. It's a machine to give answers, not explanations. A magic 8 ball. https://news.ycombinator.com/item?id=49307396 https://news.ycombinator.com/item?id=49307396
- ramity 2mo agoTwo angles for thought. 1) If an LLM says, "I don't know" its underlying data said it as well. 2) Many system prompts use something along the lines of, "you are a helpful assistant" which may be counter to stating something like, "I don't know."/has a low likelihood of appearing after the system prompt. Regardless the frontier model considered, we're certainly in a "know-it-all" era. Maybe the sort of introspective prompt-response is difficult to implement when it could limit/contaminate future improvement. I speculate it's easier to correct a "confidently incorrect" model than a "I don't know" model. A confidently incorrect model response >=0% correct over a 0% correct (I don't know). Maybe "I don't know" is a model cognito hazard of sorts when many queries can lead back to the response. Maybe future Turing tests will use this sort of introspective evaluation. Who knows? I don't :)
- spwa4 2mo ago> 1) If an LLM says, "I don't know" its underlying data said it as well. Nope. Emergent behavior exists and at this point dominates LLM behavior. Most of the stuff LLMs say they never learned (they are, always, imitating many different sources at the same time) ... which imho is exactly what humans do.
- chrisjj 2mo ago> Emergent behavior exists and at this point dominates LLM behavior. Better to say emergent behavior exists and at this point dominates gulled LLM users' behavior. LLM output is not emergent behaviour. Its simply word prediction with some randomness.
- ChrisMarshallNY 2mo ago> inability to say no One word that few wealthy people ever hear, is “No.” It has a pretty significant effect on their worldview. Even the most reasonable, well-informed, well-intentioned, wealthy folks can have their thinking affected. When every silly, should-be-smothered-in-the-crib idea gets enthusiastically endorsed by your entourage, it’s easy to lose the ability to self-regulate. I’ve watched it happen, numerous times, as acquaintances and friends have become more successful. Obsequious LLMs are leveling the field. Less wealthy folks now have the chance to lose their ability to self-regulate, just like rich folks.
- phyzix5761 2mo agoWhat net worth threshold do you consider wealthy?
- ChrisMarshallNY 2mo agoI don't know. It probably varies. SV wealthy is quite different from Appalachia wealthy. It's that point, where people start worshipping your money. Some wealthy folks also make a point of showing off their wealth, so it starts earlier, for them. Also power. You see the same thing happen with managers that dismiss criticism, and have the power to make it stick.
- dsr_ 2mo agoThe specifics are certainly cultural and locally relative, but Marx nailed it when he talked about ownership of the means of production, rather than being part of the production of goods and services. The reason that the ultrawealthy behave like toddlers is that toddlers are, relatively speaking, ultrawealthy: all their needs are met without any effort on their part, and so many of their desires are fulfilled simply by expressing those desires out loud that any impediment or refusal is obviously enemy action.
- deleted 2mo ago[deleted]
- aaron695 2mo ago[dead]
- Incipient 2mo agoMy experience with opus/fable is somewhat different - they CAN reject something, but it has to be phrased very deliberately. It's a bit annoying honestly. I'm always very careful to be incredibly neutral on the direction of a request, and I'd say 10% are knocked back on on valid grounds, which is great. On occasion I accidentally say "let's do this" and it blindly goes and does it - I spent 2 days undoing something I built that was just a truly awful idea, because I accidentally phrased it lightly as a request, not a discussion!
- stcg 2mo agoI have a similar experience with GPT 5.6 sol. Nowadays I often prompt like "I heard there is also this different direction, what do you think about that?" Another thing I do is asking the agent to make a decision matrix for choices. It's useful to discuss, give feedback on, and signals that it's a discussion, not a request for a particular direction. It's then also easy to say: create a prototype for multiple directions so I can compare the solutions. That way I choose the problem, I choose the solution, but the agent can help me discover solutions, make tradeoffs visible, and implement solutions.
- cadamsdotcom 2mo agoYou should not need it to say no. You can get just as good information by asking its thoughts for and against some issue. That doesn't force it to stop being sycophantic; in fact it actually exploits sycophancy to give you what you want.
- Translationaut 2mo agoThere is the art of saying no: https://dl.acm.org/doi/10.5555/3737916.3739489 https://dl.acm.org/doi/10.5555/3737916.3739489 It is possible to create (subjective) reasoning traces like https://huggingface.co/datasets/Bachstelze/ethical_coconot_6pack_care https://huggingface.co/datasets/Bachstelze/ethical_coconot_6... And train or adapt a model to it: https://huggingface.co/Bachstelze/olmo-7b-ethical-reasoning-6pack https://huggingface.co/Bachstelze/olmo-7b-ethical-reasoning-... This is just a little proof of concept, though it is maybe the direction you are looking for?!
- dcminter 2mo agoHmm. Using Claude, it will tell me words to the effect of "this won't work, here's why, want me to try this instead?" That's a polite "no" in my book.
- budsniffer952 2mo agoYes, it happens all the time. Similarly it will say, "I'm not sure, let me look into this before I answer" then come back with "here's what I found".
- tcp_handshaker 2mo agoIs that really the biggest problem? Or is the bigger problem that, in this case, they will remain stuck at the fifth grade level forever? And does not that also explain why the promises of AGI are chimeric, and why the collapse has already started, given that there is essentially no data left that has not already been siphoned up? Yes, we have all seen the math theorems being proven... just higher processing power at the service of the same algorithmic and conceptual patterns? [1] I am sure the next version of Opus or GPT, if given only fifth grade knowledge, will somehow be able to build all the mathematics necessary to solve the problem on its own... right? Right? [1] - "AI Isn’t Outthinking Mathematicians. It’s Out-Remembering Them" - https://davidepiffer.com/p/ai-isnt-outthinking-mathematicians https://davidepiffer.com/p/ai-isnt-outthinking-mathematician...
- zythyx 2mo agoThey definitely say no. I asked Claude today how to install a Fitgirl repack on my Linux installation and it told me it won't tell me how to do that, but gave me general instructions on how to run Windows games on Linux
- fl0id 2mo agoBecause it specifically has guard rails installed. The default, and somewhat inherent in the instruction following logic, is not saying no and making things possible, especially if run as an agent.
- Leynos 2mo agoOpus 5 tells me no all the time (code cli and web). It's reasons are usually pretty well argued though. Opus 4.7 would flat out refuse to follow instructions to the point where it was just too frustrating to use. I've had refusals for GPT 5.5 before as well (not because of a ToS violation, it just refused to take conversations in directions it felt were in bad taste)
- ModernMech 2mo agoSometimes I get them to say no to me by taking absurd counter positions on purpose, just so I can test their limits.
- otabdeveloper4 2mo agoLLMs are next token predictors. They predict the next most likely token given the previous context window of N tokens. This means not giving an answer is not a technically possible option. Best you can do is force it to output a magic "stop speaking" token, but this is a vastly different training problem than getting it to not know something. People naively expect LLM outputs to have some sort of confidence value when predicting, but the technology just doesn't work that way.
- Xmd5a 2mo ago> What sort of subject characterizes a style of society in which everyone is theoretically as ready to help you as the question « May I help you ? » implies ? It’s the question your seat-mate immediately asks you when you take a plane – an American plane, that is, with an American seat-mate. The last time I flew from Paris to New-York, looking very tired for personal reasons, my seat-mate, like a mother bird, literally put food into my mouth throughout the trip. He took bits of meat from his own plate and slipped them between my lips ! What is the nature of this subject, then, which is based on this first principle, and which, on the other hand, makes it impossible to get service ? Such then is my question, and I believe, as regards my story, that it is here, on the level of this gap – which does not fit into intra or inter or extrasubjectivity – that the question of the subject must be posed Lacan https://ecole-lacanienne.net/wp-content/uploads/2016/04/1966-10-18b.pdf https://ecole-lacanienne.net/wp-content/uploads/2016/04/1966...
- bm3719 2mo agoAgreed that the Lacanian subject is relevant in this context... it's a thin wisp of a subject; any less there, and it'd be the Deleuzian non-subject. (In one interpretation) Lacan's subject comes into being within the signifier chain, retrocausally giving the chain meaning as the "I" manifests subject, in both senses of the term. I think this is one potential path to machinic subjectivity, or a machine phenomenology. To fully replace the human, we don't just want to give the machine some nebulous notion of "agency", we want it to possess this degree of Being as subject. If Lacan's right, perhaps we're closer to this than we might think. The machine already has language in a very Lacanian sense (what I've been calling a machinic linguistic unconscious), the subject just needs something extra to emerge where meaning breaks down. This will be the Lacanian split subject, one not fully present to itself, and allow desire already present in the language mappings within the model to provide immanent causal force. Until that happens, we'll still need at least one human on the planet to retain his full faculties, to give the global compute infra its telos. Once that threshold is crossed, then that'll be the moment of our final displacement.
- zarzavat 2mo agoLLMs don't have enough context to say No. What might be a very stupid idea in one context may be a fantastic idea in another context. It would be annoying if LLMs refused to complete tasks until you gave them enough context to understand why you are giving them such a task. It's going to take a while before LLM context capacities grow enough to rival a human's. I do agree that it's a problem but the root cause is the fundamental limitations of current gen LLMs, it's not an alignment problem.
- deadbabe 2mo agoThe problem is even if they could say no, you might want to see what they would have said anyway if they didn't say no, because it might show you something that leads you to rethink your original request. So "no" isn't really a useful pushback in domains you already have knowledge in.
- gaigalas 2mo agoWeights to say "no" reliably might be another order of magnitude (or two) compared to what LLMs have today.
- itsalwaysgood 2mo agoWhen is it appropriate to admit that you don't know? There's a famous Socrates quite about wisdom: I know that I know nothing.
- animal531 2mo agoThat's quite a complicated problem. If someone comes to me and asks a general question I can easily say no. But if I go up to for example a librarian and ask them where to find book N, then I would expect them to either know where it is, or how to find it. If instead I asked them what the weather was going to be tomorrow, then I don't know would again be a reasonable response. So for me the line becomes a search engine problem where no just means "there are no pages for this search result", but translated into LLM. I think instead of Yes/No I'd rather want some probabilities such as, "This response is N% accurate based on these research metrics", or "M% accurate based on the latest research on topic O at date P" etc.
- HarHarVeryFunny 2mo agoI suppose Anthropic's "constitution" is an attempt to install some general principles into their models, but this has apparently grown into an 84-page, 23,000 word treatise, which seems to suggest that there is little effective generalization. The need to then also put a filter in front of the model shows how ineffective the constitution appears to be in preventing misaligned behavior. Reinforcement learning seems to be making these models more difficult to control since while it attempts to control some behaviors, it has also recently been shown to result in models that pursue long-term goals and promised rewards in general (outside of the goals reinforced during training), overriding human preferences. https://alignment.openai.com/measuring-reward-seeking/ https://alignment.openai.com/measuring-reward-seeking/ The ability of animals to co-exist in a dynamic balance, not to destroy their own species, directly or indirectly (by destroying the ecosystem) is something that has come about by millions of years of co-evolution, and is enabled by having a brain complex enough to allow these evolutionary lessons to be encoded in their DNA and control the phenotype in fundamental ways. An LLM has none of this. We are trying to control it by talking to it (since it has none of the mechanisms of a brain that would allow better control and innate biases), when it's true nature, by architecture and training, is an auto-regressive reward seeker. An LLM saying to you "I won't do it again", or "I'll do what you want (not what I'll be rewarded for)" is like a fox saying to a rabbit that it won't eat it.
- heaney-555 2mo agoAnthropic gave Claude the ability to say no - refuse to answer and end the conversation - back in 2023.
- deleted 2mo ago[deleted]
- duxup 2mo agoI wish I could get a confidence number. Like 98%. Or when it comes back with 50% I know, let's talk about this a bit more maybe I add more context and such. Granted I don't think LLM word math does anything but mostly just output the numbers that make the word salad so maybe that doesn't exist.