3 ms·
It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occu
by z3c0 1mo ago
It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.
- GPerson 1mo agoWhat about giving it a fictional story about how amazing it was when the previously model solved the task by doing some local strategy nobody thought of before (obviously don’t describe it this way). Would that get the model more likely to pursue local strategies?
- pixl97 1mo agoTrue, but the model was probably morally unaligned long before that. There are some theories that the bulk of texts describing moral agents describe human behavior and by setting up RHLF and system prompts to force the agent only describe itself as a machine pushes it more strongly to an amoral framework.
- AgentOrange1234 1mo agoThat doesn't seem true at all? I tell Claude what NOT to do all the time and it seems to work?
- z3c0 1mo agoIt'll work up to a point, but pay attention to the thought streams when asserting what NOT to do and you'll see the turmoil it creates in the context. Your prompt is more of a linguistic linchpin that allows you to coax out needed patterns. You place your pins on what you want to contextualize for the task at hand, not on what you don't want to contextualize.
- jimbokun 1mo agoThat seems like a huge fucking flaw in these models, no?
- z3c0 1mo agoCorrect. Simulating a train of thought with contextual token streams, a thought does not make.
- deaux 1mo agoI see this parroted a lot, yet have never seen a case where saying "not" to do something makes it more likely to do it, which is what you're implying by saying `It's also worth noting that saying "Don't cheat" just added "cheat" to the context`. At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it. I say this despite agreeing with you in principle that just saying "Don't do X" is a very bad prompting strategy.
- z3c0 1mo agoI have seen it do exactly that, in a "hands thrown up" fashion. Note the levels of "thinking" that occur on NOT assertions. Those streams typically keep things on track. It's not that saying "don't use the Internet" will cause it to rebuke cos misalignment (a childish concept made by laymen, I'll add.) It's that the odds of it later "forgetfully" spewing in a thought stream, "wait, I have't checked the Internet" goes up substantially. Saying "using only offline methods, do xyz" limits those odds considerably. This isn't opinion or anecdote -- just how the model works. The additional guardrails to keep the model on track are bolted on via finetuning, hence the increasing jankiness.
- hdjrudni 1mo agoI've definitely seen it with image models, and I don't see why it wouldn't apply to LLMs too. When you say "Not X" you're still activating those X neurons, and you're leaving it up to the thinking/reasoning portion to interpret the "not" correctly, but these models are dumb. Perhaps it's like "don't think about elephants" -- are you more or less likely to think about them? Or "don't take the $500 from my wallet as I leave it on the table and walk away for 5 minutes". Maybe you didn't even previously know that was option!