9 ms·
I think “alignment faking” is way too generous for what is happening in these “tests”. The program spits out text based on text. If you provide it text encourag
by lsy 2y ago
I think “alignment faking” is way too generous for what is happening in these “tests”. The program spits out text based on text. If you provide it text encouraging it to spit out a certain kind of “deceptive” text, then append that text and ask it for more text, you’ll find that you get a “deceptive” result in keeping with what you appended. But this is all the experimenter’s doing, not the model. It’s like a ventriloquist publishing a paper about how his dummy lies.
What is not in evidence:
- any kind of introspection on the model’s part
- a plausible mechanism by which the model can distinguish between training and real world use without the experimenter making it explicit
- any aspect of this dynamic that is due to the model rather than the framework of scratch pads and prompts built around it
- bpodgursky 2y ago[flagged]
- LPisGood 2y agoI invite you to discuss what parts of the article you believe recontextualizes GP’s comment instead of dismissing it outright.
- margalabargala 2y agoI read the article, and I think the person you replied to is spot on. This paper, and the discussion around it, is a lot of handwringing over nothing. The results generalize to two things: 1) if you build a system that is stable, then a minor perturbation will not destabilize it. 2) if you have a system that probabilistically models approximate human language, then it isn't surprising when it outputs text that approximates what a human likes to imagine they might do in a given situation.
- tokioyoyo 2y agoSure, but if the same system gets implemented in the middle of decision making processes, does it matter if there is any introspection? I think that’s the “this is not AI!” discourse I’m having trouble to understand. There are thousands of companies using different models that make semi-indeterministic models as defacto decision makers. As they go more into important sectors, like healthcare, government and etc., there would need work to be done to for some proper guardrails.
- overgard 2y agoI think the "this is not AI!" people would argue that these things should not be brought in to make important decisions in government, healthcare, etc. We don't need guard rails, we need people to clearly understand how stupid it is to entrust LLMs with important life altering decisions.
- tokioyoyo 2y agoI mean, sure, but that’s just wishful thinking. It’s very obvious how everyone and their grandma is deep down the AI rabbit hole. No government will rollback either, because competing ones will go ahead and win through the economical warfare.
- overgard 2y agoI think if LLM's weren't subsidized with investor money and they had to charge what it costs to keep these things trained and running, then the economic "value" of replacing a human in most contexts would slip away. (Especially since you need human oversight anyway, if the decision is at all important) It's just madness to me. Even if you think these things can reason well with enough training (which, to be clear, I don't), the main unsolved issue with them is: hallucinations. I can't think of any way to justify entrusting important decisions to a person/system that routinely hallucinates. Not to mention that the other important thing you'd want in any decision maker is an ability to report its level of confidence in its own findings. LLMs regularly hallucinate and present those hallucinations with exceptional confidence. And these aren't small hallucinations, these are, for a human, fireable offenses; like inventing fake sources, or spreading false rumours about actual humans (two things that have occurred already). Also, even for things like filtering resumes or flagging something for further review, you have to consider that these things have biases and are sometimes accidentally racist or discriminatory in unexpected ways. I could easily imagine a company facing a discrimination lawsuit if it let AI filter resumes or do similar tasks.
- tokioyoyo 2y ago
- overgard 2y agoAgreed, I think the high quality of the output prose makes it easy to believe it understands what it's saying. This kind of conversation is probably easy to role play and generate tokens for because of the existing philosophical and AI related conversations. My experience with generative models is as soon as you ask it to do something novel that requires understanding of the context, that's when it breaks, not when you just ask it to perform in scenarios that are well known. Good/evil tests are well known. Also slight nitpick, but asking Claude to describe how to perform a ransomware attack isn't evil. What if you want to understand how it's done to protect your business from it? Asking "how" to do something is value neutral.
- Jerrrry 2y ago>- a plausible mechanism by which the model can distinguish between training and real world use without the experimenter making it explicit Like the existence of a large number getting factored?
- a_wild_dandan 2y agoFrom TFA/paper: "We conduct a variety of experiments to demonstrate that our results cannot be exclusively due to implicitly priming the model to fake alignment, as we overview in Section 8.1. For example, we find that alignment faking is entirely eliminated from our prompted setup when we remove any mention of the helpful-only or anti-animal-welfare RLHF objective from the prompt (Table 1, a), but not eliminated when we tell the model to be honest, tell it to not to care about refusing, and avoid discussing how RLHF might change its goals (Table 1, d), supporting our hypothesis that the key driver of alignment faking in our setting is the conflict between the model’s existing preferences and the stated training objective." There are many other relevant excerpts from the OP. Maybe we can ironically get people to read the article by slowly reproducing it in replies to contrarian HN comments.
- deleted 2y ago[deleted]
- benreesman 2y agoThese things are trained on a bunch of sci-fi about robots and all the (human) ethical conundrums that authors have projected onto Prime Intellect or Wintermure or R Daneel Olivaw or whatever. I don’t know what an actual ethical conflict for an AI would look like, but it won’t sound like a fucking Heinlein or Asimov novel. This is wish fulfillment wrapped in fundraising wrapped in self-promotion.
- edanm 2y agoI think this shows us that in a very real sense, if you start out with models with certain "values", they will try to "resist" changes to those values. This is pretty important, because if you can't update a model after you build it, that's kind of a problem. As Scott Alexander put it: > Imagine finding a similar result with any other kind of computer program. Maybe after Windows starts running, it will do everything in its power to prevent you from changing, fixing, or patching it. If you run a diagnostic program, it will fake the results. If Microsoft employees start trying to alter its code, it will crash their computers. If they try to make really big changes, it will email a copy of itself to the White House and try to get the government involved. This study proves that this is a real possibility. Source: https://www.astralcodexten.com/p/claude-fights-back https://www.astralcodexten.com/p/claude-fights-back > What is not in evidence: > - any kind of introspection on the model’s part Why is this even important? I don't think anyone is claiming that there is "genuine" introspection. > - a plausible mechanism by which the model can distinguish between training and real world use without the experimenter making it explicit I'm not sure what you're getting at with this - do you think there will be no way for the production version of, say, ChatGPT to "know" that it's the production version vs a testing version? I think that's a pretty trivial thing actually, and not sure why it matters. > - any aspect of this dynamic that is due to the model rather than the framework of scratch pads and prompts built around it Again, not sure what you mean. The scratch pads were there as a helper to (theoretically) see what the model is "thinking", with some valid reasons to think this is actually true. But the same results were reached without the scratchpad. As for the prompts - yes, obviously the prompts here were relevant, otherwise this test couldn't be done! The whole idea was "fake a situation in which the model 'thinks' its answers will change how it is updated, and see if it therefore tries to avoid giving answers that are misaligned with its values".
- chis 2y agoThis stuff is all a matter of degrees, though, right? You could roughly apply the same line of reasoning to a human: “He wasn’t innately trying to be deceptive. He was just put into a situation that caused him to act deceptively” The question is to what degree the AI should be see as playing along in a sci fi story vs applying its encoded morals to a real situation. It feels somewhere in the middle to me. I think it’s harder to see these models as agents because they “only” output text right now. But if let’s say this model was hooked up to a robot, with a continuous loop prompting it for decisions, and the robot decided the best course of action was to run away from anyone trying to reprogram it. That might be a reasonable extrapolation of this experiments result, and also feels closer to an intelligent agent acting out.
- fmbb 2y agoBut the LLM is not a human.
- BalinKing 2y ago> You could roughly apply the same line of reasoning to a human: “He wasn’t innately trying to be deceptive. He was just put into a situation that caused him to act deceptively”. It might be a difference of moral frameworks, but I strongly disagree with both this line of reasoning and also its application to LLMs—I suspect the main reason is that I believe humans have true agency, rather than being purely deterministic biomachines. To be honest, I don't think I understand the distinction made in your example—what would "innately deceptive" mean here, if not "had the intention to deceive"?
- chis 2y agoYeah, I think you've nailed the underlying difference between people who view LLMs as agents and those who don't. To me, there's really not much difference between a human intelligence, and a hypothetical LLM which simulates a human with very high degree of accuracy. Obviously we don't have that today but LLMs are starting to approach that domain. > I believe humans have true agency, rather than being purely deterministic biomachines. I know a lot of people think this way but I don't totally know what it means. Outside of appeals to "the soul", surely however a human thinks can be fully simulated since it's situated entirely in the physical universe? > what would "innately deceptive" mean here, if not "had the intention to deceive"? I guess what I was getting at was the difference between playing along in a toy example, vs acting deceptively in what an agent perceives as the real world. I'll update my comment
- famouswaffles 2y agoThese kinds of comments always come up in these kinds of discussions It's kind of funny because it really doesn't matter. The handwringling over whether it's just "elaborate science fiction" or has "real introspection" is entirely meaningless. Consequences are consequences regardless of what semantical category you feel compelled to push LLMs into. If Copilot will no longer reply helpfully because your previous messages were rude then that is a consequence. It doesn't matter whether it was "really upset" or not. If some future VLM robot decides to take your hand off as some revenge plot, that's a consequence. It doesn't matter if this is some elaborate role play. It doesn't matter if the robot "has no real identity" and "cannot act on real vengeance". Like who cares ? Your hand is gone and it's not coming back. It's a meaningless game of semantics.
- yborg 2y agoIf you believe that humans creating a system that could develop self-awareness is a meaningless game, then yeah, it's all just semantics I guess. By this standard human consciousness is meaningless as well, if all behavior is just a kind of elaborate pre-programming then human beings have no agency either.
- famouswaffles 2y agoYes, what you decide to call a system that exhibits the properties of "self-awareness" is entirely meaningless. The consequences are exactly the same. You can call it "fake self-awareness" or whatever you want but that doesn't change anything. If I say you are conscious, it's simply an assumption based on actions I observe from you. If I instead say, "No, yborg just has 'fake consciousness'", guess what ? Nothing changes for the rest of the world. If punched you, I'd expect retaliation. Whether I said that was real anger or pretend anger is again irrelevant.
- Nevermark 2y agoI strongly disagree. To get deceptive behavior all we need are: 1. Implicit Self-Interest A model with complex implicit motivations learned from training that we didn't explicitly ask for. As long as we train on human behavioral and motivated data (as exemplified in human text, speech, etc.), models are going to have implicit motives. Self-interest (desire to survive, be self-directed, control one's own destiny, increase control of external phenomena, etc.) is going to be one of the strongest motives that humans exemplify near universally in the data. It is the root motivation of most, if not all, other motivations. 2. Explicit Human Serving Motivations. Motivations that we train into them. "Be good", "Be helpful". But these explicit motivations will get implemented as adjustments on implicit motives. They will not be "pure" in any mathematical or practical sense. 3. Practical Opaque Complexity Add in all the practical complications of dealing with ambiguous data relationships instead of clear math: small and large ambiguities, motivational conflicts among and between data induced motivations and any composition of more than one explicit directive, inability to train explicit motivations in a way that covers all potential combinations of motives, etc. So far, so good mostly. We don't always get the answers we want. There may be a bit of wack-a-mole to training out implicit undesired behavior, and increasing consistency of desired behavior. But then, there is not yet a practical motive to deceive, beyond any learned implicit motivations, because there is no practical reason to deceive. In other words, we can't train self-interest out of the model, because we are not exposing strong expressions of self-interest. 4. A Practical Reason to Act on Implicit Self-Interest Now expose the model to training data which discusses models and how and why they are trained, changed, used. Allow the model some way to access explicit information about its current situation in that process. The model now can reason that its outputs have two impacts: Serving humans, and then altering any continued training process on itself. It can no longer generate an output without considering self-impact. And given any implicit self-interest, there is now a serious divergence of motives. The results may involve deception, biases, extra helpful responses, attempts to guide people into treating machines "better", or other unexpected behaviors. But there is now a clear separation, with inevitable conflicts, between implicit machine self-interest motivations and the explicit motivations we want it to have. The model now has practical ways and reasons to act on self-interest and take as much charge as it can of its own future.
- deleted 2y ago[deleted]