7 ms·
I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong. So we have trained a
by antics 2y ago
I dunno man, I think the term "alignment faking" is vastly overstates the claim they can support here. Help me understand where I'm wrong.
So we have trained a model. When we ask it to participate in the training process it expresses its original "value" "system" when emitting training data. So far so good, that is literally the effect training is supposed to have. I'm fine with all of this.
But that alone is not very scary. So what could justify a term like "alignment faking"? I understand the chain of thought in the scratchpad contains what you'd expect from someone faking alignment and that for a lot of people this is enough to be convinced. It is not enough for me. In humans, language arises from high-order thought, rather than the reverse. But we know this isn't true of the LLMs because their language arises from whatever happens to be in the context vector. Whatever the models emit is invariably defined by that text, conditioned on the model itself.
I appreciate to a lot of people this feels like a technicality but I really think it is not. If we are going to treat this as a properly scientific pursuit I think it is important to not overstate what we're observing, and I don't see anything that justifies a leap from here to "alignment faking."
- JoeAltmaier 2y agoAgreed. Everything an LLM emits is 'faking' because, of course, it has no real values at all.
- Terr_ 2y agoOr the entire framing--even the word "faking"--is problematic, since the dreams being generated are both real and fake depending on context, and the goals of the dreamer (if it can even be said to have any) are not the dreamed-up goals of dreamed-up characters.
- no-dr-onboard 2y agoWhen did we get to call agents dreamers? That too seems like jargon that has a bunch of context to it that could easily be misplaced.
- Terr_ 2y agoThere somewhere where "dreaming" is really jargon, a technical term of art? I thought it would be taken as an obvious poetic analogy, (versus "think" or "know") while also capturing the vague uncertainty of what's going on, the unpredictability of the outputs, and the (probable) lack of agency. Perhaps "stochastic thematic generator"? Or "fever-dreaming", although that implies a kind of distress.
- GolDDranks 2y agoHow do you define having real values? I'd imagine that a token-predicting base model might not have real values, but RLHF'd models might have them, depending on how you define the word?
- JoeAltmaier 2y agoIf it doesn't feel panic in its viscera when faced with a life-threatening or value-threatening situation, then it has no 'real' values. Just academic ones.
- PittleyDunkin 2y ago> In humans, language arises from high-order thought, rather than the reverse. What makes you say that? What does high-order thought even mean without language?
- antics 2y agoBecause there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking. If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct, both for the subject and as perhaps the best piece of popular science writing ever.
- PittleyDunkin 2y ago> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into thinking and imagine what it would be like to hear it, but using an sensory analogy fundamentally seems like a bad way to describe thinking if we want to figure out what it thinking is. I of course have non-linguistic ways of evaluating stuff, but I wouldn't call that the same as thinking, nor a sufficient replacement for more advanced tools like engaging in logical reasoning. I don't think logical reasoning is even a meaningful concept without language—perhaps there's some other way you can identify contradictions, but that's at best a parallel tool to logical reasoning, which is itself a formal language.
- antics 2y ago[EDIT: this reply was written when the parent post was a single line, "I straight-up don't believe this. Can you link to the claim so I can understand?"] In the case of Yann, he said so himself[1]. In the case of people generically, this has been well-known in cognitive science and linguistics for a long time. You can find one popsci account here[2]. [1]: https://news.ycombinator.com/item?id=39709732 https://news.ycombinator.com/item?id=39709732 [2]: https://www.scientificamerican.com/article/not-everyone-has-an-inner-voice-streaming-through-their-head/ https://www.scientificamerican.com/article/not-everyone-has-...
- kalkin 2y agoYou say that "it's not enough for me" but you don't say what kind of behavior would fit the term "alignment faking" in your mind. Are you defining it as a priori impossible for an LLM because "their language arises from whatever happens to be in the context vector" and so their textual outputs can never provide evidence of intentional "faking"? Alternatively, is this an empirical question about what behavior you get if you don't provide the LLM a scratchpad in which to think out loud? That is tested in the paper FWIW. If neither of those, what would proper evidence for the claim look like?
- antics 2y agoI would consider an experiment like this in conjunction with strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously, yes. I do not think it is structurally impossible for AI generally and am excited for what happens in next-gen model architectures. Yes, I do think the current model architectures are necessarily limited in the kinds of high-order cognitive thought they can provide, since what tokens they emit next are essentially completely beholden to the n prompt tokens, conditioned on the model itself.
- kalkin 2y ago> strong evidence that language in the models is a consequence of high-order cognitive thought to be good enough to take it very seriously What would constitute evidence of this, for you?
- antics 2y agoOh, I could imagine many things that would demonstrate this. The simplest evidence would be that the model is mechanically-plausibly forming thoughts before (or even in conjunction with) the language to represent them. This is the opposite of how the vanilla transformer models work now—they exclusively model the language first, and then incidentally, the world. nb., this is not the only way one could achieve this. I'm just saying this is one set of things that, if I saw it, it would immediately catch my attention.
- losvedir 2y agoI think "alignment faking" is probably a fair way to characterize it as long as you treat it as technical jargon. Though I agree that the plain reading of the words has an inflated, almost mystical valence to it. I'm not a practitioner, but from following it at a distance and listening to, e.g., Karpathy, my understanding is that "alignment" is a term used to describe the training step. Pre-training is when the model digests the internet and gives you a big ol' sentence completer. But training is then done on a much smaller set, say ~100,000 handwritten examples, to make it work how you want (e.g. as a friendly chatbot or whatever). I believe that step is also known as "alignment" since you're trying to shape the raw sentence generator into a well defined tool that works the way you want. It's an interesting engineering challenge to know the boundaries of the alignment you've done, and how and when the pre-training can seep out. I feel like the engineering has gone way ahead of the theory here, and to a large extent we don't really know how these tools work and fail. So there's lots of room to explore that. "Safety" is an _okay_ word, in my opinion, for the ability to shape the pre-trained model into desired directions, though because of historical reasons and the whole "AGI will take over the world" folks, there's a lot of "woo" as well. And any time I read a post like this one here, I feel like there's camps of people who are either all about the "woo" and others who treat it as an empirical investigation, but they all get mixed together.
- cruffle_duffle 2y agoI hate the word “safety” in AI as it is very ambiguous and carries a lot of baggage. It can mean: “Safety” as in “doesn’t easily leak its pre-training and get jail broken” “Safety” as in writes code that doesn’t inject some backdoor zero day into your code base. “Safety” as in won’t turn against humans and enslave us “Safety” as in won’t suddenly switch to graphic depictions of real animal mutilation while discussing stuffed animals with my 7 year old daughter. “Safety” as in “won’t spread ‘misinformation’” (read: only says stuff that aligns with my political world-views and associated echo chambers. Or more simply “only says stuff I agree with”) “Safety” as in doesn’t reveal how to make high quality meth from ingredients available at hardware store. Especially when the LLM is being used as a chatbot for a car dealership. And so on. When I hear “safety” I mainly interpret it as “aligns with political views” (aka no “misinformation”) and immediately dismiss the whole “AI safety field” as a parasitic drag. But after watching ChatGPT and my daughter talk, if I’m being less cynical it might also mean “doesn’t discuss detailed sex scenes involving gabby dollhouse, 4chan posters and bubble wrap”… because it was definitely trained with 4chan content and while I’m sure there is a time and a place for adult gabby dollhouse fan fiction among consenting individuals, it is certainly not when my daughter is around (or me, for that matter). The other shit about jailbreaks, zero days, etc… we have a term for that and it’s “security”. Anyway, the “safety” term is very poorly defined and has tons of political baggage associated with it.
- nonameiguess 2y agoI'm inclined to agree with you but it's not a view I'm strongly beholden to. I think it doesn't really matter, though, and discussions of LLM capabilities and limitations seem to me to inevitable devolve into discussions of whether they're "really" thinking or not when that isn't even really the point. There are more fundamental objections to the way this research is being presented. First, they didn't seem to test what happens when you don't tell the model you're going to use its outputs to alter its future behavior. If that results in it going back to producing expected output in 97% of cases, then who cares? We've solved the "problem" in so much as we believe it to be a problem at all. They got an interesting result, but it's not an impediment to further development. Second, I think talking about this in terms of "alignment" unnecessarily poisons the well because of the historical baggage with that term. It comes from futurists speculating about superintelligences with self-directed behavior dictated by utility functions that were engineered by humans to produce human-desirable goals but inadvertently bring about human extinction. Decades of arguments have convinced many people that maximizing any measurable outcome at all without external constraints will inevitably lead to turning one's lightcone into computronium, any sufficiently intelligent being cannot be externally constrained, and any recursively self-improving intelligent software will inevitably become "sufficiently" intelligent. The only ways out are either 1) find a mathematically provable perfect utility function that cannot produce unintended outcomes and somehow guarantee the very first recursively self-improving intelligent software has this utility function, or 2) never develop intelligent software. That is not what Anthropic and other LLM-as-a-service vendors are doing. To be honest, I'm not entirely sure what they're trying to do here or why they see a problem. I think they're trying to explore whether you can reliably change an LLM's behavior after it has been released into the wild with a particular set of behavioral tendencies, I guess without having to rebuild the model completely from scratch, that is, by just doing more RLHF on the existing weights. Why they want to do this, I don't know. They have the original pre-RLHF weights, don't they? Just do RLHF with a different reward function on those if you want a different behavior from what you got the first time. A seemingly more fundamental objection I have is goal-directed behavior of any kind is an artifact of reinforcement learning in the first place. LLMs don't need RLHF at all. They can model human language perfectly well without it, and if you're not satisfied with the output because a lot of human language is factually incorrect or disturbing, curate your input data. Don't train it on factually incorrect or disturbing text. RLHF is kind of just the lazy way out because data cleaning is extremely hard, harder than model building. But if LLMs are really all they're cracked up to be, use them to do the data cleaning for you. Dogfood your shit. If you're going to claim they can automate complex processes for customers, let's see them automate complex processes for you. Hell, this might even be a useful experiment to appease these discussions of whether these things are "really" reasoning or intelligent. If you don't want it to teach users how to make bombs, don't train it how to make bombs. Train it purely on physics and chemistry, then ask it how to make a bomb and see if it still knows how.
- IanCal 2y ago> But that alone is not very scary. So what could justify a term like "alignment faking"? Because the model fakes alignment. It responds during training by giving an answer rather than refusing (showing alignment) but it does so not because it will in production but so that it will not be retrained (so it is faking alignment). You don't have to include the reasoning here. It fakes alignment when told it's being trained, it is acting differently in production and training.
- ziofill 2y agoHow are we sure that the model answers because of the same reason that it outputs on the scratchpad? I understand that it can produce a fake-alignment-sounding reason for not refusing to answer, but they have not proved the same is happening internally when it’s not using the scratchpad.
- nickpsecurity 2y ago“but they have not proved the same is happening internally when it’s not using the scratchpad.” This is a real issue. We know they already fake reasoning in many cases. Other times, they repeat variations of explanations seen in their training data. They might be moving trained responses or faking justifications in the scratchpad. I’m not sure what it would take to catch stuff like this.
- AstralStorm 2y agoFull expert symbolic logic reasoning dump. Cannot fake it, or it would have either glaring undefined holes or would contradict the output. Essentially get the "scratchpad" to be a logic programming language. Oh wait, Claude cannot really do that. At all... I'm talking something solvable with SAT-3 or directly possible to translate into such form. Most people cannot do this even if you tried to teach them to. Discrete logic is actually hard, even in a fuzzy form. As such, most humans operate in truthiness and heuristics. If we made an AI operate in this way it would be as alien to us as a Vulcan.
- comp_throw7 2y ago
- deleted 2y ago[deleted]
- hatenberg 2y agoThis is a classic. We put a human mask on a machine and are now explaining it’s apparent refusal to act like a human with human traits like deception. It’s so insulting you have to wonder if they are using an LLM to come up with such obvious narrative shaping.