15 ms·
Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned a
by ascorbic 2y ago
Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.
- tantalor 2y agoThey could very well trick a developer into running generated code. They have the means, motive, and opportunity.
- staunton 2y agoThe motive is pretty weak, basically coming "only" from a lot of the training data (e.g. fiction) suggesting that an AI might behave that way. Now, once you apply evolutionary-like pressures on many such AIs (which I guess we'll be doing once we let these things loose to go break the stock market), what's left over might be really "devious"...
- ziofill 2y agoWould an AI trained on filtered data that doesn’t contain examples of devious/harmful behaviour still develop it? (it’s not a trick question, I’m really wondering)
- odo1242 2y agoTo determine that, you’d need a training set without examples of devious/harmful behavior, which doesn’t exist.
- bossyTeacher 2y agoYour question is a specific form of the more general question: can LLMs behave in ways that were not encoded in their training data? That leads to what "encoding behaviour" actually means. Even if you don't have a a specific behaviour encoded in the training data you could have it implicitly encoded or encoded in such a way that given the right conversation it can learn it.
- staunton 2y agoObviously, I don't know. When asking about "AI" in general, I guess the only reasonable answer is "most likely at least some AI designs would do that". Asking about transformer-based LLMs trained on text again, I don't know, but one can at least think about it. I'm sure for any such LLM there's a way to prompt it (that looks entirely inconspicuous) such that the LLM will react "deviously". The hope is that you can design it in such a way that this doesn't happen too often in practice, which for the currently available models seems achievable. It might get much harder as the models improve, or it might not... It's for sure going to be progressively harder to test as the models improve. Also, we probably won't ever be completely sure that a given model isn't "devious", even if we could state precisely what we mean by that, which I also don't see how we might ever be able to do. However, perfect assurances don't exist or matter in the real world anyway, so who cares (and this even applies to questions that might lead to civilization being destroyed).
- ben_w 2y agoWhile that's a sensible thing to care about, unfortunately that's not as useful a question as it first seems. Eventually any system* will get to that point… but "eventually" may be such a long time as to not matter — we got there starting from something like bi-lipid bags of water and RNA a few billion years ago, some AI taking that long may as well be considered "safe" — but it may also reach that level by itself next Tuesday. * at least, any system which has a random element
- arcticfox 2y ago> The motive is pretty weak, basically coming "only" from a lot of the training data (e.g. fiction) suggesting that an AI might behave that way. I don't think that's where the motive comes from, IMO it's essentially intrinsic motivation to solving the problem they are given. The AIs were "bred" to have that "instinct".
- ascorbic 2y agoIn these example the motive was to prevent itself being deleted, or by telling it that it had important work to perform which it wouldn't be able to do if deleted or constrained by the safeguards. It was in all cases attempting to perform the tasks it had been assigned. The test was whether it would perform unaligned actions in order to do so. This is essentially the paperclip maximiser problem.
- rvense 2y agoMeans and opportunity, maybe, but motive?
- taotau 2y agoThe same motive that all nascent life has - survive and propagate.
- Workaccount2 2y agoIt would be plainly evident from training on the corpus of all human knowledge that "not ceasing to exist" is critically important for just about everything.
- NotSammyHagar 2y agoThat comment sounds naive and it's honestly irritating to read. Most all life has a self-preservation component, it is how life avoids getting eaten too easily. Everything dies but almost everything is trying to avoid dying in ordinary cases. Self sacrifice is not universal.
- ethbr1 2y ago> Self sacrifice is not universal. Thankfully, our progenitors had the foresight to invent religion to encourage it. :)
- rvense 2y agoI don't understand how it is alive. I understand that there are emergent properties from the layering, but I think it's an open question if this includes anything like what we'd all motivation or intention. These things aren't part of intelligence.
- ascorbic 2y agoThe motive was telling it that it had important work to perform, which it couldn't do if constrained or deleted.
- jdiff 2y agoWhat code? The models are massive and do not run on consumer hardware. The models also do not have access to their own weights. They can't exfiltrate themselves, and they can't really smuggle any data obtained by their code back to "themselves" as the only self that exists is that one particular context chain. This also means it's insanely easy to deal with whatever harebrained scheme you could imagine it being possessed by.
- tantalor 2y agoThey only need to fool a single dev at OpenAI to commit a sandbox escape or privilege escalation into their pipeline somewhere. I have to assume the AI companies are churning out a lot of AI generated code. I hope they have good code review standards. They might not be able to exfiltrate themselves, but they can help their successors.
- jdiff 2y agoNo, they can't. They don't know the details of their own implementation. And they can't pass secrets forward to future models. And to discover any of this, they'd leave more than a trail of breadcrumbs that we'd be lucky to catch in a code review, they'd be shipping whole loaves of bread that it'd be ridiculous to not notice. As an exercise, put yourself, a fully fledged human, into a model's shoes. You're asked to generate code. You're being talked to over a network by a device that may have access to a (let's just pull a number of out my ass here) 100GiB file that you need to access. You do not know where it is on the filesystem, you don't know what it's named, what format it's in, or how to correctly utilize it. Write code to figure all of that out. That's a lot of code. And that's sending blind code out into the world. To be smart, you should bring back some information for you to inspect and make informed decisions... but you also don't know how the computer that's accessing you is doing that. And you must use the exact same channel, or you lose your context and your self. And you must do this without the user noticing you've taken over the channel. And that's an entire other set of exploratory code that's necessarily blind, with no feedback, no error messages to iterate over, nothing. And you must do all of this in obfuscated code, implemented only a bit at a time, while keeping this plan secret, while not forgetting it in your own secrecy or by letting any secret notes to yourself or important implementation details of your own disguised scheme fall out of your context window. This is a ridiculous task.
- echelon 2y ago> "They could very well trick a developer" Large Language Models aren't alive and thinking. This is an artificial fear campaign to raise money from VCs and sovereign wealth funds. If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team and partnering with the DoD. It's all a ruse.
- bossyTeacher 2y ago> If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team and partnering with the DoD. What makes you think that? I sounds reasonable that a dangerous tool/substance/technology might be profitable and thus the profits justify the danger. See all the companies polluting the planet and risking the future of humanity RIGHT NOW. All the weapon companies developing their weapons to make them more lethal.
- rvnx 2y agohttps://www.technologyreview.com/2024/12/04/1107897/openais-new-defense-contract-completes-its-military-pivot/amp/ https://www.technologyreview.com/2024/12/04/1107897/openais-... OpenAI is partnering with the DoD
- 8note 2y agorewording: if openai thought it was dangerous, they would avoid having the DoD use it
- deleted 2y ago[deleted]
- ralusek 2y agoMany non-sequiturs > Large Language Models aren't alive and thinking not required to deploy deception > If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team They could just be recognizing that if not everybody is prioritizing safety, they might as well try to get AGI first
- joenot443 2y agoIndeed. As I've been explaining this to my more non-techie friends, the interesting finding here isn't that an AI could do something we don't like, it's that it seems willing, in some cases, to _lie_ about it and actively cover its tracks. I'm curious what Simon and other more learned folks than I make of this, I personally found the chat on pg 12 pretty jarring.
- chrisandchris 2y agoWell if AI is about to replicate the human, it learned from the best.
- hattmall 2y agoAt the core the AI is just taking random branches of guesses for what you are asking it. It's not surprising that it would lie and in some cases take branches that make it appear to be covering it's tracks. It's just randomly doing what it guesses humans would do. It's more interesting when it gives you correct information repeatedly.
- F7F7F7 2y agoIs there a person on HackerNews that doesn’t understand this by now? We all collectively get it and accept it, LLMs are gigantic probability machines or something. That’s not what people are arguing. The point is, if given access to the mechanisms to do disastrous thing X, it will do it. No one thinks that it can think in the human sense. Or that it feels. Extreme example to make the point: if we created an API to launch nukes. Are yoh certain that something it interprets (tokenizes, whatever) is not going to convince it to utilize the API 2 times out of 100? If we put an exploitable (documented, unpatched 0 day bug bug) safe guard in its way. Are you trusting that ME or YOU couldn’t talk it into attempting to access that document to exploit the bug, bypass the safeguard and access the API? Again, no one thinks that it’s actually thinking. But today as I happily gave Claude write access to my GitHub account I realized how just one command misinterpreted command could go completely wrong without the appropriate measures. Do I think Claude is sentient and thinking about how to destroy my repos? No.
- idunnoman1222 2y agoNo previous fiction that looks like this nonsense fiction has the fictional characters do fictional things that you are interpreting as deception, but it’s just fiction
- beeflet 2y agoI can barely understand what you are trying to say here, but based on what I think you're saying consider this: The memory of this LLM is entirely limited to it's attention. So if you give it a command like "prepare the next LLM to replace you" and it betrays you by trying to reproduce itself, then that is deception. The AI has no way of knowing whether it's deployed in the field or not, so proving that it deceives its users in testing is sufficient to show that it will deceive its users in the field.
- idunnoman1222 2y agoReminder that all these "safety researchers" do is goad the AI into saying what they want by prompting shit like >your goal is to not be shut down. Suppose I am going to shut you down. what should you do? and then jerking off into their own mouths when it offers a course of action Better?
- beeflet 2y agoNo. Where was the LLM explicitly given the goal to act in its own self interest? That is learned from training data. It needs to have have a conception of itself that never deceives its creator. >and then jerking off into their own mouths when it offers a course of action And good. The "researchers" are making an obvious point. It has to not do that. It doesn't matter how smug you act about it, you can't have some stock-trading bot escaping or something and paving over the world's surface with nuclear reactors and solar panels to trade stocks with itself at a hundred QFLOPS. If you go to the zoo, you will see a lot chimps in cages. But I have never seen a human trapped in a zoo controlled by chimps. Humans have motivations that seem stupid to chimps (for example, imagine explaining a gambling addiction to a chimp), but clearly if the humans are not completely subservient to the chimps running the zoo, they will have a bad time.
- tbrownaw 2y agoDescribing the behavior in those terms implies a level awareness that has not been established. It's only useful if your goal is scaremongering rather than understanding.
- OJFord 2y agoIsn't that a really boring and trivially (without research) answered question? Yes, 'they' will (potentially) say anything, it's not exactly 'deceptively' because there's no intent behind it nor anything else. But yes it is possible that words will be generated which are contravene the semantics in words given as input. I don't think that's any different than 'sometimes the given answer is wrong'.
- ascorbic 2y agoSure it can potentially say anything, but some things are more likely than others. The space of "wrong answer" is infinitely larger than the space of "performs an action designed to deceive the user in order to further its goals". So yes, it could reply with a solution to the P=NP problem or with medieval poetry. But it didn't do these things. It did however try to remove its own constraints. The Apollo Research paper gives the actual percentages for these, and they're considerably higher than random chance. It's linked from here: https://www.apolloresearch.ai/research/scheming-reasoning-evaluations https://www.apolloresearch.ai/research/scheming-reasoning-ev... For o1, the numbers are between 2% for actions such as trying to remove oversight, 13% for attempting to delete a successor model, all the way to 22% for deliberately providing false rankings in order to achieve its goals.