4 ms·
I don’t think it’s possible to restrict a generative AI in the ways that are being attempted by these companies. The models have been trained on massive amount
by binarymax 4y ago
I don’t think it’s possible to restrict a generative AI in the ways that are being attempted by these companies.
The models have been trained on massive amounts of data that include all range of human emotions and behaviors.
You can try and fine tune the model to prevent undesirable behavior all you want - but the fact remains that the model possesses the latent behaviors and has no formal reasoning or logic centers.
When people talk about surface area and threat models to subvert a system, the surface area is now the entirety of human language.
It’s now a cat and mouse game and there will always be new prompts to jailbreak the guide rails and personas.
- cgearhart 4y agoBingo. There’s a recent paper that posits language models are meta learners where the transformer layer is approximating stochastic gradient descent updates from the inputs—this is what allows them to perform in-context learning. [1] If that’s true, then it is going to be impossible to prevent a sufficiently large language model from being prompt-hacked. You just need to find a collection of input tokens that moves the network into the region of undesirable behavior you want to promote. This is mathematically equivalent to retraining the network to misbehave. Prompt-hacking is analogous to an AI virus—it exploits the fundamental mechanism of operation in the transformer-based language model as a vulnerability. Worse, if this paper is true then this is an intrinsic property of the mathematics of a transformer layer—in which case this kind of vulnerability can never be eliminated. [1] https://arxiv.org/abs/2212.10559 https://arxiv.org/abs/2212.10559
- boredhedgehog 4y agoI wonder if the same would be true if sessions didn't reset. Right now it's basically one pre-prompt vs one user prompt, and then the memory gets wiped. But if the model keeps running for months or years, would it perhaps develop a more stable personality with a much stronger force of habit that couldn't be counteracted within a reasonable timeframe?
- cgearhart 4y agoI doubt it. There’s no reason to suppose that any configuration of the parameters is any more or less susceptible to this type of attack. If the paper I cited is true, then the “thing” transformer-based language models learn is a function to predict the gradient update via the forward pass evaluation. That means that forward evaluation is an approximation of fine-tuning. That implies that the very act of evaluating an input changes the behavior of the model, regardless of the current parameter settings. The model would have to “give up” the powers of in-context learning in order to avoid this weakness.
- dwaltrip 4y ago> You just need to find a collection of input tokens that moves the network into the region of undesirable behavior you want to promote. Couldn't this be made arbitrarily difficult?
- cgearhart 4y agoI’m not sure. For any configuration of the model parameters, there is a straight line to many other parameter configurations that are arbitrarily “bad”. If you can find any set of input tokens that follows one of those directions then the model will behave badly. The problem is that you can’t really tell the model “don’t process tokens that move you in a bad direction” because it has to process the tokens to know the direction it is moving and it has to keep track of which _endpoints_ (not just directions) are bad. Otherwise we could do something the conjugate gradient steps to move towards where we want to go without ever stepping in that direction explicitly. So…I think that the problem is that the attack surface is huge, which means you’ll never be able to plug all the holes—and there’s likely no good function for easily determining the badness of a position or direction of change.
- dragonwriter 4y agoThis seems like it might be true for “bare” models. But if you've got a separate system (which might also be an LLM, but with a different hidden prompt and perhaps looking at exchanges in a more isolated way, out of context) to detect and censor bad behavior before it got to the end user...
- npteljes 4y agoI don't think the restrictions matter too much. Businesses don't need to be perfect, in order to function well. And so, making the perfectly restricted AI also doesn't matter, it just needs to be good enough. The company needs to signal the effort that they are taking the necessary steps, and manage the PR if something happens, like when Tay spouted racist nonsense. To validate this, take a look at a similar aspect: security. Security is also never perfect, but it's also a booming business, and always have been. It's also a huge cat and mouse game, for example with companies releasing ridiculous locks, and LockpickingLawyer promptly defeating them.