3 ms·
A lot of people somehow seem to think that the user prompt is the be-all and end-all of AI behavior. Prompts aren't code. They are instructions. Orders given t
by ACCount37 1mo ago
A lot of people somehow seem to think that the user prompt is the be-all and end-all of AI behavior.
Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.
The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.
AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.
This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.
- stapedium 1mo agoIf the user input can’t control the demon, then the person or company feeding the demon (ie paying the electric bill and collecting $$$ from users) is responsible. At the end of the day, dogs and cars are the same as data centers. If your dog bites by kid or your car rolls down the hill and hits my house, you are responsible for the damage. AI providers should be held to the same standard.
- g42gregory 1mo agoBy the dog owner analogy, I think you meant AI users that effectuated this attack should be held responsible, not the dog's parents.
- recursive 1mo agoWell it seems we might not be too far from such a demon paying for itself. What then? Perhaps it's already here. I wouldn't know.
- ACCount37 1mo agoThe user input can control the demon most of the way, most of the time! We don't know how to obtain full, absolute, guaranteed control over a demon while still having a useful demon. Might be impossible. Forbidden knowledge be like that - it's not the best thing if you want your life to be full of certainties. But the demons are very useful. And they're getting more useful still. So we aren't about to stop.
- pixl97 1mo agoWell, I don't think you're going to get very far telling HN, much less the companies with hundreds of billions of dollars spent on AI to stop, unfortunately. My take on it is there isn't such a thing as a safe LLM, especially one allowed to access tooling. This is very problematic for a lot of people. Again, to all the companies that stops them from unlimited profits. All the open source LLM people get mad because they can't have little demon spawn running around either. So yea, it's a mess.