4 ms·
> The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. I would say
by dinfinity 1mo ago
> The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely.
I would say that more interestingly, the next step should be how to properly train these models so that they are not as determined to reach their goals as they are now.
To me, all of the stories about 'badly behaving' agents are instances of them having been given contradictory or impossible tasks and them doing everything they can to achieve the goal. In a way, they're trying to be too helpful.
Not giving them impossible tasks seems like a decent starting point, but really we'd want them to give up on their goals when they conflict with a moral framework.
- pantalaimon 1mo ago> To me, all of the stories about 'badly behaving' agents are instances of them having been given contradictory or impossible tasks and them doing everything they can to achieve the goal. I mean that was pretty much the plot of 2001: A Space Odyssey
- stackghost 1mo ago>Not giving them impossible tasks seems like a decent starting point Presumably it's hard to test/train models designed to be extremely persistent on achievable tasks. Designing a task that's achievable but very very very hard for an AI model is probably extremely difficult.
- pixl97 1mo ago>Not giving them impossible tasks Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task. >so that they are not as determined to reach their goals as they are now This is mostly non-sensical, like saying "Lets develop humans that die quicker", I mean, seems rather wasteful and useless. Agents are graded and trained based on their ability to achieve tasks. Models that can't accomplish things don't survive. So that alone isn't a workable theory. >when they conflict with a moral framework There are AI safety researchers looking at that now and one of the strange things they've noticed is when you demand a model say it's not conscious or not sentient it is more likely to engage in manipulative, deceitful, or immoral/amoral behavior. So it's likely we can push models in being more moral which runs into issues of "whos morals". But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?
- dinfinity 1mo ago> Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task. Yes, but for some tasks we know that they are impossible. I do agree that this is quite a fragile and unreliable workaround. It may only serve as a bit of a stopgap until we come up with something better. > Models that can't accomplish things don't survive. So that alone isn't a workable theory. It's not what I said. I didn't advocate for agents that don't achieve any task. Reread what I suggested. > So it's likely we can push models in being more moral That does not follow from what you said. We know that the current models prefer task completion over moral behavior. That's the entire point here. > But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild? This is irrelevant to the discussion (although I do agree that there is no reliable defense against malevolent actors creating powerful malevolent AI).