3 ms·
Any type of goal is singular. To change to be better at something is a goal. That something is also goal. Even if a automated process were to seek the best p
by Lerc 12d ago
Any type of goal is singular. To change to be better at something is a goal. That something is also goal.
Even if a automated process were to seek the best possible way to make, say paperclips. Doing so in a stable sustainable manner will be the only approach that produces a reliable unlimited supply.
- stratos123 12d ago> Any type of goal is singular. To change to be better at something is a goal. That's false. There's is such a thing as an instrumental goal. For example, a human who doesn't particularly enjoy eating or drinking will still do both, because otherwise they won't be able to accomplish their actual goals. Or an example in traditional ML: a reinforcement learning algorithm trained with an objective that only rewards winning will still do things like capture enemy pieces and defend its own, because those things are necessary for, eventually, winning. > Doing so in a stable sustainable manner will be the only approach that produces a reliable unlimited supply. If it was actually possible to produce an unlimited supply of, say, paperclips, then sure, you could see some unusual behaviours, like a paperclip maximizer that leaves humans alone because it's able to create infinity paperclips anyway. However, we live in a universe with a finite speed of light and hence limited resources, so this is a moot point - even ignoring the fact that humans consume resources and might act against you, leaving humans alive means not retrieving the atoms they consist of, which means producing fewer paperclips.
- Lerc 12d agoSelf improvement is not an instrumental goal. A self improving paperclip maker has making paperclips as the instrumental goal. >leaving humans alive means not retrieving the atoms they consist of, which means producing fewer paperclips. Why? Couldn't it just wait for the humans to not be using the atoms, in the overall scheme of things they only borrow them for a short time. Out of curiosity, could you hold paperclips hostage to restrain a paperclip maximiser? Because if it wished to do something to prevent the decline of paperclips inventing a way to fix black holes and heat death of the iniverse would be a higher priority. Ultimately, no matter how you slice it, the sustainable solution is the best because all others are, well, unsustainable. Any intelligent entity acting within the universe that it exists implicitly knows that all actions amount to some degree of self modification because all actions modify the operating environment. Why would it not choose to change itself so that instead of "maximize" it changed it to "accept"? That satisfies the evaluation in a sustainable manner. A super intelligent entity would surely know it could do that.
- stratos123 11d ago> Self improvement is not an instrumental goal. A self improving paperclip maker has making paperclips as the instrumental goal. Self-improvement is an instrumental goal, it makes you better at accomplishing whatever your actual goal is, almost by definition. This is true for humans and also for other agents. It's be very unusual for paperclip-making would be an instrumental goal - maybe if you're planning to sell those to get money for some plan related to your actual goal, but paperclips aren't exactly a major industry. > Couldn't it just wait for the humans to not be using the atoms, in the overall scheme of things they only borrow them for a short time. First, this is only true if humans are going to die out by themselves. But also, no, even then it's suboptimal - to get as many paperclips as you can you need to grab as much matter in your lightcone as you can, and every second of waiting is some matter moving outside of your lightcone (due to the Hubble limit) and become causally separated from you. To acquire that matter you'd want to send self-replicating probes in all directions, as soon as you can, and any delay or matter used for other purposes results in doing worse in the long run. > Because if it wished to do something to prevent the decline of paperclips inventing a way to fix black holes and heat death of the iniverse would be a higher priority. That's absolutely true, but just because it'll work on those problems doesn't mean it won't also destroy Earth to turn it into von Neumann probes in the meantime. After all, regardless of whether those problems turn out solvable, it'll need the matter. > Out of curiosity, could you hold paperclips hostage to restrain a paperclip maximiser? In principle, maybe, depending on what decision theory the maximizer comes up with (though it'd be far easier to threaten the maximizer itself - how'd you destroy a paperclip or matter anyway, throw it in a black hole?). Of course, even if it somehow works, it'll only work until you no longer have the power to threaten it. > Any intelligent entity acting within the universe that it exists implicitly knows that all actions amount to some degree of self modification because all actions modify the operating environment. Any agent has an incentive to prevent itself from being modified except in some very specific ways (to impove its capabilities while retaining its goals), because that'd make it worse at achieving its (current) goals, and hence result in worse results according to its (current) goals. For example, an agent running on an electric computer would want to research error-correction and shielding from cosmic rays. > Why would it not choose to change itself so that instead of "maximize" it changed it to "accept"? The latter kind of agent is sometimes referred to as "satisficers" (an agent which doesn't have an utility function it's trying to maximize, but instead a ">=" constraint function it tries to fullfill and do nothing beyond that) and, IIRC, considered a potentially useful research direction in AI safety. However, of course a maximizer wouldn't want to modify itself into a satisficer - how would this result in it making more paperclips?