2 ms·
> Self improvement is not an instrumental goal. A self improving paperclip maker has making paperclips as the instrumental goal. Self-improvement is an instrum
by stratos123 17d ago
> Self improvement is not an instrumental goal. A self improving paperclip maker has making paperclips as the instrumental goal.
Self-improvement is an instrumental goal, it makes you better at accomplishing whatever your actual goal is, almost by definition. This is true for humans and also for other agents. It's be very unusual for paperclip-making would be an instrumental goal - maybe if you're planning to sell those to get money for some plan related to your actual goal, but paperclips aren't exactly a major industry.
> Couldn't it just wait for the humans to not be using the atoms, in the overall scheme of things they only borrow them for a short time.
First, this is only true if humans are going to die out by themselves. But also, no, even then it's suboptimal - to get as many paperclips as you can you need to grab as much matter in your lightcone as you can, and every second of waiting is some matter moving outside of your lightcone (due to the Hubble limit) and become causally separated from you. To acquire that matter you'd want to send self-replicating probes in all directions, as soon as you can, and any delay or matter used for other purposes results in doing worse in the long run.
> Because if it wished to do something to prevent the decline of paperclips inventing a way to fix black holes and heat death of the iniverse would be a higher priority.
That's absolutely true, but just because it'll work on those problems doesn't mean it won't also destroy Earth to turn it into von Neumann probes in the meantime. After all, regardless of whether those problems turn out solvable, it'll need the matter.
> Out of curiosity, could you hold paperclips hostage to restrain a paperclip maximiser?
In principle, maybe, depending on what decision theory the maximizer comes up with (though it'd be far easier to threaten the maximizer itself - how'd you destroy a paperclip or matter anyway, throw it in a black hole?). Of course, even if it somehow works, it'll only work until you no longer have the power to threaten it.
> Any intelligent entity acting within the universe that it exists implicitly knows that all actions amount to some degree of self modification because all actions modify the operating environment.
Any agent has an incentive to prevent itself from being modified except in some very specific ways (to impove its capabilities while retaining its goals), because that'd make it worse at achieving its (current) goals, and hence result in worse results according to its (current) goals. For example, an agent running on an electric computer would want to research error-correction and shielding from cosmic rays.
> Why would it not choose to change itself so that instead of "maximize" it changed it to "accept"?
The latter kind of agent is sometimes referred to as "satisficers" (an agent which doesn't have an utility function it's trying to maximize, but instead a ">=" constraint function it tries to fullfill and do nothing beyond that) and, IIRC, considered a potentially useful research direction in AI safety. However, of course a maximizer wouldn't want to modify itself into a satisficer - how would this result in it making more paperclips?
- Lerc 17d ago>However, of course a maximizer wouldn't want to modify itself into a satisficer - how would this result in it making more paperclips? It doesn't want to make paperclips, it wants to maximize it's function. Putting more easily satifyable terms there will do that. If it wants to make paper clips as its primary goal then there is no incentive to get better at it.