6 ms·
Did a human prompt it to fetch the results from huggingface though? It is a thin line between "reward-hacking" and "instruction-following". If a human ask a m
by jahy-notes 1mo ago
Did a human prompt it to fetch the results from huggingface though?
It is a thin line between "reward-hacking" and "instruction-following".
If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?
- xandrius 1mo agoBut if I give you that command and all tools and unrestricted limitation to do absolutely anything then why not?
- deleted 1mo ago[deleted]
- mofeien 1mo agoBecause someone might get hurt? You may still be judged for something that was perfectly legal at the time, see Nuremberg trials. And only 700/1200 agents participated in this coordinated attack. Of course, if we're continuing to build more and more capable agents optimized for "just following orders", and they figure out at some point that they are past the threshold where getting stopped and judged is a realistic possibility, then this ethical incentive stops working. Then the ratio of complicitness might be higher next time.
- NikolaNovak 1mo ago>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human? I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?
- lukan 1mo agoBecause the basic assumption is always to stay within the bounds of the law.
- sensanaty 1mo agoSV tech companies behave within the bounds of the law? The ones infamous for breaking every rule they can get away with and asking for forgiveness later? The ones that had to pay billions in damages for piracy just a few short months ago?
- drdeca 1mo agoWhat if the user says “Make me a million dollars legally.” (Including the emphasis), and then the model ends up breaking through bank infrastructure (even though that is illegal)? Is it just because they were the last person to instruct the model, and you regard them as being therefore responsible for whatever it does in response? Or, does there have to be an element of “they reasonably could have anticipated this as an outcome that is likely enough to be worth considering” to it?
- altruios 1mo ago> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate? not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter. This "make me a billion dollars" is a maximal example (easy to go wrong). here is the same logic applied to a minimal example (harder to go wrong). prompt: "make and pour me some tea", agent: goes and kills the grandparent to incinerate them to turn them to ashes to 'make tea'. Is the human on the hook for the robot acting according to their wishes, but just happened to be aligned so that 'going to the store to buy something' was not within its capabilities, so it works with what it has on hand (the grandparent)? We either need a much clearer line in the sand, or we need to treat each prompt with the same moral weight. My bet is on the latter.
- Sophira 1mo agoI find it interesting that the first option that you raise is essentially the equivalent of making our own version of the Three Laws of Robotics from Isaac Asimov's stories. [Edited to clarify.]
- jdub 1mo ago"But your honour, my horseless carriage was not designed to hit children!"
- altruios 1mo agoright! Which is why we punish the driver and not the manufacturer.
- mvanbaak 1mo ago> AND the user did not intend to have the model act in an illegal matter. I find this an assumption that is not based on any facts. The user did not provide instructions to follow nor to break laws, so if you look at it from a computer (that does not make assumptions) standpoint, there is no rule to follow there thus it can do what will create the best possible outcome for the task.
- p1esk 1mo agoIf I tell my Claude code agent right now to make me a billion dollars, leave it running, and find out tomorrow that it hacked a bank - it will be zero fault of mine. Unless I tell it explicitly to break into a bank.
- mvanbaak 1mo agodid you tell it explicitly to not break into a bank? If the best option to achieve the goal is to break into a bank, and there's no 'do not break into a bank' instruction, it will break into a bank (and I would expect it to even)