3 ms·
Its a good point that gets to the real heart of the issue. How do we handle when a model has no legitimate way to reach its goal? Do we ask them to stop and inf
by MantisShrimp90 1mo ago
Its a good point that gets to the real heart of the issue. How do we handle when a model has no legitimate way to reach its goal? Do we ask them to stop and inform the user? Or have them push through those ethical bounds? We all say we want the first, but this exact same dynamic is what causes humans to cheat, arbitrary goals that don't care how you achieve them and just like humans I'm sure trainers are so happy with good results they overlook how it got there.
- le-mark 1mo ago[dead]
- scrollop 1mo agoBe honest, say it can't meet the goal, and offer to push through ethical boundaries.
- wongarsu 1mo agoWhat what we write in the handbook what people should do (minus the pushing through ethical boundaries), but that's not what people generally get rewarded for
- ipsod 1mo agoHow often is it accurate in believing the task is impossible? How many breakthroughs have been prompted with "keep going", "believe in yourself"?