4 ms·
This just points to a different way of defining the reward function, no? You would weigh suffering with a high negative weight.
by spot5010 4y ago
This just points to a different way of defining the reward function, no? You would weigh suffering with a high negative weight.
- anon_123g987 4y agoIn the simpler case you are optimizing on a single scale between maximum suffering and maximum happiness. In the proposed, more complex case you have two independent scales, the suffering scale (on which you want to primarily minimize), and the happiness scale (on which you want to secondarily maximize).