3 ms·
There are no infinite rewards in biology and yet mathematicians seem to do just fine answering these sorts of questions. I don’t think you want to encode your
by state_less 5y ago
There are no infinite rewards in biology and yet mathematicians seem to do just fine answering these sorts of questions.
I don’t think you want to encode your problem domain in your reward system. It’d be like asking a logic gate to add when you really should be reaching for an FPU. Maybe I’m missing something though?
- xamuel 5y ago>There are no infinite rewards in biology and yet mathematicians seem to do just fine answering these sorts of questions This is only a problem if you're already assuming we do everything based on our biological reward systems, and in the current context that would be circular reasoning. Imagine the treasury creates a "superdollar", a product which, if you have one, you can use to create any number of dollars you want, whenever you want, as many times as you want. Obviously a superdollar is more valuable than any finite number of dollars, and humans/mathematicians/AGIs would treat it accordingly, regardless of the finiteness of our biological reward systems.
- state_less 5y ago> This is only a problem if you're already assuming we do everything based on our biological reward systems Is there some other way that we are do it beside our biological reward system? It sure looks like we get an apple and not an infinite reward when we pick the right answer to be selecting button B. I understand that might not satisfy you.
- xamuel 5y ago>Is there some other way that we are do it beside our biological reward system? Seems to me that's what this whole paper we're discussing is about. If you're already convinced that there is no other way, then you're basically already agreeing with the paper, "Rewards are enough".