4 ms·
I don't know much about RL but I was wondering has anyone tried the opposite? Like have a fixed set of actions, and a fixed/ranged movement speed, then a punish
by ffwd 3y ago
I don't know much about RL but I was wondering has anyone tried the opposite?
Like have a fixed set of actions, and a fixed/ranged movement speed, then a punishment every time it doesn't reach the goal? Or does this not work?
- eigenvalue 3y agoThat’s just a negative reward, so basically the same thing. You train most efficiently with a mix of positive and negative reinforcement, just like with children.
- ffwd 3y agoThanks, I can't write a full reply now (need to think) but for some reason my intuition was, let's say you have a constant punishment signal, and a timer, and if it doesn't solve whatever problem by the time the timer goes down, then it has to find the optimal action set, and if it reaches a goal earlier than the timer then it has to weigh that solution stronger? Like at least how I see with organisms it's about finding the optimal use of the limbs in order to solve an ongoing problem/goal state, and if you accumulate specific actions (instead of one network that optimizes one "space"), and different types of goal states, then it has to find the optimal set of actions to reach the different goals, it just seemed more efficient. But this is off the cuff a bit right now. Edit: i think what i'm thinking of are two "global" numbers. A "closer to the goal" number (higher when closer), and a countdown timer, and then it has to maximize those within the above setting
- wegfawefgawefg 3y agoif you only have negative rewards the bot will usually stop picking any actions at all. bc they all have reward values too low. sometimes a big negstive reward can scare it away from ever pressing the scary button ever again
- quickthrower2 3y agoI don't think these nets are sophisticated enough to understand complexities of what humans call punishment / reward. Therefore punishment = -reward, but this is not the case for humans. If punishment for example is prison, is reward making people "even more free" for example giving them a flying car or something :-). If money is a reward (or punishment a fine), why are people not optimizing every last cent.
- canjobear 3y ago“Reward” in RL is just a real number that tells an agent how well it’s doing. If the reward is negative that could be called punishment. Importantly the agent is choosing actions to maximize (expected estimated) reward, so only the relative reward values matter. So if an agent is choosing among exactly three actions and it knows they’ll give rewards [0, -1, +1], the agent will behave the same as if the rewards were [-100, -101, -99].