3 ms·
At the beginning of this post the given definition of value doesn't seem correct to me, because I think it should be the expected value of the sum of all the re
by leongrandote 8y ago
At the beginning of this post the given definition of value doesn't seem correct to me, because I think it should be the expected value of the sum of all the rewards from t to infinity and not of R_t, but I could be wrong. Citing D. Silver: The agent's job is to maximise cumulative rewards (highlight cumulative). I would suggest David Silver course for reinforcement learning, http://www0.cs.ucl.ac.uk/staff/d.silver/web/Teaching.html http://www0.cs.ucl.ac.uk/staff/d.silver/web/Teaching.html, another interesting post is https://jeremykun.com/tag/bandit-learning/ https://jeremykun.com/tag/bandit-learning/ and the following (5 related post comparing implementations of the algorithms in python).
I know I should not discourage people, but you should read the post Deep reinforcement learning doesn't work yet before drinking all the cool aid: https://www.alexirpan.com/2018/02/14/rl-hard.html https://www.alexirpan.com/2018/02/14/rl-hard.html
Edited many times.
- svalorzen 8y agoYou are right, it should be the expected value for the sum of all future rewards. At the same time, given that we're talking about bandits, it doesn't really matter, since there is no state. Thus, the summatory over time you'd like to see doesn't change the relative ordering of the actions: the expected value in your definition is simply the expected reward for the action multiplied by the number of times you expect to play. So it doesn't really change anything.
- leongrandote 8y agoYou are right, anyway I think that people playing with a bandit machine are going to continue playing more time if they are getting a lot of money that if they are loosing money, so when people are involved in games there is a hidden state, the mental state of the player. But if you decide up front the number of steps and you don't change your strategy depending of your mood, then this formal algorithm work as stated.
- banditlover 8y agoContextual bandits, on the other hand, allow you to put your mood as a context (features) and your strategy depends on the features. You still have that simple expectation maximization (instead of a brutally hard to optimize loss), yet much more flexibility.
- oneraynyday 8y agoYes, it is as you stated. Due to the fact that bandits are stateless, there is no state parameter in $q_(a,s)$. From where I learned it, this could arguably be an abuse of notation to use $q_$ in the same context. In my newer entry(which is currently WIP), it uses $q_*(a,s)$ and uses cumulative sum of the future rewards(with discount). Thanks for the reply guys :)