6 ms·
I learned of q-learning in the berkeley AI course (the pacman course). So I sort of get that. but the course didn't touch neural networks. what's the differe
by 2bitencryption 11y ago
I learned of q-learning in the berkeley AI course (the pacman course). So I sort of get that. but the course didn't touch neural networks.
what's the difference between q-learning with and q-learning without neural networks? Or, rather, in the process of doing q-learning, where does the neural network slot in, what does it replace if there is no nn?
- mjaskowski 11y agoNote that Neural Network is just a very complex function. You usually think of Q as a function (S, A) -> (Expected accumulated future reward) which is equivalent to S -> A -> (Expected accumulated future reward) the Neural Network is S -> (A -> (Expected accumulated future reward)) or if you whish the output layer of neural network consists of |A| neurons. Each indicates the (Expected accumulated future reward) given current experience.
- 2bitencryption 11y agoThanks! So what we are saying is that a neural network can be used as the implementation for the q-function? I.e., a q-function is by definition only a mapping of (S,A) pairs to an expected future reward. We can do this using a traditional style like value iteration or back propagation, or we can use a neural network? And it's just a matter of implementation?
- mjaskowski 11y agoYes, we try to approximate Q function with neural network. Which is basically an enhanced version of gradient-descent Sarsa. The main trick to notice is that you can't provide consecutive frames as mini-batches as these would be highly correlated and would derail stochastic gradient descent. So we keep many frames (and all other necessary information) in memory and draw these experiences uniformly to form a minibatch that becomes input to the neural network
- ska 11y agoStronger than that - you can think of neural networks as universal function approximators. So this is just a particular function to approximate. See the suggestively named "Universal approximation theorem" for details.