3 ms·
There's a lot details that need to be filled in to get a working algorithm from this idea, though. Like how to properly explore enough of the state space that y
by clickok 9y ago
There's a lot details that need to be filled in to get a working algorithm from this idea, though.
Like how to properly explore enough of the state space that you can estimate the ensuing entropy, and if it's possible to learn the utility and its variance in a sample-efficient manner using an RNN.
It might be better to start off with an environment with unknown dynamics but an exact representation before going full nonlinear function approximation.
Nonetheless it's an interesting post; I like the idea of coming up with abstract "goals" that can be applied to any environment (without having to construct a reward function) that yield complex behavior.
Even if it doesn't do precisely what you want it's useful for exploration and perhaps a good stepping stone towards the desired behavior.
On a related note, I believe you can learn to predict the entropy of a Markov process using reinforcement learning, so it might be possible to extend it towards control.
I wrote up the basic idea: http://rl.ai/posts/generalized-returns-entropy.html http://rl.ai/posts/generalized-returns-entropy.html which argues that if you had some sort of state transition model you could construct a reward function from it, and then learning the value of a state is also the "expected entropy" starting from that state.
The state transition model can itself be learned, so no simulator is required.
The reason I say "I believe" is that this is the product of original research, that is, procrastinating on my thesis.
So there's a risk that I've made an error somewhere or missed prior work.