3 ms·
It wasn't clear to me either, and I really dislike when papers expect the reader to 'notice' some rather non-trivial manipulations. So I wrote up my unpacking o
by clickok 9y ago
It wasn't clear to me either, and I really dislike when papers expect the reader to 'notice' some rather non-trivial manipulations.
So I wrote up my unpacking of the equivalence: http://rl.ai/posts/max-entropy-rl-explanation.html http://rl.ai/posts/max-entropy-rl-explanation.html
Please forgive the possibly screwy LaTeX, as I am in a battle with my static site generator.
If you're just wondering about the meaning of (18), it claims that for a given state, the entropy of the original policy plus its expected Q-value is less than the entropy + expected Q-value of the π-tilde policy.
Where they are using Q_{soft}^{\pi} to denote the state-action values (with entropy term) for an arbitrary policy.
In (17) they introduce a new policy, \tilde{\pi} defined through the Q_{soft} values for some policy π, with the probability of taking action 'a' in state 's' proportional to the exponentiated Q(s,a).
- AlexCoventry 9y agoThanks, this is really generous of you.