3 ms·
Yes. In particular, it's possible to learn the variance of the return using TD-methods with the same computational complexity as learning the expected value (t
by clickok 9y ago
Yes.
In particular, it's possible to learn the variance of the return using TD-methods with the same computational complexity as learning the expected value (the value function).
See [0] for how to do it via the squared TD-error, or [1] for how to estimate it via the second moment of the return, and my own notes (soon to be published and expanded for my thesis) here [2].
It turns out that identifying states with high variance is a great way of locating model error-- most of the real-world environments are fairly deterministic, so states with high variance tend to be "aliased" combinations of different states with wildly different outcomes.
You can use this to improve your agent via either allocating more representation power to those states to differentiate between very similar ones, or have your agent account for variance when choosing its policy.
For example, Tesla could identify when variance spikes in its model and trigger an alert to the user that they may need to take over.
Additionally, there's work by Bellemare [3] for estimating the distribution of the return, which allows for all sorts of statistical techniques for quantifying confidence, risk, or uncertainty.
---
0. https://arxiv.org/abs/1801.08287 https://arxiv.org/abs/1801.08287
1. https://arxiv.org/abs/1607.00446 https://arxiv.org/abs/1607.00446
2. http://rl.ai/posts/fun-with-the-td-error-part-ii.html http://rl.ai/posts/fun-with-the-td-error-part-ii.html
3. https://arxiv.org/abs/1707.06887 https://arxiv.org/abs/1707.06887