3 ms·
off-policy, offline, generative experience replay: these are all efficient ways to maximize our learning with finite experiences to learn from. for example, its
by jonnycomputer 2y ago
off-policy, offline, generative experience replay: these are all efficient ways to maximize our learning with finite experiences to learn from. for example, its hard to learn from rare experiences, and dangerous to learn from very hazardous situations you'd not want to be in.
However, as you pointed out, generative experience replay runs the risk of learning from counterfactual or impossible experiences. To mitigate this risk, we need a discriminator network, not only to distinguish dream from reality but to judge their plausibility (e.g. flying by flapping wings vs dreaming you took a vacation in Tahiti that you never took vs dreaming about forgetting to do your homework).