4 ms·
"EPG takes a step toward agents that are not blank slates but instead know what it means to make progress on a new task, by having experienced making progress o
by edhu2017 8y ago
"EPG takes a step toward agents that are not blank slates but instead know what it means to make progress on a new task, by having experienced making progress on similar tasks in the past."
Can someone explain to me how they take a step? It seems like they just use random search define a loss function for the sub-policy to optimize against. Is it because the loss function is "learned" over the sequence of actions, making it adaptive?