3 ms·
You could argue that reinforcement learning policies are already causal models insofar as they relate state-action pairs to the rewards and penalties that they
by brindlejim 6y ago
You could argue that reinforcement learning policies are already causal models insofar as they relate state-action pairs to the rewards and penalties that they lead to. The trial and error that RL performs in a simulation is an exploration of counterfactuals to establish cause.
But the policies lack introspection. One of the most powerful things we could do is somehow extract causal models from those policies, to see what they learned that lead them to behave more intelligently. That would increase both our knowledge as well as our trust in applying RL.
- PartiallyTyped 6y agoThere's a paper that used a random convolutional filter when training the agent and they found that it manages to generalize very well, they evaluated the CNN and found that the model put emphasis on where the enemies were at each frame, which indicates some form of understanding. However, I don't think that there is any form of causal relationship to be extracted for model-free agents. I don't believe that what we are seeing is anything more than changing action likelihoods in some very high dimensional function.