2 ms·
It's very black and white right now (no knowledge or very constrained knowledge). The same problem exists with just plain old parameter learning, with DL models
by mljoe 9y ago
It's very black and white right now (no knowledge or very constrained knowledge). The same problem exists with just plain old parameter learning, with DL models having so many free parameters it can make training more computationally expensive then it needs to be. For instance, I want to train a reinforcement learner to do some task in the real world (eg. a robot). It would be nice to be able to define a prior that puts very low to no probability on actions that are not physically possible. This would constrain the search space considerably. But it's not really clear how to do this with neural nets.
- gwern 9y ago> It would be nice to be able to define a prior that puts very low to no probability on actions that are not physically possible. This would constrain the search space considerably. But it's not really clear how to do this with neural nets. I dunno, I can think of several ways off the top of my head: 1. put large negative rewards on impossible actions; 2. only compute Q-values for feasible actions (just because DQN computes Q-values for a hardcoded set of actions doesn't mean you have to); 3. use Achiam's "Constrained Policy Optimization" https://arxiv.org/abs/1705.10528 https://arxiv.org/abs/1705.10528 ; 4. reparameterize the output to make illegal actions unrepresentable (domain-specific); 5. add illegal or not as an additional data input (ie in addition to the usual image or robot state, for the _n_ possible actions, include a _n_-long bitvector of possible/impossible), or include that as a regression target to make it predict whether each action is possible.