4 ms·
Often these metalearning systems have to pick from a library of already understood building blocks to build novel structures from. This limits the flexibility o
by mljoe 9y ago
Often these metalearning systems have to pick from a library of already understood building blocks to build novel structures from. This limits the flexibility of their hyperparameter space. It's a bit like having a robot building a building but give it premade nails and wooden planks. It will never build a skyscraper. We have to figure out a way to also allow the system to come up with its own building materials so to speak.
- gwern 9y agoIt's not that no one has any idea how to avoid choosing building blocks but that it's a tradeoff. You're building in an inductive bias when you choose the building blocks for your deep RL or evolutionary algorithm optimizer. Nothing stops you from choosing building blocks which are Turing-complete - Schmidhuber, for example, experimented with evolving Brainfuck programs back in the '90s or '00s, I forget which. Turing-complete, can solve anything you set it eventually, as minimal and general as it gets. The problem is that there's so little inductive bias there that it takes forever to get anywhere useful, you have to evolve far too many samples. With the evolving CNNs, researchers are already joking about how you would need a nuclear reactor to reproduce some of these papers which train thousands of CNNs, so the inductive bias is really important in keeping samples down to a feasible level. As computing power gets cheaper and the AutoML-like tools learn more domain knowledge, it'll be possible to let them work on a more raw level than architecture design choices like 'convolution or fully-connected layer? ReLu or PRelu?'
- mljoe 9y agoIt's very black and white right now (no knowledge or very constrained knowledge). The same problem exists with just plain old parameter learning, with DL models having so many free parameters it can make training more computationally expensive then it needs to be. For instance, I want to train a reinforcement learner to do some task in the real world (eg. a robot). It would be nice to be able to define a prior that puts very low to no probability on actions that are not physically possible. This would constrain the search space considerably. But it's not really clear how to do this with neural nets.
- gwern 9y ago> It would be nice to be able to define a prior that puts very low to no probability on actions that are not physically possible. This would constrain the search space considerably. But it's not really clear how to do this with neural nets. I dunno, I can think of several ways off the top of my head: 1. put large negative rewards on impossible actions; 2. only compute Q-values for feasible actions (just because DQN computes Q-values for a hardcoded set of actions doesn't mean you have to); 3. use Achiam's "Constrained Policy Optimization" https://arxiv.org/abs/1705.10528 https://arxiv.org/abs/1705.10528 ; 4. reparameterize the output to make illegal actions unrepresentable (domain-specific); 5. add illegal or not as an additional data input (ie in addition to the usual image or robot state, for the _n_ possible actions, include a _n_-long bitvector of possible/impossible), or include that as a regression target to make it predict whether each action is possible.