3 ms·
Is the vision network learning continuously or it has been trained with many configurations of the blocks and gives a continuous output? The post says that the
by a_d 9y ago
Is the vision network learning continuously or it has been trained with many configurations of the blocks and gives a continuous output?
The post says that the imitation network takes the input from the vision network and processes it to infer the intent of the task. Isn't the "intent" always "to stack"? Or can the imitation also be just re-arranging blocks in another configuration?
This part is interesting, if I understood it well..
> "But how does the imitation network know how to generalize?
The network learns this from the distribution of training examples. It is trained on dozens of different tasks with thousands of demonstrations for each task. Each training example is a pair of demonstrations that perform the same task. The network is given the entirety of the first demonstration and a single observation from the second demonstration. We then use supervised learning to predict what action the demonstrator took at that observation. In order to predict the action effectively, the robot must learn how to infer the relevant portion of the task from the first demonstration."
Does this mean that the imitation network has been trained on stacking, un-stacking, throwing...and other such tasks, and then it identifies that "stacking" is what is being done in order to imitate it?
Is there an ELI5 for what the 2 NNs are actually learning?
- npew 9y agoThe vision network is trained before-hand on lots of different configurations in simulation and then used to infer the block locations in the image from the camera. So it’s not learning continously. The imitation network takes the block locations predicted by the vision network, together with the demonstration trajectory in VR, and imitates the task shown in the demonstration. So, it learns to look through the demonstration to decide what action to take next given the current state (i.e. location of blocks and gripper). To keep the setup simple, we only trained the imitation network on stacking tasks (so no unstacking, throwing, etc). In future work, we want to make the setup and tasks much more general.
- a_d 9y agoThanks for the explanation. Can you also explain the significance of "one-shot imitation learning" generally (beyond the context of this experiment)?