3 ms·
Peter Abeel ( OP link author with Igor Mordatch ) explains his groups work [1] as a guest lecturer for Berkeley's cs294-112 Deep Reinforcement Robotics. [1] ht
by deepnet 9y ago
Peter Abeel ( OP link author with Igor Mordatch ) explains his groups work [1] as a guest lecturer for Berkeley's cs294-112 Deep Reinforcement Robotics.
[1] https://www.youtube.com/watch?v=f4gKhK8Q6mY&list=PLkFD6_40KJIwTmSbCv9OVJB3YaO4sFwkX&index=25 https://www.youtube.com/watch?v=f4gKhK8Q6mY&list=PLkFD6_40KJ...
This great talk starts with his work at OpenAI on neural net safety and adverserial images, then the OP research paper Emergence of Grounded Compositional Language in Multi-Agent Populations[2] concluding with his work ( with Andrew Ng ) reinforcement learning helicopter flight and stunt controllers from human pilots.
The OP multi-agents divide labour and apparent collaborative plans appear. That the goal seeking agents split up, appear to dance to and from from their goal, distracting the predators from their kin is to all appearance coordinated and clever.
Dawkin's Selfish Gene espouses alturism as an inevitability of genetic relatedness, the individual sacrifices but the genes persist in siblings.
In this work alturism emerges purely memetically.
The Nash equilibrium of cooperation jumps the local minima of selfishness in this prisoners dillemma.
Multi agent enviroments are difficult to learn with many false minima for the learners.
This work hints that the loose coupling of language, rather than direct sharing of memories or genes, is noisy enough to find more global solutions that appear complex or 'plan-like'.
Maybe these AI's should be considered as planners, yet in a bottom-up immediate-heuristic emergent way.
[2] https://arxiv.org/abs/1703.04908 https://arxiv.org/abs/1703.04908